Panoptic segmentation in forest environments is bottlenecked not by semantic quality but by instance separation; existing unsupervised panoptic approaches produce usable stuff maps but near-zero thing quality. Depth or flow-based instance discovery methods needs sensors that are not always available. We present ProGuT (Prototype Guided Training), which produces panoptic pseudo-labels without per-image training masks, needing only unlabeled images and one-time cluster-to-class mapping. ProGuT clusters CLIP patch features, then recovers trunk instances through multiscale geometric prior that falsifies non-trunk structures via structure-tensor. This is cheap compared to depth, flow or class-supervision methods to create pseudo labels. These are then used for downstream tasks which we evaluate against other unsupervised baselines. ProGuT achieves a Panoptic Quality (PQ) of 65.2 on Our-forest dataset (2.6x improvement over the initial pseudo-label quality) and reaches 65.9 mIoU on Freiburg Forest, outperforming unsupervised baselines like PiCIE (45.3 IoU) and STEGO(57.6IoU). Additionally, ProGuT outperforms existing unsupervised methods for class-agnostic trunk instance benchmark.
Figures & tables
Figure 1 : Results and overview of ProGuT . Top : Given unlabeled forest images, ProGuT produces Semantic, Instance and Panoptic segmentation. Bottom: The ProGuT pipeline. A semantic network clusters CLIP patch features into Spatially Coherent Clusters; followed by a hybrid CRF and cluster-to-class assignment yielding n -class semantic map. In parallel, a trunk-instance branch recovers individual trunks via multi-scale geometric prior. The 2 branches are merged to form panoptic pseudo-labels.
Figure 2 : Output of semantic pipeline (stages 1-3): images are clustered, refined with UNet, and mapped to n -classes, thereby producing n -class semantic map.
Figure 3 : Multi-scale coherence: the coherence map is produced by a structure tensor at each scale, followed by thresholding and intersecting across scales to retain structure that is vertically coherent at every scale.
Figure 4 : Dual-prior trunk instancing. The coherence and appearance priors are projected to per-column trunk-presence signals; the peaks of which are used for SAM2 prompting (+ve peak center, -ve edges) that segment individual trunks. Trunks from both the priors are then pooled and de-duplicated by NMS into final trunk instance pseudo-labels.
Figure 5 : Trunk-instance branch: geometric and appearance prior feed into a shared instancing (column peaks + SAM2), pooled and de-duplicated with NMS into trunk instances, then merged with n -class semantic map to produce panoptic pseudo-label.
Figure 6 : Qualitative comparison of PiCIE [ 5 ] , STEGO [ 13 ] and ProGuT on Freiburg forest dataset [ 29 ] .
Method
Sky
Trail
Grass
Veg .
mIoU
PiCIE [ 5 ]
70.41
17.69
46.68
46.24
45.25
STEGO [ 13 ]
73.50
31.20
59.05
66.53
57.57
ProGuT-UNet
79.23
32.00
54.96
66.54
58.18
ProGuT+DeepLabV3
85.29
38.88
65.51
73.89
65.89
E-Net † [ 25 ]
–
–
–
–
71.40
SegNet † [ 30 ]
–
–
–
–
74.81
Table 1 : Semantic segmentation on freiburg forest (4-class mIoU). All mask-free methods use no manual training masks. † : supervised, shown for reference.
Method
AP
AP50
AP75
MaskCut [ 33 ]
0.00
0.00
0.00
CuVLER [ 1 ]
0.22
0.25
0.25
CutLER [ 33 ]
0.53
0.82
0.52
ProGuT (coherence-only)
13.65
28.46
12.05
ProGuT (dual-pass)
16.60
34.20
14.74
Table 2 : Instance segmentation on CanaTree100 (mean AP over 5-fold CV). No method uses CanaTree100 training data.
Figure 7 : Qualitative evaluation of MaskCut [ 33 ] , CutLER [ 33 ] , CuVLER [ 1 ] and ProGuT on CanaTree100 dataset.
Method
PQ
PQ Th
PQ St
U2Seg [ 24 ]
3.75
0.00
11.33
ProGuT (pseudo-labels)
25.13
17.80
43.16
ProGuT + Mask2Former
65.18
29.21
74.17
Table 3 : Panoptic segmentation on Our forest ( 15 held-out GT images). U2Seg [ 24 ] is the unsupervised panoptic baseline (MaskCut instance discovery + STEGO semantic clustering, Hungarian-matched to Our forest classes). ProGuT pseudo-labels are the direct pipeline output; ProGuT + Mask2Former is the downstream model trained on them.
Figure 8 : Qualitative results of U2Seg [ 24 ] and ProGuT on Our forest dataset
Configuration
PQ
SQ
RQ
PQ Th
PQ St
Semantic only
11.67
61.55
18.97
0.00
28.21
+ coherence
22.99
64.28
35.76
16.67
36.55
+ UNet
20.18
64.18
31.45
14.07
34.33
+ dual-pass
20.69
63.61
32.53
15.15
34.32
+ dual-pass + CRF
25.13
69.53
36.15
17.80
43.16
Table 4 : Pseudo-label ablation on Our forest using 15 GT images. Each row adds one component over the row above.
Figure 9 : Qualitative result of ProGuT on FinnWoodlands [ 20 ] dataset
Method
PQ
PQ Th
PQ St
Ground PQ
semantic-only
9.10
0.00
52.29
79.71
coherence-only
6.09
0.09
52.29
79.71
unet-only
7.00
0.11
52.29
79.71
dual-pass
5.42
0.15
52.29
79.71
ProGuT + Mask2Former
8.17
1.33
53.14
79.71
Table 5 : Panoptic segmentation on FinnWoodlands val set. Pseudo-label modes are evaluated directly against 4-class coarse GT; ProGuT + Mask2Former denotes the downstream model.
We present SegmentAnyTreeV2, a sensor- and platform-agnostic framework for semantic and instance segmentation of forest point clouds. The model combines a serialization-based Point Transformer v3 backbone with a lightweight semantic head and a tree-focused cross-attention mask decoder. Semantic predictions restrict instance decoding to tree-class voxels, while instance-aware query initialization, one-to-many seed supervision, and asymmetric mask scoring improve separation in dense and structurally complex stands. We further introduce FOR-instance v3, an expanded benchmark comprising 427 scenes and 26,496 annotated trees across diverse biomes, forest structures, and LiDAR platforms. On the FOR-instanceV2 test split, SegmentAnyTreeV2 achieves 90.5% precision, 80.2% recall, 85.0% F1, 90.7% coverage, and 87.6% semantic mIoU, outperforming previous learning-based methods in both instance detection and mask completeness. Zero-shot evaluation on independent sites further demonstrates strong cross-domain generalization.
Maciej Wielgosz, Stefano Puliti, Rasmus Astrup
Norwegian Institute of Bioeconomy Research (NIBIO), Høgskoleveien 7, 1433 Ås, Norway
Automated instance segmentation of forest LiDAR point clouds is increasingly critical as forest monitoring moves toward scalable, detailed, 3D measurement. Yet, progress is constrained by label scarcity for tree instances; a single hectare can hold millions of points and hundreds of overlapping, complex crowns, making manual annotation from scratch with raw data laborious and error-prone. Annotations are often corrected from automatic pre-segmentations, but remain costly as these provide no interactive or AI-assisted refinement. Inspired by the promptable paradigm of foundation segmentation models, we propose SelectAnyTree, a promptable instance segmentation model that delineates any individual tree in a 3D forest point cloud from a few clicks. It introduces two key components: Click-to-query prompt encoder and Canopy Height Model (CHM)-guided first prompt. The former turns each click into a single content query, encoding its 3D position and positive/negative polarity together with a pooled local backbone feature. The latter provides treetops as a geometry- and ecologically guided first prompt without any user input. The resulting prompt query is then decoded into one tree mask by a state-space query decoder to efficiently capture long-range context in large-scale forest scenes with linear-time complexity. We evaluate SelectAnyTree in interactive and instance-level settings across seven diverse forest regions and an independent held-out test dataset, demonstrating strong generalization beyond the training domains. It segments a target tree to 78.2 Intersection over Union (IoU) from a single click, 24.8 points above the strongest promptable baseline, and reaches every accuracy target with the fewest clicks, while using far fewer parameters and less inference time than prior promptable models. The source code is available at https://github.com/thanhhff/SelectAnyTree.
Trung Thanh Nguyen, Daniel Lusk, Kilian Gerberding +10
Nagoya University, Japan · University of Freiburg, Germany · University of California, Los Angeles, USA +5
Continual Panoptic Segmentation (CPS) requires methods that can quickly adapt to new categories over time. The nature of this dense prediction task means that training images may contain a mix of labeled and unlabeled objects. As nothing is known about these unlabeled objects a priori, existing methods often simply group any unlabeled pixel into a single "background" class during training. In effect, during training, they repeatedly tell the model that all the different background categories are the same (even when they aren't). This makes learning to identify different background categories as they are added challenging since these new categories may require using information the model was previously told was unimportant and ignored. Thus, we propose a Future-Targeted Contrastive and Repulsive (FuTCR) framework that addresses this limitation by restructuring representations before new classes are introduced. FuTCR first discovers confident future-like regions by grouping model-predicted masks whose pixels are consistently classified as background but exhibit non-background logits. Next, FuTCR applies pixel-to-region contrast to build coherent prototypes from these unlabeled regions, while simultaneously repelling background features away from known-class prototypes to explicitly reserve representational space for future categories. Experiments across six CPS settings and a range of dataset sizes show FuTCR improves relative new-class panoptic quality over the state-of-the-art by up to 28%, while preserving or improving base-class performance with gains up to 4%.
Nicholas Ikechukwu, Keanu Nichols, Deepti Ghadiyaram +1