Robots need to be able to understand their surroundings in order to operate safely and robustly, and to interact with the surrounding environment. Robots deployed in unconstrained real-world scenarios must additionally be able to deal with novel situations and objects that have never been seen before. In this article, we tackle the problem of open-world panoptic segmentation, i.e., the task of discovering new semantic categories and new object instances at test time, while enforcing consistency among the categories that we incrementally discover. We present Con2MAV, a general method for open-world panoptic segmentation. Experiments across a wide range of datasets, from road scenes to underwater environments, highlight its compelling capabilities in open-world segmentation and its competitive performance on known classes. We will open-source the implementation of our approach upon acceptance. In addition, we propose PANIC (Panoptic ANomalies In Context), a benchmark for evaluating open-world segmentation tasks in autonomous driving scenarios. This dataset, recorded with a multi-modal sensor suite mounted on a car, and then manually annotated, provides high-quality, pixel-wise annotations of anomalous objects at both semantic and instance level. PANIC contains 800 images, more than 50 unknown classes, i.e., classes that do not appear in the training set, and over 4,000 object instances, providing a comprehensive benchmark for evaluating open-world segmentation methods in autonomous driving scenarios. We provide competitions for multiple open-world segmentation tasks on a hidden test set. Our dataset and competitions are available at https://www.ipb.uni-bonn.de/data/panic.
Figures & tables
Figure 1 : Our proposed approach, Con2MAV, is able to tackle multiple open-world tasks and segment unknown objects and categories. In the figure, we show predictions on (left to right) SegmentMeIfYouCan ( Chan et al., 2021 ) , SUIM ( Islam et al., 2020 ) , COCO ( Lin et al., 2014 ) , and PANIC (ours).
Figure 2 : Our dataset, PANIC, provides pixel-wise annotations of unknown semantic categories and object instances of RGB images. The images have been recorded with our own inhouse sensor suite ( Vizzo et al., 2023 ) mounted on our vehicle driving in Bonn, Germany. The dataset consists of images collected at different times of day over the span of more than a year.
Figure 3 : A visual breakdown of the four open-world segmentation tasks. Anomaly segmentation segment all anomalous areas as unknown (zebras and lion together). Open-world semantic segmentation separates classes but has no objects (zebras segmented together, lion separate). Open-set panoptic segmentation segments separate objects but has no category information. Open-world panoptic segmentation has both, classes and object information. Exemplary RGB image is generated with perplexity.ai ( Perplexity Deep Research, 2025 ) .
Figure 4 : Our network processes an RGB image via an encoder and three decoders, for semantic segmentation, anomaly segmentation and class-agnostic instance segmentation. The semantic segmentation decoder also builds class descriptors for the known categories. Results are post-processed and yield the final open-world panoptic segmentation result.
Dataset
Images
Semantic Classes
Instances
Hidden Test Set
Val
Test
Fishyscapes Lost-and-Found ( Blum et al., 2019 )
373
1203
N.A.
1864
✓
CAOS BDDAnomaly ( Hendrycks et al., 2022 )
0
810
3
1231
✗
RoadObstacle21 ( Chan et al., 2021 )
0
327
N.A.
388
✓
SegmentMeIfYouCan ( Chan et al., 2021 )
10
100
N.A.
262
✓
PANIC (ours)
131
679
58
4029
✓
Table 1 : Comparison of open-world segmentation datasets. In the semantic classes, “N.A.” means that there is no label.
Figure 5 : Sensor setup we used for recording data for the PANIC dataset. The setup includes four cameras, one GNSS/IMU device, and two 3D LiDARs. For further details, please refer to Vizzo et al. (2023) .
Approach
Pixel-Level
Component-Level
AUPR
FPR95
sIoU
PPV
mF1
Maskomaly
93.4
6.9
55.4
51.2
49.9
RbA
86.1
15.9
56.3
41.4
42.0
ContMAV
90.2
3.8
54.5
61.9
63.6
UNO
96.1
2.3
68.0
51.9
58.9
Con2MAV
90.0
2.7
59.1
68.3
69.4
Table 3 : Anomaly segmentation results on the test set of SegmentMeIfYouCan. Best results are highlighted in bold. More results available on the public leaderboard.
Approach
Pixel-Level
Component-Level
AUPR
FPR95
sIoU
PPV
mF1
ContMAV
91.7
66.4
15.0
72.1
24.2
Con2MAV
95.7
35.3
20.9
64.7
31.2
Table 4 : Anomaly segmentation results on the hidden test set of our dataset, PANIC. Best results are highlighted in bold. Public competition at codabench.org/competitions/4561 .
Approach
IoU u
mIoU u
Train
Motorcycle
Bicycle
Background + cluster
0
32.3
32.8
21.7
ContMAV
62.4
62.2
56.8
60.5
Con2MAV
66.5
64.4
53.8
61.6
Closed-world
72.3
69.3
60.9
67.5
Table 7: Open-world semantic segmentation results on BDDAnomaly. Best results are highlighted in bold.
Approach
IoU u
mIoU u
Human
Wrecks & Ruins
Robot
ContMAV
46.2
36.9
46.2
43.1
Con2MAV
69.4
64.3
53.0
62.2
Closed-world
82.0
71.6
80.2
77.9
Table 8: Open-world semantic segmentation results on the SUIM dataset. Best results are highlighted in bold.
Figure 6 : Qualitative results of our approach, Con2MAV, on open-world semantic segmentation on SUIM (top row) and PANIC (bottom row). The prediction mask is overlayed to the input RGB for clarity. In the prediction, different colors correspond to different predicted classes. We compare our approach, Con2MAV (right), with our old method, ContMAV (center).
Figure 7 : Qualitative results of our approach, Con2MAV, on open-set panoptic segmentation on COCO (top row), and PANIC (bottom row). The prediction mask is overlayed to the input RGB for clarity. In the semantic prediction, the colored area indicates the anomalous region. In the instance prediction, different colors correspond to different instance ids.
Figure 8 : Qualitative results of our approach, Con2MAV, on open-world panoptic segmentation on PANIC. The prediction mask is overlayed to the input RGB for clarity. In the semantic prediction, different colors correspond to different predicted classes. In the instance prediction, different colors correspond to different instance ids. In the top row, we show the segmentation of the unknown parts only. In the bottom row, the complete open-world panoptic segmentation results.
Figure 9 : GPT-4V and PaliGemma results on an image from SegmentMeIfYouCan ( Chan et al., 2021 ) .
Modern vision systems must operate in "open-world" settings, where models must recognize known categories and detect unseen or anomalous content. Conventional semantic segmentation models operate under a "closed-world" assumption, often producing overconfident misclassifications on novel content. We address open-world semantic segmentation, the joint task of segmenting known classes while detecting and grouping novel or anomalous content without additional supervision, by extending a dual-decoder baseline with a third, complementary decoder within a unified encoder-decoder design. The first decoder performs closed-set segmentation using Gaussian prototypes for known categories. The second uses contrastive feature learning to isolate unknown regions in embedding space. The third, our key contribution, is a sensitivity decoder that captures fine-grained texture irregularities and activation instabilities indicative of semantic uncertainty, which neither semantic prototypes nor contrastive norms can reliably detect. The three decoders provide genuinely complementary signals: class-level OOD distance in logit space, global feature energy in embedding space, and local activation instability across encoder scales. Experiments on Cityscapes and BDD-Anomaly show that our method improves anomaly segmentation and novel-class discovery while maintaining competitive closed-set accuracy, with gains of +2.4% AUROC and a 2.5 pp. reduction in FPR@95TPR on BDD-Anomaly over the baseline.
Anastasios Romanos Varvarigos, Nikos Giakoumoglou, Tania Stathaki
Department of Electrical and Electronic Engineering, Imperial College London, London, UK
Open world image segmentation aims to achieve precise segmentation and semantic understanding of targets within images by addressing the infinitely open set of object categories encountered in the real world. However, traditional closed-set segmentation approaches struggle to adapt to complex open world scenarios, while foundation segmentation models such as SAM exhibit notable discrepancies between their strong segmentation capabilities and relatively weaker semantic understanding. To bridge these discrepancies, we propose WOW-Seg, a Word-free Open World Segmentation model for segmenting and recognizing objects from open-set categories. Specifically, WOW-Seg introduces a novel visual prompt module, Mask2Token, which transforms image masks into visual tokens and ensures their alignment with the VLLM feature space. Moreover, we introduce the Cascade Attention Mask to decouple information across different instances. This approach mitigates inter-instance interference, leading to a significant improvement in model performance. We further construct an open world region recognition test benchmark: the Region Recognition Dataset (RR-7K). With 7,662 classes, it represents the most extensive category-rich region recognition dataset to date. WOW-Seg attains strong results on the LVIS dataset, achieving a semantic similarity of 89.7 and a semantic IoU of 82.4. This performance surpasses the previous SOTA while using only one-eighth the parameter count. These results underscore the strong open world generalization capabilities of WOW-Seg. The code and related resources are available at https://github.com/AAwcAA/WOW-Seg-Meta.
Danyang Li, Tianhao Wu, Bin Li +5
2VCIP, CS, Nankai University · 1NKIARI, Shenzhen Futian · 3AAIS, Nankai University +2
Recognizing unknown objects is crucial for safety-critical applications such as autonomous driving and robotics. Open-Set Panoptic Segmentation (OPS) aims to segment known thing and stuff classes while identifying valid unknown objects as separate instances. Prior OPS approaches largely treat known categories as a flat label set, ignoring the semantic hierarchy that provides valuable structural priors for distinguishing unknown objects from in-distribution classes. In this work, we propose Hyp2Former, an end-to-end framework for OPS that does not require explicit modeling of unknowns during training, and instead learns hierarchical semantic similarities continuously in hyperbolic space. By explicitly encoding hierarchical relationships among known categories, the model learns a structured embedding space that captures multiple levels of semantic abstraction. As a result, unknown objects that cannot be confidently classified as known categories still remain in close proximity to higher-level concepts (e.g., an unknown animal remains closer to "animal" or "object" than to unrelated concepts such as "electronics" or "stuff") and can therefore be reliably detected, even if their fine-grained category was not represented during training. Empirical evaluations across multiple public datasets such as MS COCO, Cityscapes, and Lost&Found demonstrate that Hyp2Former outperforms existing methods on OPS, achieving the best balance between unknown object discovery and in-distribution robustness.
Yao Lu, Rohit Mohan, Florian Drews +2
Department of Computer Science, University of Freiburg, Germany · Bosch Research, Robert Bosch GmbH, Germany