Robots need to be able to understand their surroundings in order to operate safely and robustly, and to interact with the surrounding environment. Robots deployed in unconstrained real-world scenarios must additionally be able to deal with novel situations and objects that have never been seen before. In this article, we tackle the problem of open-world panoptic segmentation, i.e., the task of discovering new semantic categories and new object instances at test time, while enforcing consistency among the categories that we incrementally discover. We present Con2MAV, a general method for open-world panoptic segmentation. Experiments across a wide range of datasets, from road scenes to underwater environments, highlight its compelling capabilities in open-world segmentation and its competitive performance on known classes. We will open-source the implementation of our approach upon acceptance. In addition, we propose PANIC (Panoptic ANomalies In Context), a benchmark for evaluating open-world segmentation tasks in autonomous driving scenarios. This dataset, recorded with a multi-modal sensor suite mounted on a car, and then manually annotated, provides high-quality, pixel-wise annotations of anomalous objects at both semantic and instance level. PANIC contains 800 images, more than 50 unknown classes, i.e., classes that do not appear in the training set, and over 4,000 object instances, providing a comprehensive benchmark for evaluating open-world segmentation methods in autonomous driving scenarios. We provide competitions for multiple open-world segmentation tasks on a hidden test set. Our dataset and competitions are available at https://www.ipb.uni-bonn.de/data/panic.
Figures & tables
Figure 1 : Our proposed approach, Con2MAV, is able to tackle multiple open-world tasks and segment unknown objects and categories. In the figure, we show predictions on (left to right) SegmentMeIfYouCan ( Chan et al., 2021 ) , SUIM ( Islam et al., 2020 ) , COCO ( Lin et al., 2014 ) , and PANIC (ours).
Figure 2 : Our dataset, PANIC, provides pixel-wise annotations of unknown semantic categories and object instances of RGB images. The images have been recorded with our own inhouse sensor suite ( Vizzo et al., 2023 ) mounted on our vehicle driving in Bonn, Germany. The dataset consists of images collected at different times of day over the span of more than a year.
Figure 3 : A visual breakdown of the four open-world segmentation tasks. Anomaly segmentation segment all anomalous areas as unknown (zebras and lion together). Open-world semantic segmentation separates classes but has no objects (zebras segmented together, lion separate). Open-set panoptic segmentation segments separate objects but has no category information. Open-world panoptic segmentation has both, classes and object information. Exemplary RGB image is generated with perplexity.ai ( Perplexity Deep Research, 2025 ) .
Figure 4 : Our network processes an RGB image via an encoder and three decoders, for semantic segmentation, anomaly segmentation and class-agnostic instance segmentation. The semantic segmentation decoder also builds class descriptors for the known categories. Results are post-processed and yield the final open-world panoptic segmentation result.
Dataset
Images
Semantic Classes
Instances
Hidden Test Set
Val
Test
Fishyscapes Lost-and-Found ( Blum et al., 2019 )
373
1203
N.A.
1864
✓
CAOS BDDAnomaly ( Hendrycks et al., 2022 )
0
810
3
1231
✗
RoadObstacle21 ( Chan et al., 2021 )
0
327
N.A.
388
✓
SegmentMeIfYouCan ( Chan et al., 2021 )
10
100
N.A.
262
✓
PANIC (ours)
131
679
58
4029
✓
Table 1 : Comparison of open-world segmentation datasets. In the semantic classes, “N.A.” means that there is no label.
Figure 5 : Sensor setup we used for recording data for the PANIC dataset. The setup includes four cameras, one GNSS/IMU device, and two 3D LiDARs. For further details, please refer to Vizzo et al. (2023) .
Approach
Pixel-Level
Component-Level
AUPR
FPR95
sIoU
PPV
mF1
Maskomaly
93.4
6.9
55.4
51.2
49.9
RbA
86.1
15.9
56.3
41.4
42.0
ContMAV
90.2
3.8
54.5
61.9
63.6
UNO
96.1
2.3
68.0
51.9
58.9
Con2MAV
90.0
2.7
59.1
68.3
69.4
Table 3 : Anomaly segmentation results on the test set of SegmentMeIfYouCan. Best results are highlighted in bold. More results available on the public leaderboard.
Approach
Pixel-Level
Component-Level
AUPR
FPR95
sIoU
PPV
mF1
ContMAV
91.7
66.4
15.0
72.1
24.2
Con2MAV
95.7
35.3
20.9
64.7
31.2
Table 4 : Anomaly segmentation results on the hidden test set of our dataset, PANIC. Best results are highlighted in bold. Public competition at codabench.org/competitions/4561 .
Approach
IoU u
mIoU u
Train
Motorcycle
Bicycle
Background + cluster
0
32.3
32.8
21.7
ContMAV
62.4
62.2
56.8
60.5
Con2MAV
66.5
64.4
53.8
61.6
Closed-world
72.3
69.3
60.9
67.5
Table 7: Open-world semantic segmentation results on BDDAnomaly. Best results are highlighted in bold.
Approach
IoU u
mIoU u
Human
Wrecks & Ruins
Robot
ContMAV
46.2
36.9
46.2
43.1
Con2MAV
69.4
64.3
53.0
62.2
Closed-world
82.0
71.6
80.2
77.9
Table 8: Open-world semantic segmentation results on the SUIM dataset. Best results are highlighted in bold.
Figure 6 : Qualitative results of our approach, Con2MAV, on open-world semantic segmentation on SUIM (top row) and PANIC (bottom row). The prediction mask is overlayed to the input RGB for clarity. In the prediction, different colors correspond to different predicted classes. We compare our approach, Con2MAV (right), with our old method, ContMAV (center).
Figure 7 : Qualitative results of our approach, Con2MAV, on open-set panoptic segmentation on COCO (top row), and PANIC (bottom row). The prediction mask is overlayed to the input RGB for clarity. In the semantic prediction, the colored area indicates the anomalous region. In the instance prediction, different colors correspond to different instance ids.
Figure 8 : Qualitative results of our approach, Con2MAV, on open-world panoptic segmentation on PANIC. The prediction mask is overlayed to the input RGB for clarity. In the semantic prediction, different colors correspond to different predicted classes. In the instance prediction, different colors correspond to different instance ids. In the top row, we show the segmentation of the unknown parts only. In the bottom row, the complete open-world panoptic segmentation results.
Figure 9 : GPT-4V and PaliGemma results on an image from SegmentMeIfYouCan ( Chan et al., 2021 ) .