All Roads Lead to Rome: Flow-driven Multi-Anchor Exploration for Open-Environment Active 3D Mapping
Authors: Yang Li, Aming Wu, Zihao Zhang, Ziju Han, Sijia Zhang, Yahong Han
Organizations: School of Artificial Intelligence, Tianjin University, China · School of Computer Science and Information Engineering, Hefei University of Technology, China
To advance the development of embodied intelligence, Open-Environment Active 3D Mapping has attracted increasing attention, aiming to perform a long-horizon and shortest trajectory exploration for reconstructing unseen scenarios. Since only limited information about unseen environments is available, methods built on the closed-set assumption, i.e., assuming that the test environments are similar to those seen during training, cannot generalize satisfactorily. In existing active mapping methods, long-horizon exploration is often guided by predicting a coarse long-range goal and then converting it into an executable path. However, this stage is usually formulated as single-point prediction. Under partial observability, the same local observation may correspond to multiple plausible exploration directions, making such deterministic prediction prone to brittle decisions and degraded performance in unseen scenarios. Our experiments further verify that this is a key factor underlying their weak generalization. To address this issue, we reformulate long-horizon target prediction as conditional multimodal anchor generation using Conditional Flow Matching.Instead of predicting a single goal, our method learns a conditional distribution over coarse exploration anchors from the current mapping state. These anchors are first converted into executable candidate paths through obstacle-aware planning. We then apply exploration-mode clustering to compress geometrically similar trajectories and reduce candidate redundancy. Finally, a hierarchical selection module selects the most promising mode and reranks paths within it to produce the final executable trajectory. Experiments show that our method improves generalization and reconstruction efficiency in open environments.
Figures & tables
Figure 1 : Motivation. Under partial observability, a similar local observation may support multiple plausible long-horizon exploration directions. Single-goal prediction collapses these futures into one dominant branch and generalizes poorly when the unseen scene continuation changes. In contrast, multimodal anchor modeling preserves alternative hypotheses before downstream planning, leading to better generalization and higher coverage.
Figure 2 : Conditional Flow Matching for multimodal long-horizon anchors. CFM transports samples from a simple prior x0∼N(0,I) toward the target anchor distribution x1∼p1(⋅∣zt) , yielding multiple plausible long-horizon exploration hypotheses under the same partial observation. Unlike directly generating full trajectories, this formulation models coarse exploration intent in a compact space and remains naturally compatible with downstream planning.
Figure 3 : Overview of our method. Given the current mapping state Et , a mapping-progress encoder produces the context feature zt . Conditional Flow Matching (CFM) then generates a multimodal set of long-horizon exploration anchors At , which are converted into executable candidate paths and clustered into exploration modes {Mt(m)}m=1Mt . Finally, a hierarchical selection module first selects the most promising mode and then reranks paths within that mode to output the final trajectory τt⋆ .
Simple
Normal
Hard
Insane
Final Cov.
AUCs
Final Cov.
AUCs
Final Cov.
AUCs
Final Cov.
AUCs
Random
0.323 ±0.156
0.270 ±0.135
0.190 ±0.124
0.152 ±0.103
0.124 ±0.082
0.088 ±0.060
0.074 ±0.048
0.050 ±0.035
FBE [ 18 ]
0.760 ±0.174
0.605 ±0.171
0.565 ±0.139
0.415 ±0.109
0.425 ±0.114
0.311 ±0.080
0.330 ±0.097
0.239 ±0.079
SCONE [ 19 ]
0.577 ±0.173
0.483 ±0.138
0.412 ±0.114
0.313 ±0.087
0.290 ±0.093
0.210 ±0.072
0.196 ±0.079
0.140 ±0.060
MACARONS [ 7 ]
0.599 ±0.200
0.479 ±0.172
0.418 ±0.120
0.314 ±0.088
0.302 ±0.097
0.218 ±0.070
0.192 ±0.078
0.139 ±0.058
NBP [ 8 ]
0.879 ±0.142
0.692 ±0.156
0.734 ±0.142
0.526 ±0.112
0.618 ±0.153
0.432 ±0.115
0.472 ±0.095
0.312 ±0.073
Table 1 : Evaluation results on AiMDoom Dataset. For each difficulty level, all baseline models, including ours, are trained from scratch on the corresponding training set to ensure a fair comparison.
Method
Comp. (%) ↑
Comp. (cm) ↓
Random
45.67
26.53
FBE [ 18 ]
71.18
9.78
UPEN [ 3 ]
69.06
10.60
OccAnt [ 6 ]
71.72
9.40
ANM [ 4 ]
73.15
9.11
NBP [ 8 ]
79.38
6.78
Table 2 : Comparison on the MP3D dataset.
Method
Comp. (%) ↑
Comp. (cm) ↓
Random
32.12
29.36
FBE [ 18 ]
56.73
14.45
UPEN [ 3 ]
54.70
13.09
OccAnt [ 6 ]
58.49
14.18
ANM [ 4 ]
60.22
11.78
NBP [ 8 ]
61.81
10.22
Table 3 : Train on AiMDoom test on MP3D.
Figure 4 : Qualitative comparison between the SOTA method NBP (Top) and Ours (Bottom). For each scene, both methods start from the same initial pose. The dashed boxes highlight representative regions that remain unexplored by NBP but are successfully recovered by our method. This difference is not merely local: NBP tends to commit early to a single dominant long-horizon branch, which induces a narrower global coverage pattern and leaves weakly indicated but still informative continuations unexplored. In contrast, our method preserves multiple plausible long-horizon exploration hypotheses before downstream planning and selection, enabling more balanced expansion across alternative scene continuations and ultimately yielding more complete scene coverage.
Simple
Normal
Hard
Insane
Protocol
Method
Final Cov.
AUC
Final Cov.
AUC
Final Cov.
AUC
Final Cov.
AUC
Protocol A: Train on Simple, test on all difficulty levels
Random
0.323 ± 0.156
0.270 ± 0.135
0.174 ± 0.112
0.132 ± 0.091
0.101 ± 0.075
0.072 ± 0.054
0.062 ± 0.041
0.041 ± 0.030
FBE [ 18 ]
0.760 ± 0.174
0.605 ± 0.171
0.512 ± 0.155
0.364 ± 0.124
0.368 ± 0.132
0.255 ± 0.098
0.284 ± 0.110
0.188 ± 0.088
SCONE [ 19 ]
0.577 ± 0.173
0.483 ± 0.138
0.354 ± 0.128
0.245 ± 0.104
0.213 ± 0.111
0.142 ± 0.082
0.145 ± 0.092
0.094 ± 0.071
MACARONS [ 7 ]
0.599 ± 0.200
0.479 ± 0.172
0.368 ± 0.141
0.252 ± 0.115
0.225 ± 0.124
0.148 ± 0.087
0.151 ± 0.096
0.098 ± 0.075
Table 4 : Generalization under difficulty shift on AiMDoom. We report two complementary protocols. Protocol A evaluates cross-difficulty transfer by training on the Simple split only and directly testing on all difficulty levels . Protocol B evaluates leave-one-difficulty-out generalization by training on Simple + Normal + Hard and directly testing on the held-out Insane split.
Method
OOD Setting
Baseline
MAG
EMC
HMS
Comp. (%) ↑
Comp. (cm) ↓
✓
61.81
10.22
✓
✓
66.21
9.64
✓
✓
✓
67.60
9.36
✓
✓
✓
✓
68.45
9.11
Table 5 : Progressive validation of the proposed pipeline under the cross-dataset OOD setting.
Current 3D mapping pipelines generally assume static environments, which limits their ability to accurately capture and reconstruct moving objects. To address this limitation, we introduce the novel task of active mapping of moving objects, in which a mapping agent must plan its trajectory while compensating for the object's motion. Our approach, Paparazzo, provides a learning-free solution that robustly predicts the target's trajectory and identifies the most informative viewpoints from which to observe it, to plan its own path. We also contribute a comprehensive benchmark designed for this new task. Through extensive experiments, we show that Paparazzo significantly improves 3D reconstruction completeness and accuracy compared to several strong baselines, marking an important step toward dynamic scene understanding. Project page: https://davidea97.github.io/paparazzo-page/
Davide Allegro, Shiyao Li, Stefano Ghidoni +1
University of Padova · 2LIGM, ´Ecole Nationale des Ponts et Chauss´ees, IP Paris, Univ Gustave Eiffel, CNRS
Long-horizon online visual mapping requires continuous camera-motion and scene-geometry estimation under bounded computation. Recent feed-forward 3D reconstruction models provide strong geometric priors, but streaming variants often predict poses in a fixed or historically maintained coordinate system, leading to train--test mismatch, early-anchor attention bias, and accumulated drift. We propose \emph{Anchor3R}, a current-centric streaming 3D reconstruction framework that predicts window-relative poses and local geometry in the current-frame coordinate system. Overlapping predictions form a dense relative-pose graph, supporting online pose updates and loop-aware motion averaging for global reconstruction. Experiments on indoor, outdoor, driving, and RGB-D benchmarks demonstrate improved long-horizon pose accuracy and dense reconstruction quality over existing streaming baselines. Despite being trained only on 48-frame sequences, Anchor3R directly generalizes to streams exceeding 10,000 frames while maintaining bounded GPU memory during online inference. Code is available at https://github.com/polar-explorer/Anchor3R.
We present Semantic-Aware Guided Exploration, SAGE, a system for open-vocabulary exploration in unknown 3D indoor environments that preserves coverage-oriented behavior while allowing semantic cues to reprioritize frontier selection. Building on the FALCON volumetric explorer, SAGE integrates Contrastive Language-Image Pre-training (CLIP) via four key components: object-centric embedding storage, a temporal cache that projects recent observations onto the free-unknown boundary, object frontiers for high-similarity detections, and a unified semantic-geometric planning cost. This cost function bounds semantic reweighting influence, ensuring frontiers are prioritized without sacrificing total coverage. In Matterport3D-based simulations, SAGE outperforms FALCON and a semantic-only ablation in object discovery across map-query pairs. Compared to Finding Things in the Unknown (FTU), SAGE completes exploration 9.0 to 25.9 times faster across the nine shared map-query pairs, achieving a mean speedup of 13.7. Furthermore, SAGE achieves substantially higher volumetric throughput than FTU. Finally, we deploy SAGE in five real-world flights in two environments on a Modal AI Starling 2 quadrotor with onboard sensing and planning, and offboard CLIP inference. Comparing SAGE and FALCON, we find that while FALCON results in faster exploration and shorter mapping trajectories, SAGE outperforms FALCON in terms of object discovery.
Nitin Vegesna, Avideh Zakhor
Department of Electrical Engineering and Computer Sciences University of California, Berkeley Berkeley, CA, USA