Recent developments in feed-forward 3D reconstruction resulted in models which can recover dense scene representations and camera motion solely from an image stream. However, such predictions are prone to becoming inconsistent over long trajectories, specifically in demanding environments with repetitive structures, weak textures and dynamic objects or people. One way to mitigate those challenges is to use an omnidirectional camera, which provides wide spatial coverage and captures richer visual information. Yet, the majority of models do not offer support for 360-degree imagery or require additional fine-tuning. To bridge these two aspects, we present RIGOR: a large-scale reconstruction pipeline for gravity-aligned omnidirectional videos that retains a frozen feed-forward perspective backbone and exploits each panorama as a four-view virtual rig. The rig structure is used to detect and repair locally inconsistent predictions, to retrieve loop closures through cyclic four-view consensus, and to geometrically verify candidate revisits before global optimization. Verified constraints drive a Sim(3) pose graph that corrects accumulated rotation, translation, and scale drift along the sequence. We demonstrate that the proposed consistency mechanisms improve both trajectory accuracy and reconstructed geometry over a feed-forward baseline on challenging construction-site sequences. The code is made available under this link: https://github.com/TangentH/RIGOR.
Figures & tables
Figure 1 : Conceptual illustration of the RIGOR principle, providing consistency constraints for local and global reconstruction.
Method
Pano.
Dense
Long seq.
Loop
Pano train.
Feed-forward reconstruction
DUSt3R, MASt3R,
✗
✓
✗
✗
✗
VGGT, DA3, π3 , MapAnything
MUSt3R
✗
✓
✓
✗
✗
Persistent reconstruction / learned SLAM
Spann3R, CUT3R, SLAM3R,
✗
✓
✓
✗
✗
Table 1 : Comparison of representative reconstruction systems, categorized by the support of panoramic images ( Pano. ), dense reconstruction ( Dense ), long sequences ( Long seq. ), explicit revisit detection with corrective optimization ( Loop ), and whether panorama-specific learned training was required ( Pano train. ). “–” denotes that the criterion is not applicable.
Figure 2 : RIGOR retains a frozen perspective predictor and uses the known four-view panorama rig for validated local repair and cyclic multi-view loop retrieval. Joint feed-forward prediction verifies the retrieved candidates before accepted loop measurements enter a global Sim(3) pose graph for reconstruction fusion.
Figure 3 : Per-sequence distributions of trajectory and geometry metrics.
Method
ATE RMSE [m] ↓
RPE-10s RMSE [m] ↓
DA3-Sequential
2.841
1.827
DA3-Legacy
2.704
1.797
PanoVGGT
5.206
3.991
Ours w/o Rig Repair
2.841
1.827
Ours
2.312
1.712
Table 2 : Quantitative results. We report medians over 30 trajectories and 29 runs with LiDAR geometry. Every cloud is independently registered under the same scale-adjusting evaluation protocol. Best values within each block are bold.
Figure 4 : Qualitative comparison on two challenging sequences. For each sequence, the upper row shows the estimated trajectory and the lower row shows the reconstructed point cloud. Methods that did not consistently produce complete long-sequence reconstructions are included for qualitative analysis but omitted from the main quantitative table.
Retrieval
Pair AP ↑
R@P99 ↑
W/T/L
Single-view max
0.530
0.178
–
Cyclic four-view
0.609
0.190
24/0/4
Rig repair
ROI CD ↓
ROI F@25 cm↑
W/T/L
Off
0.230
0.718
–
On
0.132
0.873
28/0/1
Graph update
ATE RMSE ↓
RPE-10s RMSE ↓
W/T/L
Table 3 : Controlled component experiments. Retrieval uses 28 runs with GT-labeled revisits, geometry uses 29 LiDAR runs, and trajectory uses all 30 runs. We additionally report sequence-wise win/tie/loss (W/T/L) counts for the first metric in each ablation block. A win indicates that the full method outperforms the compared variant on that sequence according to the metric direction, a loss indicates the opposite, and values differing by less than 10−12 are counted as ties. The rig-repair block is an end-to-end ablation: loop verification and graph construction are recomputed after repair is disabled.
Figure 5 : Parameter choices. (a) The fraction of 400 cyclically adjacent-view pairs per FoV returning a non-empty homography consensus increases with wider crops. At 90∘ , the crops have no designed area overlap; non-empty responses can arise from boundary sampling or ambiguous matches. (b) Raw DA3 pose consistency on the same 100 panoramas peaks at 95∘ under the common 0.1 m center and 5∘ rotation diagnostic. These rates describe the parameter comparison, not accepted repairs. (c) On one 229-panorama run under the earlier sequential/no-loop protocol, yaw-four-aligned layouts avoid incomplete capture groups, and 32/16 gives the lowest ATE in this check.
Recent 3D geometric foundation models, such as VGGT, provide robust feed-forward 3D reconstruction by directly predicting camera poses and 3D scene points from input images. However, their results remain inaccurate, and scaling them to long sequences or large unordered image sets typically requires chunk-wise processing, which can introduce drift and inconsistency. We present Glob3R, a global SfM-style reconstruction built on 3D foundation models. Our key idea is to explicitly optimize feed-forward geometric predictions. To this end, we augment a frozen Pi3X backbone with a lightweight dense matching head that predicts image warps between selected reference frames and neighboring views. These dense warps are converted into sparse but reliable multi-view feature tracks, which provide correspondence constraints for global optimization. We further introduce a keyframe-based sliding-window association strategy that propagates tracks and relative poses across overlapping windows, enabling scalable reconstruction. Finally, we perform global motion averaging and bundle adjustment to refine camera poses, reduce scale inconsistencies, and recover dense scene geometry. Extensive experiments on indoor, outdoor, large-scale driving, and unordered SfM benchmarks demonstrate that Glob3R achieves robust and accurate reconstruction. It consistently improves over feed-forward foundation-model baselines and recent scalable reconstruction methods, while being more robust than classical SfM pipelines. The refined poses also lead to higher-quality neural rendering, validating the benefit of combining foundation-model priors with global geometric optimization. Project page: https://junyuandeng.github.io/Glob3r
Junyuan Deng, Heng Li, Kejie Qiu +7
The Hong Kong University of Science and Technology · Tongyi Lab, Alibaba Group · Nanjing University +1
Feed-forward 3D reconstruction provides an efficient paradigm for scene modeling from image sequences. Scaling these models to large monocular scenarios are constrained by excessive GPU memory footprint, degraded local geometry, and long-term trajectory drift. Existing chunk-based optimization strategies provide limited geometric constraints and fail to maintain global consistency over extended trajectories. We present a unified framework for stable and scalable feed-forward 3D reconstruction from long monocular sequences. Our approach builds on coarse-to-fine trajectory alignment augmented by lightweight geometric prior injection. Distilling monocular geometric cues into the feed-forward backbone via LoRA adaptation improves depth accuracy on fine structures while preserving inference efficiency. We introduce a hybrid-weight sparse ray-field optimization that leverages high-frequency geometric features to guide local point-cloud refinement and enforce consistent inter-frame ray constraints. Unlike prior chunk-based methods, this establishes strong cross-frame geometric coupling while maintaining scalability. Finally, an efficient trajectory stitching strategy with joint ray-error optimization explicitly reduces accumulated drift. Extensive experiments show that our approach achieves competitive trajectory accuracy compared with representative SLAM systems, while maintaining globally consistent 3D reconstruction in large-scale scenarios.
Enpeng Li, Yunzhou Zhang, Zhiyao Zhang +4
College of Information Science and Engineering, Northeastern University, China
This paper presents MAGiSt3R, a multi-agent 3D reconstruction framework performing reconstruction and camera tracking for monocular RGB videos at almost 10 FPS. MAGiSt3R relies on a feed-forward model from the 3R family to process RGB videos and regress local point maps, and on a merging model, MAGMA, that combines local maps at both intra-agent and inter-agent levels to obtain the final global point map. Furthermore, MAGiSt3R performs pose graph optimization to mitigate cumulative camera drift occurring along the feed-forward pipeline. We evaluate MAGiSt3R on both synthetic and real-world datasets, demonstrating its superior reconstruction and camera tracking accuracy compared to state-of-the-art approaches.
Ziren Gong, Xiaohan Li, Fabio Tosi +4
University of Bologna, Italy · Faculty of Dentistry, The University of Hong Kong, China · School of Instrument Science and Engineering, Southeast University, China +1