Mobile robots and vehicles carry synchronized multi-camera rigs, yet many streaming 3D foundation models are designed for monocular input, leaving efficient use of rig geometry a challenge. We present StreamRig, a freeze-and-stream framework that builds causal streaming odometry for calibrated rigs on a frozen multi-view 3D foundation model. The frozen front-end jointly perceives the synchronized views using rig calibration. A Rig-Resampler compresses their features, a CausalBridge applies causal attention with a key-value cache, and a lightweight head regresses rig poses. A periodic re-anchoring protocol supports stable pose estimation over long sequences. Only these modules are trained, 74.6M parameters in total, with relative poses as the sole supervision. Our two-stage training strategy combines group relocalization pretraining with causal rig training to transfer the geometric priors of the frozen front-end and the alignment ability of the pretrained modules to streaming odometry. We evaluate on NCLT, TartanGround, KITTI-360, and our self-collected humanoid-robot dataset ZJH, where training uses only simulation and real-world evaluation is zero-shot. Across all four datasets, StreamRig achieves lower translation and rotation drift than the evaluated non-oracle monocular streaming and rig-aware offline models, while maintaining low inference cost. Ablations and controlled camera-count experiments identify the sources of these gains. We further examine how longer training windows affect inference over longer horizons. Code has been released at https://github.com/WeiYuFei0217/StreamRig.
Figures & tables
Fig. 1: Accurate and efficient surround-view odometry. (a) StreamRig follows the NCLT ground-truth trajectory more closely than G2G-chain and CUT3R. (b) RGB-colored LiDAR points stitched using our predicted poses visualize the boxed region. (c) Five-camera StreamRig achieves lower ATE, GPU memory use, and latency than monocular CUT3R. Accuracy and costs are measured in separate configurations (Sec. IV-E ).
Fig. 2: StreamRig overview. (a) The frozen front-end perceives each rig with its calibration, the Rig-Resampler compresses every camera into L latent tokens, the CausalBridge attends causally over the compressed rigs, and the pose head regresses Ta←t from the bridge outputs. Bottom: inside the bridge, tokens are tagged with camera, group, and anchor embeddings, attention across rigs is causal with the keys and values of earlier rigs cached, and every query rig receives its own anchor snapshot; the optional long-history mode prunes the latents of distant rigs. (b) The two training stages, group relocalization with bidirectional attention and causal training on rig windows with the same modules, and the streaming loop that re-anchors every N rigs.
NCLT
TartanGround
KITTI-360
ZJH (real robot) ‡
5-cam surround, co-centered
4 × 90 ∘ co-located ring
stereo 0.59 m + 2 fisheye
stereo 7 cm + 2 oblique
2 traj. / 12.29 km
10 scenes / 12.60 km
2 scenes / 13.63 km
3 scenes / 0.36 km
Method
In.
Tr.
trel↓
rrel↓
ATE ↓
trel↓
rrel↓
ATE ↓
trel↓
rrel↓
ATE ↓
trel↓
rrel↓
ATE ↓
Non-streaming (offline / pairwise)
Reloc3R-rig † [ 15 ]
full
M
24.26
11.06
207.4
36.39
79.18
28.6
47.12
15.02
345.4
16.49
163.17
3.77
Rig3R † [ 11 ]
full
M
48.80
24.31
216.7
158.89
231.79
65.4
471.96
34.78
4257.0
28.02
409.79
5.23
TABLE I: STREAMING ODOMETRY ON FOUR RIGS ( trel IN %, rrel IN ∘ /100 M, ATE IN M).
Cameras
trel↓
rrel↓
SE3 ↓
Sim3 ↓
Mono
4.84
1.92
85.4
78.5
Stereo
3.59
1.33
67.2
65.6
Stereo + 2 side
2.59
0.98
63.7
59.1
TABLE II: MONO, STEREO, AND FULL RIG ON KITTI-360.
Fig. 3: Example trajectories from the Table I models, after full-sequence Sim(3) alignment and translation to a common start; KITTI-360 follows the same GT gap-bridging protocol as Table I .
Fig. 4: Re-anchoring interval and training horizon on NCLT. Translation drift (filled, left axis) and rotation drift (open, right axis) versus re-anchoring interval. Stars mark extrapolation beyond training limits. Top bars show estimated training costs in RTX 4090 GPU-hours.
Online 3D reconstruction requires estimating camera pose and scene geometry under strict causal and bounded-memory constraints. Existing methods often suffer from drift, jitter, or collapse on long sequences. We trace these failures to a fundamental mismatch. Streaming geometry is inherently temporally heterogeneous, with evidence ranging from short-lived correspondences to persistent global scale. However, current architectures impose uniform and pathological influence patterns. For example, sliding windows enforce hard cutoffs, while ungated recurrence and causal attention cause cache saturation and spike-like attention sinks. To resolve this, we formalize geometric propagation as an \emph{evidence influence kernel} and propose HorizonStream, a long-horizon Transformer that explicitly factorizes this kernel. For the long-range temporal factor, Geometric Linear Attention learns channel-wise decay rates to enable bounded, multi-timescale propagation of geometric evidence. For the short-range spatial factor, Geometric Local Attention with Spatiotemporal RoPE performs reliable 3D matching while suppressing attention sinks. Finally, Metric Readout Tokens recover stable scale and rigid pose directly from the persistent geometric state. Extensive experiments show that HorizonStream, trained on only 48-frame clips, generalizes stably to sequences exceeding 10,000\ frames with constant memory and linear time, achieving state-of-the-art streaming 3D reconstruction performance. Project Page: https://3dagentworld.github.io/horizonstream/
Recent feed-forward geometry foundation models have demonstrated impressive generalization by recovering depth and poses in a single forward pass. However, these models are typically constrained by a global coordinate frame assumption. This dependency becomes a significant bottleneck for long-context and streaming reconstruction, as it forces the network to maintain an arbitrary temporal origin and handle translation magnitudes that grow unbounded over time. Our solution, which we call R3, employs relative regression. We employ a lightweight MLP to predict confidence-weighted relative constraints. These confidences serve as a unified anchor: weighting losses during training and guiding pose aggregation during inference. R3 supports both full-context offline reconstruction and causal, bounded-memory streaming. Our evaluation in both offline and streaming settings validates the effectiveness of our relative mechanism. Project page: https://kevinxu02.github.io/r3-site
Congrong Xu, Huachen Gao, Xingyu Chen +3
University of Michigan · Westlake University · NVIDIA
Long-horizon online visual mapping requires continuous camera-motion and scene-geometry estimation under bounded computation. Recent feed-forward 3D reconstruction models provide strong geometric priors, but streaming variants often predict poses in a fixed or historically maintained coordinate system, leading to train--test mismatch, early-anchor attention bias, and accumulated drift. We propose \emph{Anchor3R}, a current-centric streaming 3D reconstruction framework that predicts window-relative poses and local geometry in the current-frame coordinate system. Overlapping predictions form a dense relative-pose graph, supporting online pose updates and loop-aware motion averaging for global reconstruction. Experiments on indoor, outdoor, driving, and RGB-D benchmarks demonstrate improved long-horizon pose accuracy and dense reconstruction quality over existing streaming baselines. Despite being trained only on 48-frame sequences, Anchor3R directly generalizes to streams exceeding 10,000 frames while maintaining bounded GPU memory during online inference. Code is available at https://github.com/polar-explorer/Anchor3R.