Streaming 3D reconstruction must preserve evidence from each frame while processing an expanding scene online. Spatial memory is a natural fit because it organizes history by reconstructed 3D location. Yet Point3R uses spatial proximity both to associate a new observation with an existing memory entry and to decide whether to fuse it, conflating co-location with state identity. Because pointers summarize image patches, nearby pointers may encode distinct surfaces, viewpoints, or visibility conditions; averaging them can destroy complementary evidence before later frames disambiguate it. We argue that location should determine address, not whether observations must merge. Slot3R realizes this principle as a training-free, set-associative retrofit that lets multiple states coexist at a shared address while keeping the pretrained Point3R backbone frozen. A bounded sparse readout further decouples persistent storage from per-frame decoder access. At 300-500 sampled frames, Slot3R reduces Point3R's point-cloud accuracy error (Acc) by 57.1%-63.1% on 7Scenes and 64.0%-72.0% on NeuralRGBD, lowers Sim(3)-aligned absolute trajectory error (ATE) on all three pose benchmarks, and remains competitive on video-depth estimation. It completes all evaluated settings from 600 to 1000 sampled frames at about 19 FPS under the same protocol, whereas Point3R and InfiniteVGGT run out of memory at 800 frames and beyond.
Figures & tables
Fig. 1 : Reconstruction quality, viewpoint-guided pose estimation, and streaming scalability. (a) Local point-cloud comparison with Point3R. (b) Indoor and outdoor reconstructions produced by Ours. (c) Relative to Point3R, both VPC variants reduce all three pose errors on Sintel and TUM-Dynamic; Ours-VPC-A lowers ATE by 41.8%/53.1% and translation RPE by 58.6%/79.6%, respectively. (d) The 640-token readout keeps decoder-visible memory bounded: Slot3R completes 1000-frame streams with Acc ≤0.062 , while Point3R fails with OOM at 800 frames.
Fig. 2 : Overview of the core Slot3R pipeline. (a) The frozen Point3R backbone interacts image tokens with selected memory states and predicts geometry and camera pose (Sec. 3.1 ). (b) K-way set-associative spatial memory retains multiple states at each discrete 3D address (Sec. 3.2 ). (c) Confidence and feature similarity govern the memory-write path through rejection, insertion, fusion, and replacement (Sec. 3.3 ). (d) Previous-frame spatial priors retrieve neighboring buckets; the resulting local states are combined with uniformly sampled global anchors under a fixed readout budget (Sec. 3.4 ).
Fig. 3 : Qualitative point-cloud reconstruction comparison on four 500-frame NeuralRGBD sequences. Our variants recover more complete and coherent scene geometry than Point3R and more closely approach the ground truth.
Method
Acc ↓
Comp ↓
NC ↑
Avg. FPS ↑
300
400
500
300
400
500
300
400
500
7Scenes
CUT3R
0.1362
0.1697
0.1926
0.0744
0.1078
0.0894
0.5513
0.5447
0.5409
13.89
TTT3R †
0.0402
0.0502
0.0651
0.0245
0.0263
0.0295
0.5647
0.5575
0.5522
13.68
StreamVGGT
OOM
OOM
OOM
OOM
OOM
OOM
OOM
OOM
OOM
OOM
STream3R
0.0962
0.1013
OOM
0.0317
0.0426
OOM
0.5769
0.5744
OOM
5.82
Table 1: Point-cloud reconstruction on 7Scenes and NeuralRGBD with k=2 at 300–500 frames. Pink , Orange , and Yellow denote the top three results per block; ties share and occupy the corresponding ranks. Δ reports the relative change from Point3R to Ours (Core), and OOM denotes failure. Avg. FPS averages completed runs. † marks reconstruction values reported by RetrieveVGGT [ 16 ] ; FPS is measured in our setup.
Method
Alignment
ScanNet
Bonn
KITTI
AbsRel ↓
δ<1.25↑
AbsRel ↓
δ<1.25↑
AbsRel ↓
δ<1.25↑
CUT3R
Per-sequence
0.056
97.1
0.074
94.4
0.111
88.3
TTT3R
Per-sequence
0.049
97.8
0.064
95.8
0.100
91.3
StreamVGGT
Per-sequence
0.268
57.1
0.059
97.2
0.173
72.2
STream3R
Per-sequence
0.037
98.9
0.058
97.5
0.077
95.2
InfiniteVGGT
Per-sequence
0.269
57.0
0.060
97.3
0.172
72.4
Table 2: Video depth estimation on ScanNet, Bonn, and KITTI under per-sequence alignment and direct metric-scale evaluation. δ<1.25 is reported as a percentage.
Method
ScanNet (Static)
Sintel
TUM-Dynamic
ATE RMSE ↓
RPE trans ↓
RPE rot ↓
ATE RMSE ↓
RPE trans ↓
RPE rot ↓
ATE RMSE ↓
RPE trans ↓
RPE rot ↓
CUT3R
0.0960
0.0180
0.4830
0.2020
0.0560
0.5290
0.0460
0.0120
0.3750
TTT3R
0.0640
0.0170
0.4590
0.2050
0.0740
0.5930
0.0290
0.0110
0.3270
StreamVGGT
0.1607
0.0515
3.2820
0.2589
0.1296
1.6673
0.0604
0.0303
2.7908
STream3R
0.1191
0.0407
2.3151
0.3721
0.0923
1.1694
0.0294
0.0125
0.3098
InfiniteVGGT
0.1618
0.0515
3.3262
0.2589
0.1296
1.6673
0.0607
0.0305
2.7962
Table 3: Camera pose estimation after Sim(3) alignment on 94 ScanNet scenes, 14 Sintel sequences, and 8 TUM-Dynamic sequences.
Component
Acc ↓
Comp ↓
NC ↑
FPS ↑
Final tokens ↓
Point3R
0.086
0.032
0.635
17.943
836.0
+ Set-associative K-way memory
0.044
0.017
0.676
16.346
1165.2
+ Sparse local readout (640)
0.042
0.017
0.670
19.411
1162.3
+ Filtering and anchors (Ours)
0.039
0.018
0.680
19.283
694.2
Table 4: Component ablation on nine 200-frame NeuralRGBD sequences ( k=2 ). Ours uses q=0.25 and a 640-token readout; Final tokens is the average sequence-end memory size.
Method
Acc ↓
Comp ↓
NC ↑
FPS ↑
Final tokens ↓
Sparse640 ( q=0.25 )
0.039
0.018
0.680
19.28
694.2
+ Predicted 3D position prior
0.039
0.018
0.676
12.29
695.4
Table 5: Spatial-prior ablation on nine 200-frame NeuralRGBD sequences ( k=2 ). The current-frame alternative adds a preliminary prediction pass; Final tokens follows Table 4 .
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Fig. 4 : Auxiliary viewpoint-guided pose conditioning. Spatial priors retrieve nearby viewpoint-bank entries, whose features are pooled using the pose query to condition its input to the shared decoder. Ours-VPC-M uses motion- and agreement-gated mixing, whereas Ours-VPC-A uses agreement-only mixing with per-frame bank refresh. The auxiliary bank leaves the structure and update rule of the core Slot3R memory unchanged, and neither variant requires a second decoder forward.
Method
Acc ↓
Comp ↓
NC ↑
Avg. FPS ↑
600
700
800
900
1000
600
700
800
900
1000
600
700
800
900
1000
7Scenes
CUT3R
0.2117
0.2200
0.2381
0.2658
0.2674
0.1263
0.1286
0.1279
0.1236
0.1010
0.5620
0.5602
0.5547
0.5502
0.5333
17.37
TTT3R
0.0901
0.1106
0.1338
0.1588
0.1638
0.0477
0.0616
0.0670
0.0728
0.0728
0.6217
0.6157
0.6075
0.6018
0.6013
17.04
StreamVGGT
OOM
OOM
OOM
OOM
OOM
OOM
OOM
OOM
OOM
OOM
OOM
OOM
OOM
OOM
OOM
OOM
InfiniteVGGT
0.0533
0.0545
OOM
OOM
OOM
0.0356
0.0378
OOM
OOM
OOM
0.6815
0.6772
OOM
OOM
OOM
3.44
Appendix
Table 6: Extended-sequence point-cloud reconstruction on 7Scenes and NeuralRGBD with k -frame sampling k=1 . “OOM” indicates out-of-memory failure. Pink , Orange , and Yellow cells mark the best, second-best, and third-best results within each dataset–length–metric block and within each dataset for Avg. FPS; ties occupy the corresponding ranks. † RetrieveVGGT exhausts CPU memory because its full-history KV repository grows with the input sequence.
Filtering policy
Acc ↓
Comp ↓
NC ↑
Avg. FPS ↑
200
300
400
500
1000
200
300
400
500
1000
200
300
400
500
1000
C–D, bucket 25%
0.0440
0.0450
0.0690
0.0686
OOM
0.0180
0.0150
0.0227
0.0204
OOM
0.6745
0.6620
0.6692
0.6673
OOM
19.52
C–D, bucket 12.5%
0.0390
0.0460
0.0672
0.0708
OOM
0.0180
0.0160
0.0227
0.0217
OOM
0.6810
0.6675
0.6733
0.6695
OOM
18.99
C–D, global 10%
0.0420
0.0620
0.1109
0.0861
OOM
0.0180
0.0190
0.0362
0.0250
OOM
0.6670
0.6490
0.6545
0.6554
OOM
19.26
C–D, global 15%
0.0530
0.1430
0.1753
0.1242
OOM
0.0200
0.0290
0.0660
0.0362
OOM
0.6625
0.6435
0.6446
0.6469
OOM
19.33
No dropping
0.0390
0.0450
0.0617
0.0650
OOM
0.0180
0.0160
0.0214
0.0203
OOM
0.6795
0.6635
0.6737
0.6736
OOM
18.91
Appendix
Table 7: Ablation of the ordering and scope of confidence-guided memory filtering on NeuralRGBD with k=2 and n=9 sequences. All variants retain the same K-way memory, confidence-aware fusion, and 640-token sparse readout. C–D and D–C denote Compare-then-Drop and Drop-then-Compare, respectively. Avg. FPS is computed over the matched 200–500-frame runs. This diagnostic sweep was run separately from the main comparison in Table 1 . OOM denotes failure on at least one sequence.
q
Acc ↓
Comp ↓
NC ↑
FPS ↑
Final tokens ↓
0.00
0.0649
0.0227
0.6713
17.949
1367.3
0.05
0.0637
0.0228
0.6678
19.140
1291.4
0.10
0.0621
0.0211
0.6713
19.230
1210.9
0.15
0.0624
0.0220
0.6698
19.141
1163.4
0.20
0.0620
0.0220
0.6702
19.253
1101.4
0.25
0.0589
0.0236
0.6729
19.216
1046.2
Appendix
Table 8: A separate sensitivity sweep of the pre-write drop quantile q on 500-frame NeuralRGBD reconstruction with n=9 sequences. All variants within this sweep use the same K-way memory and 640-token sparse readout.
Alignment
q
ScanNet
Bonn
KITTI
AbsRel ↓
δ<1.25↑
AbsRel ↓
δ<1.25↑
AbsRel ↓
δ<1.25↑
Metric-scale
0.00
0.06219
96.54
0.08453
96.08
0.18206
77.19
Metric-scale
0.05
0.06225
96.53
0.08534
95.98
0.18179
77.07
Metric-scale
0.10
0.06179
96.58
0.08514
96.04
0.18216
77.33
Metric-scale
0.15
0.06251
96.45
0.08521
95.97
0.18135
77.11
Metric-scale
0.20
0.06230
96.43
0.08547
96.01
0.17915
77.72
Appendix
Table 9: Depth sensitivity to the pre-write drop quantile q on ScanNet, Bonn, and KITTI under metric-scale and per-sequence alignment. These results are from a diagnostic sweep separate from the main depth evaluation in Table 2 . δ<1.25 is reported as a percentage.
Method
K
Acc ↓
Comp ↓
NC ↑
FPS ↑
Avg. final tokens ↓
Point3R
–
0.047003
0.018037
0.578354
17.746
918.0
Ours (Core)
1
0.089329
0.053222
0.569152
19.911
219.2
Ours (Core)
2
0.029775
0.017823
0.580439
19.935
250.9
Ours (Core)
4
0.028163
0.016960
0.583035
19.816
440.4
Ours (Core)
8
0.026758
0.017059
0.584917
19.405
754.3
Ours (Core)
16
0.026092
0.016093
0.585791
18.958
1235.2
Appendix
Table 10: Bucket-capacity ablation on 7Scenes with k=2 and 200 sampled frames, where K is the maximum number of slots per spatial bucket. Avg. final token count is averaged across evaluated sequences.
Method
K
Acc ↓
Comp ↓
NC ↑
FPS ↑
Avg. final tokens ↓
Point3R
–
0.034096
0.024413
0.649663
13.618
1406.9
Ours (Core)
1
0.056078
0.039312
0.636682
16.914
168.3
Ours (Core)
2
0.042358
0.029619
0.641545
16.199
271.4
Ours (Core)
4
0.034110
0.023624
0.646727
17.047
482.4
Ours (Core)
8
0.031013
0.021317
0.648767
16.459
847.0
Ours (Core)
16
0.031309
0.021690
0.648118
16.513
1415.4
Appendix
Table 11: Bucket-capacity ablation on 7Scenes when sampling every 20th frame ( k=20 ). Avg. final token count is averaged across evaluated sequences.
Sparse budget
Acc ↓
Comp ↓
NC ↑
FPS ↑
128
0.043
0.017
0.653
20.000
256
0.041
0.017
0.665
20.068
384
0.040
0.017
0.677
20.312
512
0.039
0.018
0.680
19.756
640
0.038
0.018
0.677
19.114
768
0.038
0.017
0.678
18.925
Appendix
Table 12: Sparse-readout budget sensitivity on NeuralRGBD point-cloud reconstruction with k=2 and 200 sampled frames. This diagnostic sweep was run separately from the component ablation in Table 4 .
Sparse budget
AbsRel ↓
δ<1.25↑
128
0.06508
0.96185
256
0.06427
0.96323
384
0.06453
0.96245
512
0.06405
0.96303
640
0.06326
0.96399
768
0.06323
0.96429
Appendix
Table 13: Sparse-readout budget sensitivity for depth estimation.
Hash
n
Acc ↓
Comp ↓
NC ↑
Acc med. ↓
Spherical
7
0.0491
0.0280
0.7451
0.0318
Voxel
7
0.1409
0.0564
0.6076
0.0991
Appendix
Table 14: Hash parameterization ablation for sparse-512 readout.
Method
Acc ↓
Comp ↓
NC ↑
FPS ↑
Avg. stored tokens ↓
Point3R w/o fused RoPE3D
0.084
0.030
0.753
4.547
1571.0
Point3R w/ fused RoPE3D
0.083
0.030
0.752
14.750
1572.6
Ours (Core) w/o sparse readout
0.061
0.027
0.753
16.658
1204.1
Ours (Core)
0.063
0.029
0.753
18.269
1220.1
Appendix
Table 15: Effect of fused RoPE3D and sparse readout on NeuralRGBD diagnostic runs with approximately 50-frame sequences. “Ours (Core) w/o sparse readout” uses K-way memory and pre-write filtering with q=0.25 , but exposes the full stored memory to the decoder; Ours (Core) instead applies a bounded readout of at most 640 tokens, comprising up to 128 uniformly sampled global anchors and up to 512 additional local tokens after deduplication.
Dense 3D reconstruction from continuous image streams requires both accurate geometric aggregation and stable long-term memory management. Recent feed-forward reconstruction frameworks integrate observations through persistent memory representations, yet most rely primarily on appearance-based similarity when updating memory. Such appearance-driven integration often leads to redundant accumulation of observations and unstable geometry when viewpoint changes occur. In this work, we propose a ray-aware pointer memory for streaming 3D reconstruction that explicitly models both spatial location and viewing direction within a unified memory representation. Each memory pointer stores its 3D position, associated ray direction, and feature embedding, allowing the system to reason jointly about geometric proximity and viewpoint consistency. Based on this representation, we introduce an adaptive pointer update strategy that replaces traditional fusion-based memory compression with a retain-or-replace mechanism. Instead of averaging nearby observations, the system selectively retains informative pointers while discarding redundant ones, preserving distinctive geometric structures while maintaining bounded memory growth. Furthermore, the joint reasoning over spatial distance and ray-direction discrepancy enables the system to distinguish between local redundancy, novel observations, and potential loop revisits in a unified manner. When loop candidates are detected, pose refinement is triggered to enforce global geometric consistency across the reconstruction. Extensive experiments demonstrate that the proposed ray-aware memory design significantly improves long-term reconstruction stability and camera pose accuracy while maintaining efficient streaming inference. Our approach provides a principled framework for scalable and drift-resistant online 3D reconstruction from image streams.
Feifei Li, Qi Song, Chi Zhang +1
The Chinese University of HongKong, Shenzhen · Tsinghua University
Streaming 3D reconstruction relies on a compact recurrent scene state to process long image streams in linear time and bounded memory. However, repeated updates can gradually corrupt this state, causing reliable historical information to be overwritten by noisy or ambiguous observations. We introduce ReCal3R, a reliability-calibrated learning rate method for recurrent 3D reconstruction. Instead of directly applying a candidate learning rate, our method estimates state token reliability from the maintained scene state and uses it to calibrate a candidate learning rate derived from token alignment, state reconstruction residual, and recent update pressure. The resulting token-wise learning rate interpolates between a conservative base rate and the candidate rate, suppressing aggressive updates on unreliable tokens while preserving adaptation to informative frames. Applied to CUT3R as a training-free calibration rule, ReCal3R reaches strong performance on long sequences in pose, depth, and reconstruction quality, including a 3.7× reduction in ATE, with comparable runtime and memory. Code is available at: https://github.com/Powertony102/ReCal3R.
Xinze Li, Yiyuan Wang, Pengxu Chen +4
Beijing Normal-Hong Kong Baptist University · Hong Kong Baptist University · Jilin University +2
Streaming recurrent models enable efficient 3D reconstruction by maintaining persistent state representations. However, they suffer from catastrophic forgetting over long sequences due to balancing historical information with new observations. Recent methods alleviate this by deriving adaptive signals from the attention perspective, but they operate on single dimensions without considering temporal and spatial consistency. To this end, we propose a training-free framework termed TTSA3R that leverages both temporal state evolution and spatial observation quality for adaptive state updates in 3D reconstruction. In particular, we devise a Temporal Adaptive Update Module that regulates update magnitude by analyzing temporal state evolution patterns. Then, a Spatial Contextual Update Module is introduced to localize spatial regions that require updates through observation-state alignment and scene dynamics. These complementary signals are finally fused to determine the state updating strategies. Extensive experiments show that TTSA3R achieves competitive performance on standard short-sequence benchmarks and provides substantially stronger robustness on extended sequences. On NRGBD, as sequences extend from 50 to 250 frames, TTSA3R exhibits only a 1.33x error increase, compared with over 4x degradation for CUT3R. This highlights the practical value of temporal-spatial adaptive updates for long-term reconstruction stability. Our code is available at https://github.com/anonus2357/ttsa3r.