Streaming 3D reconstruction requires more than a sequence of geometric predictions: it requires a persistent scene state that can incorporate new evidence and remain renderable as observations arrive. Latent spatial tokens offer a promising representation for this purpose, but constructing them from an image collection leaves open how to maintain them online, where each observation may both revisit known regions and reveal new content. We introduce S2Tok, a feed-forward framework that maintains a size-adaptive, persistent scene state from uncalibrated image streams. Its central idea is to distinguish updates to the existing representation from selective expansion. A spatially informed transformer integrates each incoming observation with the persistent scene tokens, while a learned admission module selectively expands the representation to limit redundant storage. A hierarchical decoder and Gaussian head convert the evolving state into non-pixel-aligned 3D Gaussians, enabling novel-view rendering without caching previous frames. Experiments across four benchmarks demonstrate competitive streaming rendering quality with compact Gaussian representations. These results support latent spatial tokens as a persistent computational state for online 3D reconstruction, combining learned scene updates with explicit Gaussian rendering.
Figures & tables
Figure 1. S2Tok on Various Scenarios. We show the final 3DGS scenes reconstructed by S2Tok in a streaming manner across diverse scenarios. More streaming reconstruction samples are in Fig. 12 .
Figure 2. S2Tok Pipeline. The encoder maps the new frame Ii to tokens fi . The SS-Transformer jointly updates frame tokens and existing spatial tokens Si−1 . The admission module appends selected frame tokens to form Si , which is carried to the next step. The Hierarchical Decoder and Gaussian Head decode and predict Gaussian attributes. ⊕ denotes concatenation.
Figure 3. Stream-Spatial Transformer. Scale token s , spatial tokens Si−1 , and frame tokens fi attend to one another in Aglobal . Scene-local attention Ascene then updates s and Si−1 , while frame-local attention Aframe updates fi .
Figure 4. Hierarchical Decoder. Hierarchical tokens SIII , SII , and SI from the SS-Transformer are processed individually through self-attention layers and then fused through residual addition + into the decoding stream, which is initialized from SIII . A self-attention layer allows one more update after each residual addition.
Method
Stream- ing
Cam- free
Final
Intermediate
Time (ms/v)
Mem. (GiB)
12 views
50 views
Avg. #GS/PT
12 views
50 views
12 / 50 views
ZipSplat r=1
\ding 55
\ding 51
23.04/0.69/0.24
21.94/0.63/0.28
124K/518K
23.26/0.70/0.25
22.13/0.62/0.29
–
–
ZipSplat †
\ding 55
\ding 51
21.79/0.64/0.31
20.88/0.58/0.35
36K/66K
21.74/0.64/0.32
20.81/0.57/0.37
–
–
TokenGS
\ding 55
\ding 55
21.74/0.69/0.40
20.53/0.63/0.50
262K/262K
22.64/0.73/0.34
21.54/0.65/0.45
–
–
GlobalSplat
\ding 55
\ding 55
17.71/0.47/0.55
15.92/0.40/0.63
32K/32K
18.65/0.52/0.49
17.01/0.44/0.58
–
–
Table 1. Quantitative Comparison on DL3DV. Each metric cell reports PSNR ↑ / SSIM ↑ / LPIPS ↓ . r=1 denotes the default compression ratio of ZipSplat. † adjusts the ratio to match our token budget. Avg. represents the average across all sequences. Time and Mem. report per-view reconstruction time and peak GPU memory of the streaming methods in the 50-view setting, measured on a single NVIDIA A100 80GB. Notably, MonoGS is not feed-forward.
Figure 5. Qualitative Rendering Results on DL3DV (Final). “GT depth” denotes the pseudo-GT depth produced by DA3 ( Lin et al., 2025 ) .
Method
Stream- ing
Cam- free
RealEstate10K (12 Views)
Mip-NeRF 360 (50 Views)
Final
Avg. #GS/PT
Intermediate
Final
Avg. #GS/PT
Intermediate
ZipSplat r=1
\ding 55
\ding 51
27.65/0.87/0.13
124K
27.95/0.88/0.12
21.46/0.54/0.38
518K
21.39/0.53/0.38
ZipSplat †
\ding 55
\ding 51
25.47/0.82/0.19
17K
26.26/0.84/0.17
20.66/0.50/0.43
82K
20.68/0.49/0.43
TokenGS
\ding 55
\ding 55
32.00/0.94/0.15
262K
30.79/0.90/0.17
22.29/0.64/0.48
262K
22.88/0.67/0.43
GlobalSplat
\ding 55
\ding 55
29.85/0.91/0.13
32K
29.23/0.89/0.14
17.25/0.39/0.65
32K
17.85/0.42/0.59
AnySplat
\ding 55
\ding 51
23.39/0.80/0.14
663K
25.87/0.85/0.12
18.85/0.45/0.36
2850K
19.18/0.51/0.32
Table 2. Quantitative Comparison on RealEstate10K and Mip-NeRF 360. Each metric cell reports PSNR ↑ / SSIM ↑ / LPIPS ↓ . Results are evaluated under 12-view and 50-view settings. Detailed comparisons are in Tab. 5 and Tab. 6 in the Appendix. r=1 is the default compression ratio of ZipSplat. † adjusts the compression ratio to match our token budget. Notably, MonoGS is not feed-forward.
Figure 6. Qualitative Rendering Results on ScanNet++ (Final). “GT depth” denotes the pseudo-GT depth produced by DA3 ( Lin et al., 2025 ) .
Figure 7. Qualitative Rendering Results on Mip-NeRF 360 (Intermediate). “GT depth” denotes the pseudo-GT depth produced by DA3 ( Lin et al., 2025 ) .
Variant
Final
Intermediate
Avg. #GS
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
w/o RoPE3D
17.74
0.4117
0.5166
18.85
0.4452
0.4794
65K
w/o Admission
17.50
0.4076
0.5476
20.39
0.5087
0.4149
518K
w/o SCST
18.13
0.4281
0.4997
19.67
0.4757
0.4463
120K
w/o SCST(matched budget)
17.90
0.4216
0.5112
19.03
0.4541
0.4730
66K
w/o scale token s
17.89
0.4214
0.5126
19.01
0.4543
0.4765
66K
Table 3. Ablation experiments on the 50-view DL3DV evaluation set.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Streaming Setting
Method
Streaming
Latent Spatial Token/Memory
Cache/Window- free
Pixel- free Prediction
Reconstruction Kernel
Pose- free
Training Dataset(s)
AnySplat
\ding 55
\ding 55
non- streaming
\ding 55
GS
\ding 51
Hypersim, ARKitScenes, BlendedMVS, ScanNet++, CO3D-v2, Objaverse, Unreal4K, WildRGB-D, and DL3DV
Spann3R
\ding 51
\ding 51
\ding 55
\ding 55
Point
\ding 51
Habitat, ScanNet, ScanNet++, ARKitScenes, BlendedMVS, and CO3D-v2 †
Table 4. Detailed Comparison of Related Feed-forward 3D Reconstruction Methods. Each method is shown separately. Training datasets refer to the datasets used for training or fine-tuning the reported model.
Method
Stream- ing
Cam- free
Final
Intermediate
12 views
50 views
Avg. #GS/PT
12 views
50 views
12 / 50 views
ZipSplat r=1
\ding 55
\ding 51
27.65/0.87/0.13
29.13/0.91/0.10
124K/518K
27.95/0.88/0.12
29.29/0.91/0.10
ZipSplat †
\ding 55
\ding 51
25.47/0.82/0.19
26.49/0.86/0.17
17K/21K
26.26/0.84/0.17
26.89/0.86/0.17
TokenGS
\ding 55
\ding 55
32.00/0.94/0.15
31.96/0.94/0.14
262K/262K
30.79/0.90/0.17
32.47/0.94/0.14
GlobalSplat
\ding 55
\ding 55
29.85/0.91/0.13
29.04/0.89/0.16
32K/32K
29.23/0.89/0.14
29.37/0.90/0.15
Appendix
Table 5. Quantitative Comparison on RealEstate10K. Each metric cell reports PSNR ↑ / SSIM ↑ / LPIPS ↓ . Final evaluates all target views after observing the complete context stream, whereas Intermediate evaluates each target view when it first becomes renderable using only previously observed frames. r=1 denotes the default compression ratio of ZipSplat. † adjusts the compression ratio to match our token budget. Notably, MonoGS is not feed-forward.
Method
Stream- ing
Cam- free
Final
Intermediate
12 views
50 views
Avg. #GS/PT
12 views
50 views
12 / 50 views
ZipSplat r=1
\ding 55
\ding 51
20.71/0.50/0.40
21.46/0.54/0.38
124K/518K
20.83/0.50/0.40
21.39/0.53/0.38
ZipSplat †
\ding 55
\ding 51
20.40/0.48/0.43
20.66/0.50/0.43
43K/82K
20.48/0.49/0.43
20.68/0.49/0.43
TokenGS
\ding 55
\ding 55
22.55/0.67/0.42
22.29/0.64/0.48
262K/262K
21.68/0.66/0.41
22.88/0.67/0.43
GlobalSplat
\ding 55
\ding 55
18.12/0.41/0.58
17.25/0.39/0.65
32K/32K
17.50/0.40/0.57
17.85/0.42/0.59
Appendix
Table 6. Quantitative Comparison on Mip-NeRF 360. Each metric cell reports PSNR ↑ / SSIM ↑ / LPIPS ↓ . Final evaluates all target views after observing the complete context stream, whereas Intermediate evaluates each target view when it first becomes renderable using only previously observed frames. r=1 denotes the default compression ratio of ZipSplat. ZipSplat † adjusts its compression ratio to match the primitive budget of our method. Notably, MonoGS is not feed-forward.
Method
Stream- ing
Cam- free
Final
Intermediate
12 views
50 views
Avg. #GS/PT
12 views
50 views
12 / 50 views
ZipSplat r=1
\ding 55
\ding 51
22.25/0.74/0.30
21.46/0.73/0.35
124K/518K
23.21/0.76/0.27
22.84/0.76/0.31
ZipSplat †
\ding 55
\ding 51
21.99/0.74/0.31
21.23/0.73/0.35
41K/71K
22.92/0.76/0.29
22.55/0.76/0.32
TokenGS
\ding 55
\ding 55
28.71/0.91/0.22
28.99/0.91/0.25
262K/262K
29.26/0.92/0.20
30.15/0.91/0.23
GlobalSplat
\ding 55
\ding 55
22.63/0.79/0.34
19.33/0.73/0.39
32K/32K
23.39/0.80/0.30
20.43/0.73/0.35
Appendix
Table 7. Quantitative Comparison on ScanNet++. Each metric cell reports PSNR ↑ / SSIM ↑ / LPIPS ↓ . Final evaluates target views after observing the complete context stream, while Intermediate evaluates each target view using only the context frames available at that step. ZipSplat † uses the S2Tok-derived reduced-budget setting. Notably, MonoGS is not feed-forward.
Figure 8. Qualitative Rendering Results on Mip-NeRF 360 (Final). “GT depth” denotes the pseudo-GT depth produced by DA3 ( Lin et al., 2025 ) .
Figure 9. Qualitative Rendering Results on DL3DV (Intermediate). “GT depth” denotes the pseudo-GT depth produced by DA3 ( Lin et al., 2025 ) .
Figure 10. Qualitative Rendering Results on ScanNet++ (Intermediate). “GT depth” denotes the pseudo-GT depth produced by DA3 ( Lin et al., 2025 ) .
Figure 11. Scene Representation Size Comparisons at Each Step. We report the storage of the representation components identified in the legend, averaged over the evaluated sequences and shown on a logarithmic scale. The curves characterize storage growth over this finite evaluation window, rather than end-to-end GPU memory consumption.
Figure 12. Evolution of the decoded scene on Mip-NeRF 360. For each scene, the top row shows the 12 input frames in streaming order. The bottom row shows the accumulated Gaussian scene, rendered from a fixed overview camera after steps 2, 5, 8, and 11. At inference, each step uses the incoming image and the carried scene state, without rereading previous images. The snapshots illustrate changes in the decoded scene as additional observations are incorporated.
In this work, we revisit several key design choices of modern Transformer-based approaches for feed-forward 3D Gaussian Splatting (3DGS) prediction. We argue that the common practice of regressing Gaussian means as depths along camera rays is suboptimal, and instead propose to directly regress 3D mean coordinates using only a self-supervised rendering loss. This formulation allows us to move from the standard encoder-only design to an encoder-decoder architecture with learnable Gaussian tokens, thereby unbinding the number of predicted primitives from input image resolution and number of views. Our resulting method, TokenGS, demonstrates improved robustness to pose noise and multiview inconsistencies, while naturally supporting efficient test-time optimization in token space without degrading learned priors. TokenGS achieves state-of-the-art feed-forward reconstruction performance on both static and dynamic scenes, producing more regularized geometry and more balanced 3DGS distribution, while seamlessly recovering emergent scene attributes such as static-dynamic decomposition and scene flow.
Jiawei Ren, Michal Jan Tyszkiewicz, Jiahui Huang +1
Streaming recurrent models enable efficient 3D reconstruction by maintaining persistent state representations. However, they suffer from catastrophic forgetting over long sequences due to balancing historical information with new observations. Recent methods alleviate this by deriving adaptive signals from the attention perspective, but they operate on single dimensions without considering temporal and spatial consistency. To this end, we propose a training-free framework termed TTSA3R that leverages both temporal state evolution and spatial observation quality for adaptive state updates in 3D reconstruction. In particular, we devise a Temporal Adaptive Update Module that regulates update magnitude by analyzing temporal state evolution patterns. Then, a Spatial Contextual Update Module is introduced to localize spatial regions that require updates through observation-state alignment and scene dynamics. These complementary signals are finally fused to determine the state updating strategies. Extensive experiments show that TTSA3R achieves competitive performance on standard short-sequence benchmarks and provides substantially stronger robustness on extended sequences. On NRGBD, as sequences extend from 50 to 250 frames, TTSA3R exhibits only a 1.33x error increase, compared with over 4x degradation for CUT3R. This highlights the practical value of temporal-spatial adaptive updates for long-term reconstruction stability. Our code is available at https://github.com/anonus2357/ttsa3r.
Feed-forward 3D Gaussian Splatting (3DGS) reconstructs a scene in a single forward pass, replacing per-scene optimization with a network trained across many scenes. Its quality, however, degrades sharply as the number of input images drops. The bottleneck is upstream of the reconstruction heads: from a few unposed views, the internal representation they read carries no evidence for unobserved regions, leaving holes, floaters, and blur. The common remedy supplies that evidence as pixels, synthesizing extra views with an image or video generator and re-encoding them, which is costly and not 3D-consistent by construction. We instead densify the evidence itself. We present DensiTok, a plug-in module for pretrained feed-forward 3DGS models that densifies their internal geometry tokens directly, making a frozen backbone behave as though it had observed many more views than it was given. DensiTok compresses those tokens into a compact latent space, completes the latents of the unobserved viewpoints in a single flow-matching step conditioned on camera geometry, and decodes them back into tokens that the original reconstruction heads. The same module design can be integrated into different pretrained predictors while keeping each backbone and its reconstruction heads frozen. Completion in a low-dimensional latent space requires no image synthesis or additional encoder passes. Across three pretrained backbones and two benchmarks, DensiTok consistently improves sparse-view reconstruction and recovers much of the gap to dense-view reconstruction.
Minhyeok Lee, Jungho Lee, Minseok Kang +3
Yonsei University · NAVER AI Lab · Korea Institute of Science and Technology (KIST)