Streaming 3D reconstruction requires more than a sequence of geometric predictions: it requires a persistent scene state that can incorporate new evidence and remain renderable as observations arrive. Latent spatial tokens offer a promising representation for this purpose, but constructing them from an image collection leaves open how to maintain them online, where each observation may both revisit known regions and reveal new content. We introduce S2Tok, a feed-forward framework that maintains a size-adaptive, persistent scene state from uncalibrated image streams. Its central idea is to distinguish updates to the existing representation from selective expansion. A spatially informed transformer integrates each incoming observation with the persistent scene tokens, while a learned admission module selectively expands the representation to limit redundant storage. A hierarchical decoder and Gaussian head convert the evolving state into non-pixel-aligned 3D Gaussians, enabling novel-view rendering without caching previous frames. Experiments across four benchmarks demonstrate competitive streaming rendering quality with compact Gaussian representations. These results support latent spatial tokens as a persistent computational state for online 3D reconstruction, combining learned scene updates with explicit Gaussian rendering.
Figures & tables
Figure 1. S2Tok on Various Scenarios. We show the final 3DGS scenes reconstructed by S2Tok in a streaming manner across diverse scenarios. More streaming reconstruction samples are in Fig. 12 .
Figure 2. S2Tok Pipeline. The encoder maps the new frame Ii to tokens fi . The SS-Transformer jointly updates frame tokens and existing spatial tokens Si−1 . The admission module appends selected frame tokens to form Si , which is carried to the next step. The Hierarchical Decoder and Gaussian Head decode and predict Gaussian attributes. ⊕ denotes concatenation.
Figure 3. Stream-Spatial Transformer. Scale token s , spatial tokens Si−1 , and frame tokens fi attend to one another in Aglobal . Scene-local attention Ascene then updates s and Si−1 , while frame-local attention Aframe updates fi .
Figure 4. Hierarchical Decoder. Hierarchical tokens SIII , SII , and SI from the SS-Transformer are processed individually through self-attention layers and then fused through residual addition + into the decoding stream, which is initialized from SIII . A self-attention layer allows one more update after each residual addition.
Method
Stream- ing
Cam- free
Final
Intermediate
Time (ms/v)
Mem. (GiB)
12 views
50 views
Avg. #GS/PT
12 views
50 views
12 / 50 views
ZipSplat r=1
\ding 55
\ding 51
23.04/0.69/0.24
21.94/0.63/0.28
124K/518K
23.26/0.70/0.25
22.13/0.62/0.29
–
–
ZipSplat †
\ding 55
\ding 51
21.79/0.64/0.31
20.88/0.58/0.35
36K/66K
21.74/0.64/0.32
20.81/0.57/0.37
–
–
TokenGS
\ding 55
\ding 55
21.74/0.69/0.40
20.53/0.63/0.50
262K/262K
22.64/0.73/0.34
21.54/0.65/0.45
–
–
GlobalSplat
\ding 55
\ding 55
17.71/0.47/0.55
15.92/0.40/0.63
32K/32K
18.65/0.52/0.49
17.01/0.44/0.58
–
–
Table 1. Quantitative Comparison on DL3DV. Each metric cell reports PSNR ↑ / SSIM ↑ / LPIPS ↓ . r=1 denotes the default compression ratio of ZipSplat. † adjusts the ratio to match our token budget. Avg. represents the average across all sequences. Time and Mem. report per-view reconstruction time and peak GPU memory of the streaming methods in the 50-view setting, measured on a single NVIDIA A100 80GB. Notably, MonoGS is not feed-forward.
Figure 5. Qualitative Rendering Results on DL3DV (Final). “GT depth” denotes the pseudo-GT depth produced by DA3 ( Lin et al., 2025 ) .
Method
Stream- ing
Cam- free
RealEstate10K (12 Views)
Mip-NeRF 360 (50 Views)
Final
Avg. #GS/PT
Intermediate
Final
Avg. #GS/PT
Intermediate
ZipSplat r=1
\ding 55
\ding 51
27.65/0.87/0.13
124K
27.95/0.88/0.12
21.46/0.54/0.38
518K
21.39/0.53/0.38
ZipSplat †
\ding 55
\ding 51
25.47/0.82/0.19
17K
26.26/0.84/0.17
20.66/0.50/0.43
82K
20.68/0.49/0.43
TokenGS
\ding 55
\ding 55
32.00/0.94/0.15
262K
30.79/0.90/0.17
22.29/0.64/0.48
262K
22.88/0.67/0.43
GlobalSplat
\ding 55
\ding 55
29.85/0.91/0.13
32K
29.23/0.89/0.14
17.25/0.39/0.65
32K
17.85/0.42/0.59
AnySplat
\ding 55
\ding 51
23.39/0.80/0.14
663K
25.87/0.85/0.12
18.85/0.45/0.36
2850K
19.18/0.51/0.32
Table 2. Quantitative Comparison on RealEstate10K and Mip-NeRF 360. Each metric cell reports PSNR ↑ / SSIM ↑ / LPIPS ↓ . Results are evaluated under 12-view and 50-view settings. Detailed comparisons are in Tab. 5 and Tab. 6 in the Appendix. r=1 is the default compression ratio of ZipSplat. † adjusts the compression ratio to match our token budget. Notably, MonoGS is not feed-forward.
Figure 6. Qualitative Rendering Results on ScanNet++ (Final). “GT depth” denotes the pseudo-GT depth produced by DA3 ( Lin et al., 2025 ) .
Figure 7. Qualitative Rendering Results on Mip-NeRF 360 (Intermediate). “GT depth” denotes the pseudo-GT depth produced by DA3 ( Lin et al., 2025 ) .
Variant
Final
Intermediate
Avg. #GS
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
w/o RoPE3D
17.74
0.4117
0.5166
18.85
0.4452
0.4794
65K
w/o Admission
17.50
0.4076
0.5476
20.39
0.5087
0.4149
518K
w/o SCST
18.13
0.4281
0.4997
19.67
0.4757
0.4463
120K
w/o SCST(matched budget)
17.90
0.4216
0.5112
19.03
0.4541
0.4730
66K
w/o scale token s
17.89
0.4214
0.5126
19.01
0.4543
0.4765
66K
Table 3. Ablation experiments on the 50-view DL3DV evaluation set.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Streaming Setting
Method
Streaming
Latent Spatial Token/Memory
Cache/Window- free
Pixel- free Prediction
Reconstruction Kernel
Pose- free
Training Dataset(s)
AnySplat
\ding 55
\ding 55
non- streaming
\ding 55
GS
\ding 51
Hypersim, ARKitScenes, BlendedMVS, ScanNet++, CO3D-v2, Objaverse, Unreal4K, WildRGB-D, and DL3DV
Spann3R
\ding 51
\ding 51
\ding 55
\ding 55
Point
\ding 51
Habitat, ScanNet, ScanNet++, ARKitScenes, BlendedMVS, and CO3D-v2 †
Table 4. Detailed Comparison of Related Feed-forward 3D Reconstruction Methods. Each method is shown separately. Training datasets refer to the datasets used for training or fine-tuning the reported model.
Method
Stream- ing
Cam- free
Final
Intermediate
12 views
50 views
Avg. #GS/PT
12 views
50 views
12 / 50 views
ZipSplat r=1
\ding 55
\ding 51
27.65/0.87/0.13
29.13/0.91/0.10
124K/518K
27.95/0.88/0.12
29.29/0.91/0.10
ZipSplat †
\ding 55
\ding 51
25.47/0.82/0.19
26.49/0.86/0.17
17K/21K
26.26/0.84/0.17
26.89/0.86/0.17
TokenGS
\ding 55
\ding 55
32.00/0.94/0.15
31.96/0.94/0.14
262K/262K
30.79/0.90/0.17
32.47/0.94/0.14
GlobalSplat
\ding 55
\ding 55
29.85/0.91/0.13
29.04/0.89/0.16
32K/32K
29.23/0.89/0.14
29.37/0.90/0.15
Appendix
Table 5. Quantitative Comparison on RealEstate10K. Each metric cell reports PSNR ↑ / SSIM ↑ / LPIPS ↓ . Final evaluates all target views after observing the complete context stream, whereas Intermediate evaluates each target view when it first becomes renderable using only previously observed frames. r=1 denotes the default compression ratio of ZipSplat. † adjusts the compression ratio to match our token budget. Notably, MonoGS is not feed-forward.
Method
Stream- ing
Cam- free
Final
Intermediate
12 views
50 views
Avg. #GS/PT
12 views
50 views
12 / 50 views
ZipSplat r=1
\ding 55
\ding 51
20.71/0.50/0.40
21.46/0.54/0.38
124K/518K
20.83/0.50/0.40
21.39/0.53/0.38
ZipSplat †
\ding 55
\ding 51
20.40/0.48/0.43
20.66/0.50/0.43
43K/82K
20.48/0.49/0.43
20.68/0.49/0.43
TokenGS
\ding 55
\ding 55
22.55/0.67/0.42
22.29/0.64/0.48
262K/262K
21.68/0.66/0.41
22.88/0.67/0.43
GlobalSplat
\ding 55
\ding 55
18.12/0.41/0.58
17.25/0.39/0.65
32K/32K
17.50/0.40/0.57
17.85/0.42/0.59
Appendix
Table 6. Quantitative Comparison on Mip-NeRF 360. Each metric cell reports PSNR ↑ / SSIM ↑ / LPIPS ↓ . Final evaluates all target views after observing the complete context stream, whereas Intermediate evaluates each target view when it first becomes renderable using only previously observed frames. r=1 denotes the default compression ratio of ZipSplat. ZipSplat † adjusts its compression ratio to match the primitive budget of our method. Notably, MonoGS is not feed-forward.
Method
Stream- ing
Cam- free
Final
Intermediate
12 views
50 views
Avg. #GS/PT
12 views
50 views
12 / 50 views
ZipSplat r=1
\ding 55
\ding 51
22.25/0.74/0.30
21.46/0.73/0.35
124K/518K
23.21/0.76/0.27
22.84/0.76/0.31
ZipSplat †
\ding 55
\ding 51
21.99/0.74/0.31
21.23/0.73/0.35
41K/71K
22.92/0.76/0.29
22.55/0.76/0.32
TokenGS
\ding 55
\ding 55
28.71/0.91/0.22
28.99/0.91/0.25
262K/262K
29.26/0.92/0.20
30.15/0.91/0.23
GlobalSplat
\ding 55
\ding 55
22.63/0.79/0.34
19.33/0.73/0.39
32K/32K
23.39/0.80/0.30
20.43/0.73/0.35
Appendix
Table 7. Quantitative Comparison on ScanNet++. Each metric cell reports PSNR ↑ / SSIM ↑ / LPIPS ↓ . Final evaluates target views after observing the complete context stream, while Intermediate evaluates each target view using only the context frames available at that step. ZipSplat † uses the S2Tok-derived reduced-budget setting. Notably, MonoGS is not feed-forward.
Figure 8. Qualitative Rendering Results on Mip-NeRF 360 (Final). “GT depth” denotes the pseudo-GT depth produced by DA3 ( Lin et al., 2025 ) .
Figure 9. Qualitative Rendering Results on DL3DV (Intermediate). “GT depth” denotes the pseudo-GT depth produced by DA3 ( Lin et al., 2025 ) .
Figure 10. Qualitative Rendering Results on ScanNet++ (Intermediate). “GT depth” denotes the pseudo-GT depth produced by DA3 ( Lin et al., 2025 ) .
Figure 11. Scene Representation Size Comparisons at Each Step. We report the storage of the representation components identified in the legend, averaged over the evaluated sequences and shown on a logarithmic scale. The curves characterize storage growth over this finite evaluation window, rather than end-to-end GPU memory consumption.
Figure 12. Evolution of the decoded scene on Mip-NeRF 360. For each scene, the top row shows the 12 input frames in streaming order. The bottom row shows the accumulated Gaussian scene, rendered from a fixed overview camera after steps 2, 5, 8, and 11. At inference, each step uses the incoming image and the carried scene state, without rereading previous images. The snapshots illustrate changes in the decoded scene as additional observations are incorporated.