Organizations: Department of Precision Engineering, Graduate School of Engineering, The University of Tokyo, 5-1-5 Kashiwanoha, Kashiwa, Chiba 277-8563, Japan. · Department of Human and Engineered Environmental Studies, Graduate School of Frontier Sciences, The University of Tokyo, Kashiwa, Chiba, Japan
Tracking any point of a dynamic scene in metric 3D - in absolute meters, not up to an unknown scale - underpins 3D and 4D reconstruction, robot navigation, and autonomous driving, where decisions are made in meters, not pixels. Our objective is a 3D point tracker accurate in those absolute terms and operating within a single commodity GPU, pose-free, monocular budget. Our method rests on one observation: once a point's 2D image trajectory is fixed, the quantity that governs its metric accuracy is the depth along its pixel ray. Rather than learning tracking end-to-end, we therefore compose two frozen front-ends - dense optical flow for 2D correspondence and a monocular metric-depth network for the third dimension - and learn only the residual they cannot supply: that depth, refined by a compact state space model (Mamba-3) conditioned on appearance features (DINOv3). A state space model rather than the transformers the strongest 3D trackers adopt is what makes a single-GPU budget attainable: it summarises a track in a fixed-size recurrent state whose memory cost is constant in the number of frames, whereas attention requires a key-value cache that grows linearly with them. On the TAPVid-3D minival benchmark our best configuration attains the highest absolute metric accuracy among methods evaluated under identical conditions (mean metric Average Jaccard, 0.256), exceeding strong feed-forward trackers, while a companion analysis, reproduced with each competitor's own evaluator, explains why several published trackers lose most of their accuracy under this budget.
Figures & tables
Figure 1 : Our metric 3D point tracking on a driving scene. (a) Query points are propagated in 2D and overlaid on the video: each track is drawn as one curve from its start up to the frame shown, ending in a dot at the point’s position in that frame; the panel is therefore the tracker’s state at one instant rather than a summary of the whole clip. (b) The same points are lifted and refined into metric 3D. The axis values are absolute meters in the DriveTrack [ 1 ] ground-truth frame, read off the data and not re-centred: Z is distance from the camera, so this clip lies at 13 – 15 m, while X and Y are offsets about the optical axis and stay near zero. Solid = predicted, dashed = ground truth, dots = anchor frame; the view is cropped to a 3 m cube centred on the anchor points so that individual trajectories separate, and tracks that run outside it are clipped. Predicted tracks closely follow ground truth along all three axes X , Y , Z , not only in the X - Y image-plane projection that a method with correct 2D tracking but wrong depth could also get right. Our depth-along-ray refiner corrects per-point depth so that the recovered tracks lie on the true surfaces in all three dimensions, not just when reprojected to 2D. The 3D view is oriented like the image of (a): X to the right, Y downward, Z depth.
Figure 2 : Overall pipeline. The flow front-end (SEA-RAFT or WAFT) and the metric-depth backbone (DA3-l or DA3-g, § 4 ) are both frozen and swappable: swapping either only changes which frozen network fills these two grey boxes, and re-training uses an identical architecture, only the frozen depth cache changes. Flow chaining composes the front-end’s consecutive-frame flow into a full-length 2D track and, by a forward–backward consistency check, sets each point’s visibility v —the “ + vis” of its output box (§ 3.4 ). Grey = frozen / training-free; the blue box is the only learned module (under 1 M params); its internals are the Mamba-3 refiner (Fig. 5 5(a) ) and the vmamba3 refiner (Fig. 6 6(a) ). The Mamba-3 refiner leaves the frozen flow front-end’s 2D track ut (and visibility) untouched; the vmamba3 refiner additionally applies a small bounded 2D correction to it.
Figure 3 : Mamba-3 state space duality. The three-term recurrence of Eq. ( 2 ) (left) is exactly a masked linear attention (right); “ ≡ ” denotes this duality—the two sides are the same computation, not merely similar. The structured mask L is the product of a full lower-triangular decay matrix ∏α (every cell generically non-zero) and a two-band matrix carrying a main ( γ ) diagonal and, unique to Mamba-3, a sub-diagonal ( β ) band, Eq. ( 5 )—so L itself is densely lower-triangular, not two isolated bands, and replaces the softmax of ordinary attention. Blank (upper-triangular) cells are zero in both matrices.
Figure 4 : The causal mask, and the four ways of 3.2.2 to remove causality from Eq. ( 4 ). (a) Causal SSD masks the future. (b) Summing forward and reverse scans yields a symmetric, all-pairs mask—non-causal, but still weighted by distance along the scan order ( 3.2.2 ). (c) Four-directional scanning sums a second, column-major symmetric mask on top of (b)’s row-major one, removing the direction bias a single scan order leaves behind. (d) VSSD-1pool [ 25 ] collapses the T×T mask to a single T -vector m , so Eq. ( 6 ) runs in O(TND) with an O(ND) state independent of sequence length, but cannot represent Mamba-3’s β -band. (e) VSSD-2pool stacks a second, independently-read vector m(2) under m(1) , recovering β -band-like capacity at 2× VSSD-1pool’s cost, Eq. ( 9 ); derived in 3.2.2 and measured in Table 1 . Coloured cells denote non-zero mask entries; blank cells are zero (masked out). Where a panel sums two contributions, the second is drawn in orange—the same orange Fig. 3 uses for the β -band, which in (e) is literally what m(2) is motivated by.
Figure 5 : Mamba-3 3D refiner. (a) Architecture — the inside of the blue box in Fig. 2 . r=(rayx,rayy) is the frozen pixel ray from the 2D track ut ; z/zr=zraw/zref ; v= visibility. Bypass : zraw (the frozen metric-depth backbone’s raw depth) skips the SSM entirely and is multiplied back in via zt=zraweΔlogz , Eq. ( 11 ); the SSM is therefore a relative corrector , not an absolute-depth predictor. causal × 2 : two causal Mamba-3 layers (past-only; enables streaming). SSD splits the 64-dim state into 4 heads (16-dim each), analogous to multi-head attention but computed as a linear recurrence. (b) Ray-invariance guarantee. Red marks what the refiner changes, black what it leaves alone — the same convention as Fig. 6 6(b) . The grey wedge is the camera: apex at the camera centre, right edge the image plane; the ray leaves the centre through the tracked pixel and continues into the scene. The pixel, and hence the ray direction, is frozen, so only the depth slides along the ray and the 2D projection is unchanged: the refiner moves depth alone, leaving the 2D track and visibility invariant.
Figure 6 : vmamba3 3D refiner. (a) Architecture, extending Fig. 5 5(a) — again the inside of the blue box in Fig. 2 . Orange nodes are new relative to the Mamba-3 refiner. The pixel position u0 (from the frozen optical-flow front-end; Fig. 2 ) sets the ray direction r(u0) , the depth-patch centre, and the DINOv3 sample location. The three inputs are embedded and fed through the same causal Mamba-3 SSD stack, now carrying appearance and local-geometry context. A second, zero-initialised head predicts Δu=2tanh(⋅) (bounded to ±2 px); depth is re-sampled at ut=u0+Δu before Δlogz rescales it. Both heads start at zero, and the vmamba3 refiner therefore starts exactly at the Mamba-3 refiner’s output. (b) Safe, appearance-conditioned 2D correction. Red marks what changes, black what does not : the shifted ray and zt are both red because this variant moves both, whereas Fig. 5 5(b) moves only zt . The camera wedge is as in Fig. 5 5(b) ; both rays leave the centre through their own pixel. Two new inputs — DINOv3 appearance and a local 5×5 depth patch, both orange — give the model the context to predict a 2D shift Δu safely: unlike a naive nudge, which crosses depth discontinuities, this correction is conditioned on local appearance and geometry. Depth is then re-sampled at the shifted pixel u0+Δu before Δlogz rescales it.
Figure 7 : Depth scale refiner (§ 3.9 ). It reads the metric depth map directly: no tracker appears in the path, and the front-end cannot change what it learns. Two independent branches summarise each frame’s log depth: a small CNN over it pooled to 642 , giving 64 features, and the 5 th, 15 th, …, 95 th percentiles taken over every pixel, one vector of ten numbers per frame. Pooling and then averaging inside the CNN discards the shape of the depth distribution, which the percentiles restore. The two concatenate to 74 and project to one 128 -dimensional token per frame; the state space layer mixes the F tokens and a linear head emits one scalar Δsf , rescaling that frame’s depth z←zraweΔsf before the refiner of § 3.8 runs.
Figure 8 : Visibility computation (§ 3.10 ). Both paths read one pair of flow vectors per point per frame and nothing else — no depth, no appearance, nothing from the refiner. Upper, the rule every earlier result was scored with : follow the point one frame forward by ft and back by bt+1 , and call it occluded if it misses its start by more than α(∥ft∥+∥bt+1∥)+β , with α and β two hand-chosen constants ( 0.05 and 1 px) rather than anything fitted; the backward sweep applies the same test to ft−1 and bt . The result is latched, vt+1=vt∧okt , and one failure therefore fixes the point as occluded for every later frame. Lower, the learned head : eight scalars per point per frame, all functions of ft and bt , embedded to 64 ; two state space layers mix along time only, never across points; a linear head and a sigmoid follow. The outputs differ in kind — one bit against a probability, read as visible above 0.5 — and only the head can mark a point visible again after an occlusion, which 52 – 66% of ground-truth points require.
Method
drivetrack
pstudio
adt
mean
Published methods (external)
SpatialTrackerV2 ∗
0.017
0.012
0.025
0.018
DELTA + DA3-l §
0.130
0.141
0.152
0.141
DELTAv2 + DA3-l §
0.130
0.136
0.155
0.140
TAPIP3D + DA3-l †
0.065
0.023
0.004
0.030
TAPIP3D + MegaSaM ‡
0.000
0.003
0.004
0.002
Table 1 : TAPVid-3D minival, median-scaled 3D-AJ (standard leaderboard metric; higher better) — DA3-l depth ( DA3Metric-Large , 0.35 B). Bold = best per column. In the row names, vmamba3 is the appearance- and geometry-conditioned refiner of § 3.8 (the v is for vision) and vmamba3-2pool is that refiner with its temporal mixer replaced by VSSD-2pool ( 3.2.2 ). All numbers are produced by our own evaluation pipeline on the same hardware. DA3-g counterpart: Table 3 .
Method
drivetrack
pstudio
adt
mean
Published methods (external)
SpatialTrackerV2 ∗
0.008
0.192
0.179
0.126
DELTA + DA3-l §
0.006
0.199
0.344
0.183
DELTAv2 + DA3-l §
0.005
0.190
0.345
0.180
TAPIP3D + DA3-l
0.006
0.131
0.164
0.100
TAPIP3D + MegaSaM †
0.000
0.000
0.000
0.000
Table 2 : TAPVid-3D minival, absolute metric-AJ (no scaling, fixed-meter thresholds 1 cm–2.56 m; higher better) — DA3-l depth ( DA3Metric-Large , 0.35 B). Bold = best per column. Row names and evaluation conditions as in Table 1 . Highest mean across all four tables: WAFT+DA3-l+vmamba3-2pool with the visibility head ( 0.256 ); no DA3-g pipeline reaches it (§ 5.1 ). DA3-g counterpart: Table 4 .
Method (DA3-g)
drivetrack
pstudio
adt
mean
Published methods (external)
DELTA + DA3-g
0.137
0.049
0.141
0.109
DELTAv2 + DA3-g
0.131
0.046
0.148
0.108
SpatialTrackerV2 + DA3-g
0.018
0.008
0.027
0.017
TAPIP3D + DA3-g
0.057
0.011
OOM
0.034
Ours
Table 3 : TAPVid-3D minival, median-scaled 3D-AJ (leaderboard metric; higher better) — DA3-g depth ( DA3Nested-Giant-Large , 1.40 B): every DA3-consuming method re-run on the DA3-g backbone, all other components unchanged. Bold = best per column. “OOM” = TAPIP3D adt out-of-memory. DA3-l counterpart: Table 1 .
Method (DA3-g)
drivetrack
pstudio
adt
mean
Published methods (external)
DELTA + DA3-g
0.164
0.194
0.324
0.227
DELTAv2 + DA3-g
0.159
0.171
0.319
0.216
SpatialTrackerV2 + DA3-g
0.043
0.184
0.170
0.132
TAPIP3D + DA3-g
0.110
0.148
OOM
0.129
Ours
Table 4 : TAPVid-3D minival, absolute metric-AJ (fixed-meter thresholds; higher better) — DA3-g depth ( DA3Nested-Giant-Large , 1.40 B): every DA3-consuming method re-run on the DA3-g backbone, all other components unchanged. Bold = best per column. Best DA3-g mean: our WAFT+DA3-g+scale refiner+2pool with the visibility head ( 0.229 ) , which § 4.3 reports as matching, not exceeding, the best external DA3-g method, DELTA + DA3-g ( 0.227 ). Every DA3-g configuration here remains below the DA3-l best ( 0.256 , Table 2 ), which is this table’s purpose. “OOM” = TAPIP3D adt out-of-memory. DA3-l counterpart: Table 2 .
Method
fps
trainable params
Published methods (external)
TAPIP3D ⋆
3.1
25.8 M
SpatialTrackerV2 (s_wind = 60)
4.0
≈1.23 B ‡
DELTA (DenseTrack3D) ⋆
7.7
59.2 M
DELTAv2 (DenseTrack3Dv2) ⋆
6.9
51.4 M
TrackCraft3R §
0.06
1.3 B §
Table 5 : Inference throughput on RTX 4080 (12 GB), measured on TAPVid-3D minival with the DA3-l depth backbone ( DA3Metric-Large , 0.35 B). Our lightweight SSM refiner adds modest overhead over the baseline while remaining comparable in throughput to the DELTA family, 1.6 – 2× faster than TAPIP3D and SpatialTrackerV2, and two orders of magnitude faster than the video-diffusion tracker TrackCraft3R; its learned module ( 0.44 M, 0.61 M with the two-pool mixer, 0.79 M with the visibility head) is 40 – 3000× smaller than the external trackers ( 25.8 M– 1.3 B). External param counts are summed from the official released checkpoints (no paper states them). Bold marks the best value in each column. Throughput is an end-to-end average over the evaluation run; a row whose figure is instead carried from another row, measured on one subset, or derived is marked in the note below. DA3-g counterpart: Table 6 .
Method
fps
trainable
Published methods (external)
TAPIP3D + DA3-g ⋆
3.1
25.8 M
SpatialTrackerV2 (s_wind = 60)
4.0
≈1.23 B ‡
DELTA + DA3-g ⋆
7.7
59.2 M
DELTAv2 + DA3-g ⋆
6.9
51.4 M
Ours & training-free baselines
Table 6 : Inference throughput and model size on the DA3-g pipeline ( DA3Nested-Giant-Large , 1.40 B) — identical to Table 5 except for the depth backbone. Online fps is unchanged from the DA3-l table because depth is read from the offline cache; only the one-time depth precompute is slower for the giant backbone. Every method here reads the same 1.40 B DA3-g depth, so that backbone is common to all of them and is not what separates them; what differs is the trainable count, where ours ranges from 0.44 M for the refiner alone to 1.37 M with the scale refiner and visibility head added. Bold marks the best value in each column. Throughput is an end-to-end average over the evaluation run; a row whose figure is instead carried from another row, measured on one subset, or derived is marked in the note below. Columns and external counts as in Table 5 ; TrackCraft3R could not be re-run on DA3-g ( 12 GB OOM) and is omitted here.
Figure 9 : 3D trajectory comparison, DELTA + DA3-l (left of each pair) against ours (right; the panel labels name the configuration), one clip per subset: (a) drivetrack, (b) pstudio, (c) adt. Both use the same DA3-l depth, which isolates the tracker. Solid = predicted, dashed = ground truth; a prediction that leaves the plotted volume simply does not appear. The red square marks the track our method improves the most over DELTA, the same one in both panels, and the inset shows that region at twice the size. The difference is plain in (a) and (b) but hard to see in (c): adt is the one subset of the three where DELTA scores above us (Table 2 ). § 4.6 discusses each subset.
Figure 10 : DA3-l vs. DA3-g depth error at ground-truth pixels, TAPVid-3D minival. Here scale is the single positive scalar s by which every depth in a frame (panel a) or a whole clip (panel b) is multiplied to best match ground truth in a least-squares sense; dividing it out isolates spatial shape error from any error in that one overall multiplier. DA3-g’s spatial shape is as good as or better than DA3-l’s, but unlike DA3-l, its scale flickers from frame to frame. Throughout, the y -axis is a scale-invariant log error, ∣log(scale⋅pred/gt)∣ : 0 means the prediction exactly matches ground truth once a scale is removed, and larger values mean more error. (a) Median spatial shape error vs. distance (per-frame scale removed): DA3-g (solid) matches or exceeds DA3-l (dashed) almost everywhere. (b) Mean error under per-frame vs. per-clip scale alignment. Read each backbone across the two alignments: per-frame ≈ per-clip means the scale is temporally stable, and a large gap between them means the per-frame scale drifts. DA3-l barely changes, whereas DA3-g’s pstudio error rises 4.3× when only one scale per clip is allowed, i.e. its per-frame scale drifts over time. The near-field defect is temporal scale instability, not spatial depth noise.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Table 1 : CIFAR-10 from scratch, one fixed training protocol, single seed per cell, at T=65 (patch size 4 ) and T=1025 (patch size 1 ). Every operator is crossed with every positional encoding and their combinations, so that each encoding’s own effect can be read separately. Enc. takes one of: “—” for no encoding; “2-D RoPE” for the data- independent 2-D rotary embedding on B,C ( A.2 ); “rotary” for Mamba-3’s own data- dependent complex-SSM rotary; “1-turn rotary” for that same rotary with its angular spread pinned to one turn instead of growing with T ; and the two names joined by “ + ” for the corresponding combination. The 1-turn rows are motivated at T=1025 , where the growing form collapses ( A.4 ), and are now measured at both lengths: pinning costs about five points at T=65 , where the growing rotary is still short enough to work, and gains about five at T=1025 . Repeat runs of one configuration differed by 0.67 points; accuracy gaps below ∼0.7 are not meaningful; accuracy is therefore printed to one decimal. “n/a” marks the CNN’s T=1025 cells: a CNN has no notion of tokens and so no sequence length, and is reported at T=65 only. Latency and peak memory are a single forward pass at batch 128 on one RTX 4080, with the optimiser released, i.e. an inference footprint.
angular spread
VSSD-1pool
VSSD-2pool
n=0.1
66.0
65.4
n=0.25
65.1
66.6
n=0.5
64.7
64.8
n=1
59.0
59.7
n=2
58.8
59.9
n=4
60.6
59.4
Appendix
Table 2 : Pinned rotary spread on CIFAR-10 at T=1025 , under Table 1 ’s protocol (patch size 1 , 30 epochs, no 2-D RoPE, single seed per cell). n is the total angular spread in turns across the sequence, θj=2πnj/T . Bold marks the best pinned spread in each column. The last three rows are the corresponding cells of Table 1 , repeated for comparison: default rotary is Mamba-3’s own rotary left unmodified, whose angle accumulates to a spread of n≈160 at this length ( A.4 ); that is why it collapses both operators to near chance and is the reason this sweep exists.
Mixer
abs. rel. ↓
RMSE (m) ↓
log10↓
δ<1.25↑
softmax (DA3-SMALL) ⋆
0.032
0.061
0.014
1.000
VSSD-2pool (iv)
0.051
0.101
0.022
0.996
bidirectional (i)
0.055
0.103
0.024
0.999
VSSD-1pool (iii)
0.060
0.118
0.026
0.976
Appendix
Table 3 : Metric depth on the held-out ETH3D [ 24 ] terrains scene at matched parameters ( ≈22 M): DA3-SMALL against the same network with only its self-attention mixer replaced. One training run per row, scored on the same twelve views. The three operator rows fall between 0.051 and 0.060 absolute relative error, a range narrower than the largest within-row standard deviation across the twelve evaluated views ( 0.017 ); they are not separated by this scene and their ordering is not a ranking. Bold marks the best value in each column. Writing d for predicted and d∗ for ground-truth depth, and averaging over the valid pixels of an image and then over images: abs. rel. =mean∣d−d∗∣/d∗ , a relative error, under which a given metric error counts less at large depth; RMSE =(mean(d−d∗)2)1/2 in metres, which by squaring weights the worst pixels most; log10=mean∣log10(d/d∗)∣ , which scores a factor-of- k error equally at every depth; and δ<1.25 , the fraction of pixels with max(d/d∗,d∗/d)<1.25 , i.e. an inlier rate rather than an error. Every row is median-aligned first— d is scaled so that its median matches d∗ ’s—because monocular depth is recovered only up to a global factor. The rows did not receive the same training. DA3-SMALL is evaluated zero-shot, with no ETH3D training of any kind, whereas each vmamba3 row had a mixer distilled, a bridge trained and the depth head fine-tuned on ETH3D ( A.5 ); their higher δ<1.25 follows from that extra training and is not evidence of a better backbone.
University of Bologna, Italy · Faculty of Dentistry, The University of Hong Kong, China · School of Instrument Science and Engineering, Southeast University, China +1