cs.CVSep 27, 2026

3D Point Tracking with State Space Models

Authors: Masahiro Ogawa, Qi An, Atsushi Yamashita

Organizations: Department of Precision Engineering, Graduate School of Engineering, The University of Tokyo, 5-1-5 Kashiwanoha, Kashiwa, Chiba 277-8563, Japan. · Department of Human and Engineered Environmental Studies, Graduate School of Frontier Sciences, The University of Tokyo, Kashiwa, Chiba, Japan

Abstract

Tracking any point of a dynamic scene in metric 3D - in absolute meters, not up to an unknown scale - underpins 3D and 4D reconstruction, robot navigation, and autonomous driving, where decisions are made in meters, not pixels. Our objective is a 3D point tracker accurate in those absolute terms and operating within a single commodity GPU, pose-free, monocular budget. Our method rests on one observation: once a point's 2D image trajectory is fixed, the quantity that governs its metric accuracy is the depth along its pixel ray. Rather than learning tracking end-to-end, we therefore compose two frozen front-ends - dense optical flow for 2D correspondence and a monocular metric-depth network for the third dimension - and learn only the residual they cannot supply: that depth, refined by a compact state space model (Mamba-3) conditioned on appearance features (DINOv3). A state space model rather than the transformers the strongest 3D trackers adopt is what makes a single-GPU budget attainable: it summarises a track in a fixed-size recurrent state whose memory cost is constant in the number of frames, whereas attention requires a key-value cache that grows linearly with them. On the TAPVid-3D minival benchmark our best configuration attains the highest absolute metric accuracy among methods evaluated under identical conditions (mean metric Average Jaccard, 0.256), exceeding strong feed-forward trackers, while a companion analysis, reproduced with each competitor's own evaluator, explains why several published trackers lose most of their accuracy under this budget.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations

    Sep 24, 2026Ayush Jain, Sreeharsha Paruchuri, Ishita Gupta +53D Tracking3D Scene Understanding

  2. TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking

    May 12, 2026Jisu Nam, Jahyeok Koo, Soowon Son +43D TrackingVideo Diffusion Models

  3. MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos

    Jul 16, 2026Ziren Gong, Xiaohan Li, Fabio Tosi +4Feed-Forward 3D ReconstructionCamera Control