cs.CVSep 29, 2026

TrackFish3D: Self-Supervised 3D Tracking of Schooling Fish from Multi-view Videos

Authors: Patt Phurtivilai, Zhiyang Dou, Yifan Wu, Kinfung Chu, Yuan Liu, Lei Yang, Wenping Wang, Taku Komura

Organizations: The University of Hong Kong, Hong Kong, China · Centre for Transformative Garment Production · Hong Kong University of Science and Technology, Hong Kong, China · Texas A&M University, College Station, TX, US

Abstract

Quantifying collective fish behavior requires accurate trajectories, yet multi-view 3D tracking remains challenging due to frequent occlusions, visually similar individuals, and the long-standing scarcity of identity annotations. We present TrackFish3D, a geometry-driven self-supervised framework for dense multi-camera 3D tracking of schooling fish. Instead of relying on appearance-based re-identification or manually annotated identities, TrackFish3D turns calibrated multi-view geometry into supervision: triangulation and reprojection consistency provide pseudo-associations, while a geometric encoder and global association transformer learn all-to-all cross-view correspondence within each frame. To make these associations identity-aware, TrackFish3D introduces a self-supervised contrastive objective that separates co-visible individuals in the embedding space, together with a temporal predictor that preserves identities and bridges short occlusions across frames. The resulting model is trained once on unlabeled footage and applied directly to unseen test videos, requiring no cross-view identity labels, temporal annotations, 3D ground truth, appearance features, or test-time optimization. On our benchmark, TrackFish3D improves 3D Multi-Object Tracking Accuracy from 87.7% for the strongest baseline to 95.8%. On the 3D-ZeF zebrafish benchmark, it achieves 81.1% MOTA, compared with 77.4% for the best geometric baseline. TrackFish3D also generalizes beyond fish, achieving strong results on real-world bird tracking.

Figures & tables

Appendix figures & tables21 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking

    May 12, 2026Jisu Nam, Jahyeok Koo, Soowon Son +43D TrackingVideo Diffusion Models

  2. BEAST3D: Animal behavioral analysis and neural encoding from multi-view video via Gaussian splatting

    Jun 1, 2026Yanchen Wang, Lenny Aharon, Wangshu Zhu +73D Representation

  3. Fully Distributed Multi-View 3D Tracking in Real-Time

    Jun 11, 2026Byron Hernandez, Fangyu Li, Aotian Wu +3Collaborative Fusion