Organizations: The University of Hong Kong, Hong Kong, China · Centre for Transformative Garment Production · Hong Kong University of Science and Technology, Hong Kong, China · Texas A&M University, College Station, TX, US
Quantifying collective fish behavior requires accurate trajectories, yet multi-view 3D tracking remains challenging due to frequent occlusions, visually similar individuals, and the long-standing scarcity of identity annotations. We present TrackFish3D, a geometry-driven self-supervised framework for dense multi-camera 3D tracking of schooling fish. Instead of relying on appearance-based re-identification or manually annotated identities, TrackFish3D turns calibrated multi-view geometry into supervision: triangulation and reprojection consistency provide pseudo-associations, while a geometric encoder and global association transformer learn all-to-all cross-view correspondence within each frame. To make these associations identity-aware, TrackFish3D introduces a self-supervised contrastive objective that separates co-visible individuals in the embedding space, together with a temporal predictor that preserves identities and bridges short occlusions across frames. The resulting model is trained once on unlabeled footage and applied directly to unseen test videos, requiring no cross-view identity labels, temporal annotations, 3D ground truth, appearance features, or test-time optimization. On our benchmark, TrackFish3D improves 3D Multi-Object Tracking Accuracy from 87.7% for the strongest baseline to 95.8%. On the 3D-ZeF zebrafish benchmark, it achieves 81.1% MOTA, compared with 77.4% for the best geometric baseline. TrackFish3D also generalizes beyond fish, achieving strong results on real-world bird tracking.
Figures & tables
No labels / GT required
Temp.
Spat.
No XV-ID
No T-ID
No 3D GT
No Vis.
SVMOT
✓
✗
✓
✗
✓
✗
Supervised MVA
✗
✓
✗
✓
✓
✗
Self-supervised MVA
✗
✓
✓
✓
✓
✗
MVMOT
✓
✓
✗
✗
✓
✗
TrackFish3D
✓
✓
✓
✓
✓
✓
Tab. 1: Method families for dense 3D tracking. Temp./Spat. : temporal/spatial association. The central columns indicate whether a family can operate without cross-view identity labels (No XV-ID), temporal identity labels (No T-ID), or 3D ground-truth trajectories (No 3D GT). No Vis. : no visual appearance features.
Fig. 2: Overview of the TrackFish3D framework. Synchronised multi-view video is processed by (A) a Geometric Encoder and (B) an Association Transformer that predicts cross-view pairwise scores, supervised by (C) Geometric Pseudo-Supervision from triangulation consistency through (D) a Self-Supervised Association Loss . Detections are fused via (E) Multi-View Grouping & Validation to produce 3D hypotheses, whose embeddings are separated by (F) a Self-Supervised Contrastive Loss . Across frames, (G) Temporal Linking matches identities and (H) a Self-Supervised Temporal Loss enforces cross-frame consistency, enabling stable 3D trajectory recovery at inference.
Figure 3Figure 4
Fig. 6: Qualitative 3D trajectory comparison on SynFish test scenes. Columns: ground truth, TrackFish3D, and representative baselines. Colour encodes fish identity; trajectory breaks indicate identity switches or fragmentations.
Method
MOTA ↑
MT% ↑
ML% ↓
IDSw ↓
Frag ↓
MTBF ↑
Hungarian + nearest
77.4%
90.0%
0.0%
6.0
73.0
42.5
Greedy + nearest
76.9%
90.0%
0.0%
10.0
76.0
38.7
3D-SORT
75.6%
90.0%
0.0%
13.5
65.5
43.4
Self-MVA [ 12 ]
50.8%
60.0%
0.0%
21.5
73.5
48.5
ASNet [ 7 ]
0.3%
0.0%
0.0%
72.5
175.0
10.6
SambaMOTR [ 54 ]
− 26.4%
0.0%
100.0%
0.5
1.0
6.5
Tab. 3: 3D tracking results on the 3D-ZeF test set, averaged over 2 sequences (zebra2: 5 fish, zebra4: 5 fish). Best result in each column (among methods with MOTA >0 ) is bolded .
Fig. 7: Qualitative 3D trajectory comparison on 3D-ZeF test sequences. TrackFish3D maintains identity consistency through real-world occlusions that fragment or swap baseline trajectories. ReST produces no trajectories because its cross-view association fails to match detections correctly.
Variant
MOTA ↑
MT% ↑
ML% ↓
IDSw ↓
Frag ↓
MTBF ↑
Ablation 1: loss components
TrackFish3D (full)
97.7%
98.3%
0.0%
0.5
3.1
283.1
w/o Ltemp
86.0%
85.2%
8.0%
2.0
5.2
243.1
w/o Lctr
36.8%
23.8%
36.4%
3.0
8.0
99.1
w/o both
36.9%
25.6%
34.2%
3.8
9.8
88.8
Ablation 2: visual features
Tab. 4: Ablation studies on SynFish. Top: removing contrastive ( Lctr ) and/or temporal ( Ltemp ) losses. Bottom: concatenating DINOv3 [ 56 ] visual features with geometric features.
Variant
MOTA ↑
MT% ↑
IDSw ↓
Frag ↓
MTBF ↑
TrackFish3D (SAM3)
95.8%
95.8%
0.0
6.5
266.3
MLP-only + all losses
78.4%
64.2%
2.0
16.3
172.9
SAM3 + Hungarian
83.5%
71.9%
7.5
12.8
183.8
Tab. 5: Controlled Association Transformer ablation on SynFish.
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Module
Parameters
%
Geometric Feature Encoder ϕ
9,216
1.4%
Association Transformer ψ
579,457
88.5%
Temporal Predictor ρ
66,432
10.1%
Total (trainable)
≈ 655K
100%
Appendix
Tab. 6: TrackFish3D parameter budget (geometry-only mode, no visual backbone).
Category
Hyperparameter
Value
Optimiser
Algorithm
AdamW
Learning rate
1×10−4
Weight decay
1×10−4
Gradient clipping (max norm)
5.0
Loss weights
λasso (association)
1.0
λctr (contrastive)
0.5
Appendix
Tab. 7: Complete TrackFish3D training hyperparameters.
Split
Scene
#Fish
#Frames
#Cams
Resolution
OcclLevel
%Overlap
SynFish (train)
train1
4
359
3
1920 × 1080
Low
0.2%
SynFish (train)
train2
6
349
3
1920 × 1080
Mild
10.2%
SynFish (train)
train3
8
362
3
1920 × 1080
Moderate
81.1%
SynFish (train)
train4
10
359
3
1920 × 1080
Moderate
75.1%
SynFish (train)
train5
14
359
3
1920 × 1080
Heavy
91.9%
SynFish (train)
train6
14
362
3
1920 × 1080
Mild
66.9%
Appendix
Tab. 8: Dataset summary and difficulty statistics. Occlusion level is derived from the OccScore, which is computed per scene by taking, for each (frame, camera) pair, the maximum pairwise 2D bounding-box IoU among all tracked objects, then averaging these maxima across all frames and camera views. %Overlap is the fraction of (frame, camera) instances with any overlap (IoU > 0.01).
Rig
Train scenes
Test scenes
0
—
test1, test2
1
train1
—
2
train2
—
3
train4
—
4
train5, train6
—
5
train3
test3
Appendix
Tab. 9: SynFish scene-to-rig mapping. Scenes sharing a rig use identical camera placements (same R and t for all 3 cameras). Rig 0 appears only in the test set.
Rig
Cam
R (row-major)
t
0
0
− 0.6129
0.7902
− 0.0000
7.48
18.22
43.67
0.1589
0.1232
− 0.9796
− 0.7740
− 0.6004
− 0.2011
1
− 0.9866
− 0.1633
0.0000
− 1.35
13.66
73.22
− 0.1450
0.8759
− 0.4602
0.0752
− 0.4540
− 0.8878
Appendix
Tab. 10: SynFish camera rig extrinsics. Each rig defines 3 cameras (0, 1, 2) with rotation R (shown row-by-row) and translation t . All cameras share the intrinsic Ksyn above.
Sequence
Cam
R (row-major)
t
zebra01 (train)
0
0.9999
− 0.0143
0.0054
− 14.68
− 16.34
31.24
0.0139
0.9974
0.0708
− 0.0064
− 0.0707
0.9975
1
0.9999
− 0.0077
− 0.0067
− 14.63
− 9.80
48.50
0.0079
0.1591
0.9872
− 0.0065
− 0.9872
0.1591
Appendix
Tab. 11: 3D-ZeF camera extrinsics per sequence. All sequences share the intrinsics K0 , K1 above.
MOTA (%) ↑
MT (%) ↑
ML (%) ↓
Method
test1
test2
test3
test4
test5
test6
Avg
test1
test2
test3
test4
test5
test6
Avg
test1
test2
test3
test4
test5
test6
Avg
TrackFish3D (SAM3)
99.6
99.9
96.2
86.2
94.8
98.0
95.8
100
100
100
81.2
93.8
100
95.8
0
0
0
0
0
0
0.0
TrackFish3D (YOLO26+SAM2)
100
100
97.7
86.3
96.7
99.1
96.6
100
100
100
81.2
93.8
100
95.8
0
0
0
0
0
0
0.0
SAM3+Hung.
100
100
93.8
55.2
75.1
77.1
83.5
100
100
100
31.2
43.8
56.2
71.9
0
0
0
0
0
0
0.0
SAM3+Greedy
100
100
93.8
55.2
75.1
77.1
83.5
100
100
100
31.2
43.8
56.2
71.9
0
0
0
0
0
0
0.0
YOLO26+SAM2+Hung.
99.7
99.8
95.9
53.7
69.0
75.0
82.2
100
100
100
25.0
43.8
50.0
69.8
0
0
0
0
0
0
0.0
Appendix
Tab. 12: Per-scene accuracy metrics on the SynFish test set (6 held-out scenes). Scene names correspond to Table 8 : test1 (4 fish), test2 (5 fish), test3 (8 fish), test4–test6 (16 fish each).
IDSw ↓
Frag ↓
MTBF ↑
Method
test1
test2
test3
test4
test5
test6
Avg
test1
test2
test3
test4
test5
test6
Avg
test1
test2
test3
test4
test5
test6
Avg
TrackFish3D (SAM3)
0
0
0
0
0
0
0.0
0
0
2
18
9
10
6.5
359.0
358.6
276.4
158.6
224.7
220.5
266.3
TrackFish3D (YOLO26+SAM2)
0
0
0
0
0
1
0.2
0
0
0
11
5
6
3.7
359.0
359.0
350.6
198.5
271.2
240.5
296.5
SAM3+Hung.
0
0
5
12
17
11
7.5
0
0
6
27
25
19
12.8
359.0
359.0
142.4
66.0
77.8
98.3
183.8
SAM3+Greedy
0
0
5
12
17
11
7.5
0
0
6
27
25
19
12.8
359.0
359.0
142.4
66.0
77.8
98.3
183.8
YOLO26+SAM2+Hung.
0
0
3
21
13
15
8.7
0
0
7
27
18
21
12.2
358.0
358.4
154.1
54.1
88.2
85.9
183.1
Appendix
Tab. 13: Per-scene identity stability metrics on the SynFish test set (6 held-out scenes). Same scene abbreviations as Table 12 .
MOTA% ↑
MT% ↑
ML% ↓
IDSw ↓
Frag ↓
MTBF ↑
Method
z2
z4
Avg
z2
z4
Avg
z2
z4
Avg
z2
z4
Avg
z2
z4
Avg
z2
z4
Avg
TrackFish3D
82.7
79.4
81.1
100
80
90.0
0
0
0.0
2
6
4.0
24
37
30.5
103.9
83.3
93.6
Hung.+NN
77.1
77.7
77.4
100
80
90.0
0
0
0.0
6
6
6.0
60
86
73.0
43.8
41.2
42.5
Greedy+NN
76.3
77.4
76.9
100
80
90.0
0
0
0.0
8
12
10.0
64
88
76.0
40.3
37.1
38.7
Self-MVA
84.7
16.8
50.8
100
20
60.0
0
0
0.0
2
41
21.5
33
114
73.5
81.2
15.8
48.5
ASNet
13.9
− 13.4
0.3
0
0
0.0
0
0
0.0
55
90
72.5
152
198
175.0
11.5
9.7
10.6
Appendix
Tab. 14: Per-scene 3D tracking results on the 3D-ZeF test set (zebra2: 5 fish, zebra4: 5 fish). ReST fails to produce any valid cross-camera associations on zebrafish sequences.
α
β
MOTA ↑
MT% ↑
ML% ↓
IDSw ↓
Frag ↓
MTBF ↑
1.0
0.0
95.8%
95.8%
0.0%
0.0
6.5
266.3
0.8
0.2
95.8%
95.8%
0.0%
0.0
6.5
266.3
0.6
0.4
95.8%
95.8%
0.0%
0.0
6.5
266.3
0.4
0.6
95.8%
95.8%
0.0%
0.0
6.5
266.3
0.2
0.8
95.8%
95.8%
0.0%
0.2
6.8
262.5
0.0
1.0
95.8%
95.8%
0.0%
0.3
7.5
257.3
Appendix
Tab. 15: Ablation of temporal linking weights (α,β) on the SynFish test set (6 scenes, averaged). The cost is robust for α≥0.4 ; removing the distance component ( α=0 ) degrades identity stability. The selected setting is bolded .
σR (°)
σt
MOTA ↑
MT% ↑
ML% ↓
IDSw ↓
Frag ↓
MTBF ↑
0
0
95.8%
95.8%
0.0%
0.0
6.5
266.3
0.05
0.05
94.1%
95.8%
0.0%
0.0
5.0
274.0
0.1
0.1
92.1%
93.8%
0.0%
0.0
9.2
246.6
0.2
0.2
81.3%
82.3%
1.0%
0.2
20.3
206.0
Appendix
Tab. 16: Robustness of TrackFish3D (SAM 3) to camera extrinsic perturbation on the SynFish test set (6 scenes, averaged). σR is the rotation noise standard deviation (degrees) and σt is the translation noise standard deviation (world units). The unperturbed baseline ( σ=0 ) is shown in the first row.
σ
MOTA ↑
MT% ↑
ML% ↓
IDSw ↓
Frag ↓
MTBF ↑
0
95.8%
95.8%
0.0%
0.0
6.5
266.3
0.05
90.3%
95.8%
0.0%
3.8
12.2
175.2
0.10
84.1%
94.8%
0.0%
7.0
15.3
126.0
0.20
72.3%
86.5%
0.0%
11.2
24.8
86.4
0.40
45.4%
40.6%
3.1%
15.3
63.5
37.1
Appendix
Tab. 17: Robustness of TrackFish3D (SAM 3) to detection noise on the SynFish test set (6 scenes, averaged). σ is the noise level controlling both the false-negative drop rate and the false-positive injection rate. The unperturbed baseline ( σ=0 ) is shown in the first row.
Metric
Value
Parameters
655 K
Training time
17h 36m
Inference (ms / frame)
124.8
Peak GPU memory (GB)
0.02
Appendix
Tab. 18: TrackFish3D runtime and memory profile. Training time is for the full 400-epoch schedule on SynFish (8 training scenes). Inference speed is measured per frame (all cameras), averaged over the SynFish test set.
Family
Methods
Temp
Spat
No Sup
No Vis
How 3D tracks are produced
Det. + geometric MVA
SAM3 / YOLO26+SAM2 + {Hung., Greedy}
–
–
✓
✓
Geometric MVA → triangulate → NN linking
Det. + geometric MVA + Kalman
SAM3 + 3D-SORT
–
–
✓
✓
Geometric MVA → triangulate → Kalman tracking
Det. + learned MVA
Self-MVA
–
✓
✓
–
Learned MVA → triangulate → NN linking
ASNet
–
✓
–
–
Learned MVA → triangulate → NN linking
2D MOT → 3D
SambaMOTR, MOTIP
✓
–
–
–
2D tracking per view → MVA at t=0→ triangulate
Multi-camera MOT
ReST, MCTR
✓
✓
–
–
Cross-view 2D tracking → triangulate
Appendix
Tab. 19: Summary of baseline families, capability profiles, and the fair 3D tracking protocol. Temp = learned temporal association; Spat = learned spatial (cross-view) association; No Sup = no identity supervision required; No Vis = no visual appearance features used. All methods receive the same inputs (multi-view detections and calibrated cameras) and use the same post-processing for fairness.
Figure 24
Method
MOTA ↑
Recall ↑
Prec. ↑
IDF1 ↑
IDSW ↓
Frag ↓
MT ↑
PT ↓
ML ↓
Greedy
86.0%
92.6%
95.2%
63.3%
60
36
8
2
0
Hungarian
91.5%
93.9%
97.6%
92.0%
3
6
10
0
0
ASNet [ 7 ]
85.6%
91.6%
95.3%
63.0%
46
40
9
1
0
Self-MVA [ 12 ]
68.7%
78.4%
90.7%
47.8%
50
61
6
4
0
ReST [ 13 ]
88.3%
91.5%
98.5%
58.0%
56
35
8
2
0
MCTR [ 43 ]
28.3%
58.2%
67.7%
35.7%
66
87
1
7
2
Appendix
Tab. 21: Cross-species evaluation on the held-out Pigeon10 sequence seq48 from 3D-POP. TrackFish3D is trained on four real-world pigeon sequences and evaluated on the fifth. Results show that the same geometry-driven formulation transfers beyond fish to another visually similar animal collective.
Dataset
Method
MOTA ↑
MT% ↑
ML% ↓
IDSw ↓
Frag ↓
MTBF ↑
SynFish
SAM3 + Hungarian
83.5%
71.9%
0.0%
7.5
12.8
183.8
SynFish
3D-SORT
87.7%
77.1%
0.0%
4.0
9.0
209.4
SynFish
TrackFish3D (SAM3)
95.8%
95.8%
0.0%
0.0
6.5
266.3
3D-ZeF
Hungarian
77.4%
90.0%
0.0%
6.0
73.0
42.5
3D-ZeF
3D-SORT
75.6%
90.0%
0.0%
13.5
65.5
43.4
3D-ZeF
TrackFish3D
81.1%
90.0%
0.0%
4.0
30.5
93.6
Appendix
Tab. 22: Comparison with the 3D-SORT kinematic baseline.
Dense 3D tracking from monocular video is fundamental to dynamic scene understanding. While recent 3D foundation models provide reliable per-frame geometry, recovering object motion in this geometry remains challenging and benefits from strong motion priors learned from real-world videos. Existing 3D trackers either follow iterative paradigms trained from scratch on synthetic data or fine-tune 3D reconstruction models learned from static multi-view images, both lacking real-world motion priors. Pre-trained video diffusion transformers (video DiTs) offer rich spatio-temporal priors from internet-scale videos, making them a promising foundation for 3D tracking. However, their frame-anchored formulation, which generates each frame's content, is fundamentally mismatched with reference-anchored dense 3D tracking, which must follow the same physical points from a reference frame across time. We present TrackCraft3R, the first method to repurpose a video DiT as a feed-forward dense 3D tracker. Given a monocular video and its frame-anchored reconstruction pointmap, TrackCraft3R predicts a reference-anchored tracking pointmap that follows every pixel of the first frame across time in a single forward pass, along with its visibility. We achieve this through two designs: (i) a dual-latent representation that uses per-frame geometry latents and reference-anchored track latents as dense queries, and (ii) temporal RoPE alignment, which specifies the target timestamp of each track latent. Together, these designs convert the per-frame generative paradigm of video DiTs into a reference-anchored tracking formulation with LoRA fine-tuning. TrackCraft3R achieves state-of-the-art performance on standard sparse and dense 3D tracking benchmarks, while running 1.3x faster and using 4.6x less peak memory than the strongest prior method. We further demonstrate robustness to large motions and long videos.
Multi-view video recordings are increasingly used to capture the 3D movements of animals in experimental settings, yet extracting rich 3D representations from these recordings remains challenging. Supervised pose estimation requires extensive manual annotation, while general-purpose 3D reconstruction models trained on generic scene datasets fail on the specialized imagery and sparse-view setting of laboratory experiments. We address these limitations with BEAST3D, a self-supervised pretraining framework that learns 3D visual representations from unlabeled, calibrated multi-view video. BEAST3D uses a vision transformer to predict 3D Gaussian splats that reconstruct held-out views through differentiable rendering, while simultaneously segmenting the animal from the background. BEAST3D reconstructs 3D structure with as few as four views by conditioning directly on known camera parameters--unlike general-purpose models, which must estimate camera geometry from dense overlapping viewpoints that are seldom available in lab settings. Through comprehensive evaluation across four species, we demonstrate that BEAST3D produces rich, viewpoint-invariant features that transfer effectively to three downstream tasks: novel view synthesis, which validates the quality of the learned 3D representations; multi-view pose estimation, which provides the sparse keypoint trajectories widely used in behavioral analysis; and neural encoding, which relates 3D behavioral features to simultaneously recorded neural activity. BEAST3D thus establishes a versatile framework for behavioral analysis that leverages 3D structure in modern multi-view laboratory recordings.
Yanchen Wang, Lenny Aharon, Wangshu Zhu +7
Columbia University · Cold Spring Harbor · Stanford University
Multi-camera tracking with overlapping fields of view typically relies on centralized fusion, which creates computational bottlenecks that prevent deployment at scale. We present MV3DT, a fully distributed framework for real-time multi-view 3D tracking that achieves accurate identity propagation and occlusion recovery through peer-to-peer coordination, eliminating the need for central aggregation. Each camera node executes a lightweight modular pipeline comprising monocular 3D perception, distributed multi-view association, and collaborative fusion via lightweight messaging. MV3DT achieves 96.5% IDF1, 93.1% MOTA, and 94.6% MOTP on WILDTRACK, competitive with state-of-the-art centralized methods, and unprecedented 41.7% IDF1 and 50.9% MOTA on SCOUT while demonstrating superior scalability: sustaining 30 FPS on 100 cameras with <10ms inter-camera latency and only 2.2% communication overhead. MV3DT operates in a zero-shot regime given camera calibrations, requiring no scene-specific learning and making it directly deployable in new environments. These results establish MV3DT as a practical solution for real-time multi-view tracking in large-scale overlapping camera networks.
Byron Hernandez, Fangyu Li, Aotian Wu +3
University of Florida, Gainesville, FL, USA · NVIDIA Corporation, Santa Clara, CA, USA