MOTIP2: Spatial Priors for End-to-End Multi-Object Tracking
Organizations: École Centrale de Lyon LIRIS, UMR CNRS 5205 Lyon, France · IDEMIA Courbevoie, France
Abstract
End-to-end multi-object trackers have narrowed the gap with classical tracking-by-detection on association-difficult benchmarks. Yet they still make spatially implausible errors no classical tracker would, such as assigning one identity to objects on opposite sides of the frame. A model could learn to avoid them, but tracking annotations are scarce, so we encode spatial priors explicitly instead, while keeping inference fully end-to-end with no post-hoc association. We propose three spatial priors, at the data, loss, and representation stages. Spatial ID Switches bias trajectory permutations toward spatially overlapping objects, reducing the mismatch between training and inference confusions. Spatial ID Loss scales each identity's penalty by its box distance, so a distant switch costs more than a nearby one. Spatial Anchor gives each track token its frame position, an explicit spatial cue for attention. We instantiate the three priors in MOTIP2, a tracker adapted from MOTIP and built on the real-time DEIM detection transformer. Trained without extra data, its main model, MOTIP2-L, sets a new state of the art: 73.4 HOTA on DanceTrack, 76.0 on SportsMOT, and 71.1 IDF1 on PersonPath22. MOTIP2 is a family of models spanning the speed-accuracy trade-off: a lighter model, MOTIP2-S, matches the original MOTIP at over 3x the speed, and MOTIP2-X reaches 74.8 HOTA on DanceTrack.
Figures & tables
| Method | HOTA | DetA | AssA | MOTA | IDF1 |
|---|---|---|---|---|---|
| Tracking-by-detection | |||||
| FairMOT [ Zhang et al.(2021b)Zhang, Wang, Wang, Zeng, and Liu ] | 39.7 | 66.7 | 23.8 | 82.2 | 40.8 |
| CenterTrack [ Zhou et al.(2020)Zhou, Koltun, and Krähenbühl ] | 41.8 | 78.1 | 22.6 | 86.8 | 35.7 |
| ByteTrack [ Zhang et al.(2022)Zhang, Sun, Jiang, Yu, Weng, Yuan, Luo, Liu, and Wang ] | 47.7 | 71.0 | 32.1 | 89.6 | 53.9 |
| OC-SORT [ Cao et al.(2023)Cao, Pang, Weng, Khirodkar, and Kitani ] | 55.1 | 80.3 | 38.3 | 92.0 | 54.6 |
| StrongSORT [ Du et al.(2023)Du, Zhao, Song, Zhao, Su, Gong, and Meng ] | 55.6 | 80.7 | 38.6 | 91.1 | 55.2 |
| Method | HOTA | DetA | AssA | MOTA | IDF1 |
|---|---|---|---|---|---|
| Tracking-by-detection | |||||
| FairMOT [ Zhang et al.(2021b)Zhang, Wang, Wang, Zeng, and Liu ] | 49.3 | 70.2 | 34.7 | 86.4 | 53.5 |
| GTR [ Zhou et al.(2022)Zhou, Yin, Koltun, and Krähenbühl ] | 54.5 | 64.8 | 45.9 | 67.9 | 55.8 |
| QDTrack [ Pang et al.(2021)Pang, Qiu, Li, Chen, Li, Darrell, and Yu ] | 60.4 | 77.5 | 47.2 | 90.1 | 62.3 |
| CenterTrack [ Zhou et al.(2020)Zhou, Koltun, and Krähenbühl ] | 62.7 | 82.1 | 48.0 | 90.8 | 60.0 |
| ByteTrack [ Zhang et al.(2022)Zhang, Sun, Jiang, Yu, Weng, Yuan, Luo, Liu, and Wang ] | 62.8 | 77.1 | 51.2 | 94.1 | 69.8 |
| Method | IDF1 | MOTA | IDSW | FP | FN |
| Tracking-by-detection | |||||
| CenterTrack [ Zhou et al.(2020)Zhou, Koltun, and Krähenbühl ] | 46.4 | 59.3 | 10 319 | 24K | 72K |
| FairMOT [ Zhang et al.(2021b)Zhang, Wang, Wang, Zeng, and Liu ] | 61.1 | 61.8 | 5 095 | 15K | 80K |
| ByteTrack [ Zhang et al.(2022)Zhang, Sun, Jiang, Yu, Weng, Yuan, Luo, Liu, and Wang ] | 66.8 | 75.4 | 5 931 | 17K | 40K |
| OC-SORT [ Cao et al.(2023)Cao, Pang, Weng, Khirodkar, and Kitani ] | 68.2 | 74.0 | 4 931 | 17K | 41K |
| Path Consistency [ Lu et al.(2024)Lu, Shuai, Chen, Xu, and Modolo ] | 69.4 | 75.8 | 4 663 | 15K | 38K |
| Accuracy | V100 PT FP32 | T4 TRT FP16 | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Method | HOTA | DetA | AssA | Det. | ID Dec | Total | Det. | ID Dec | Total |
| MOTIP [ Gao et al.(2025)Gao, Qi, and Wang ] | 64.0 | 75.4 | 54.7 | 72.7 | 14.2 | 86.9 | – | ||
| MOTIP2-X | 70.1 | 80.7 | 61.1 | 71.2 | 9.9 | 81.1 | 38.8 | 1.3 | 40.1 |
| MOTIP2-L | 68.0 | 77.3 | 60.0 | 45.0 | 9.8 | 54.8 | 21.9 | 1.3 | 23.2 |
| MOTIP2-M | 67.3 | 76.4 | 59.7 | 32.5 | 6.8 | 39.3 | 15.8 | 0.9 | 16.7 |
| MOTIP2-S | 64.5 | 74.0 | 56.5 | 21.8 | 5.3 | 27.1 | 9.4 | 0.7 | 10.1 |
| PersonPath22 | DanceTrack | |||||||
| HOTA | DetA | AssA | IDF1 | HOTA | DetA | AssA | IDF1 | |
| Full model | 61.6 | 63.8 | 60.3 | 70.4 | 64.5 | 74.0 | 56.5 | 68.8 |
| uniform IDSW | 56.1 | 58.6 | 54.4 | 64.0 | 63.7 | 73.0 | 55.8 | 68.1 |
| w/o IDSW | 59.5 | 64.2 | 55.8 | 66.8 | 60.1 | 73.3 | 49.7 | 63.0 |
| w/o Spatial ID Loss | 61.4 | 64.2 | 59.4 | 70.1 | 63.5 | 73.3 | 55.5 | 67.3 |
| w/o Spatial Anchor | 61.6 | 64.3 | 59.9 | 70.4 | 64.0 | 73.3 | 56.1 | 68.6 |
| Method | HOTA | DetA | AssA | MOTA | IDF1 | |
|---|---|---|---|---|---|---|
| Rel. PE | – | 61.3 | 64.1 | 59.3 | 73.4 | 70.2 |
| RoPE | 1.0 | 60.4 | 63.7 | 58.0 | 72.7 | 68.7 |
| 1.5 | 61.0 | 63.7 | 59.1 | 72.8 | 69.6 | |
| 2.0 | 61.6 | 63.8 | 60.3 | 72.9 | 70.4 | |
| 2.5 | 61.6 | 63.8 | 60.3 | 72.9 | 70.4 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | HOTA | DetA | AssA | MOTA | IDF1 |
|---|---|---|---|---|---|
| Tracking-by-detection | |||||
| CenterTrack [ Zhou et al.(2020)Zhou, Koltun, and Krähenbühl ] | 52.2 | 53.8 | 51.0 | 67.8 | 64.7 |
| FairMOT [ Zhang et al.(2021b)Zhang, Wang, Wang, Zeng, and Liu ] | 59.3 | 60.9 | 58.0 | 73.7 | 72.3 |
| DeepSORT [ Wojke et al.(2017)Wojke, Bewley, and Paulus ] | 61.2 | 63.1 | 59.7 | 78.0 | 74.5 |
| SORT [ Bewley et al.(2016)Bewley, Ge, Ott, Ramos, and Upcroft ] | 63.0 | 64.2 | 62.2 | 80.1 | 78.2 |
| ByteTrack [ Zhang et al.(2022)Zhang, Sun, Jiang, Yu, Weng, Yuan, Luo, Liu, and Wang ] | 63.1 | 64.5 | 62.0 | 80.3 | 77.3 |
| Training | Inference | ||||
| Sampling | Swap% | LR% | HOTA | LR-sw | |
| DanceTrack | |||||
| None | – | – | – | 60.1 | 120 |
| Uniform | 37% | 85% | 63.7 | 260 | |
| Ours | 10% | 26% | 64.5 | 134 | |
| PersonPath22 | |||||
| PersonPath22 | DanceTrack | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| HOTA | DetA | AssA | MOTA | IDF1 | IDSW | HOTA | DetA | AssA | MOTA | IDF1 | IDSW | |
| 1 | 61.5 | 63.7 | 60.0 | 72.5 | 70.3 | 5 612 | 64.3 | 73.7 | 56.4 | 81.9 | 67.7 | 3 287 |
| 2 | 61.6 | 63.8 | 60.3 | 72.9 | 70.4 | 5 069 | 64.5 | 74.0 | 56.5 | 82.9 | 68.8 | 2 489 |
| 3 | 61.6 | 63.8 | 60.3 | 73.1 | 70.4 | 4 911 | 64.5 | 74.2 | 56.3 | 83.2 | 68.3 | 2 356 |
| Method | HOTA | Latency (ms) | Resolution |
|---|---|---|---|
| End-to-end | |||
| FastTrackTr-R18-Dec3 [ Liao et al.(2026)Liao, Yang, Wu, Yu, Li, and Zhang ] | 47.4 | 17.1 † | 640 640 |
| FastTrackTr-Dec3 [ Liao et al.(2026)Liao, Yang, Wu, Yu, Li, and Zhang ] | 51.2 | 27.2 † | 640 640 |
| PuTR [ Liu et al.(2025a)Liu, Li, Wang, and Xu ] | 54.5 | 39.9 † | 800 1333 |
| FastTrackTr (640) [ Liao et al.(2026)Liao, Yang, Wu, Yu, Li, and Zhang ] | 54.8 | 31.5 † | 640 640 |
| FastTrackTr (1333) [ Liao et al.(2026)Liao, Yang, Wu, Yu, Li, and Zhang ] | 56.9 | 53.6 † | 800 1333 |