cs.CVSep 28, 2026

VastMAT: A Large-Scale Multi-Category Benchmark for Multi-Animal Tracking

Authors: Zhizhen Li, Zan Wang, Huidong Peng, Bohan Tan, Shimin Shan, Yu Liu, Liang Peng

Organizations: Dalian University of Technology · Wuhan University · University of North Texas · The Hong Kong University of Science and Technology

Abstract

Multi-animal tracking (MAT) supports the study of animal movement, behavior, and group interactions. However, general multi-object tracking (MOT) benchmarks primarily focus on pedestrians and vehicles, whereas dedicated MAT benchmarks remain limited in jointly supporting broad animal coverage, large-scale video data, and extensive within-video multi-instance association. To address this gap, we introduce VastMAT, which has four key characteristics: (1) Large scale. It comprises 2,947 videos with 1,002,562 annotated frames, totaling 27.85 hours. (2) Broad category coverage. These videos cover 337 animal categories with diverse morphologies and motion patterns. (3) Extensive instance annotations. It provides 3,663,248 bounding boxes and 22,883 identity trajectories---to our knowledge, the largest numbers of both among dedicated MAT benchmarks. (4) High-quality annotations. To ensure reliability, annotations undergo iterative expert review and correction, and quality is assessed through an independent reannotation audit. To systematically assess tracking performance and cross-category generalization, we establish Seen-category and category-disjoint Unseen-category protocols, and evaluate eight representative MOT methods under both protocols. Under these protocols, the highest baseline HOTA scores are 66.37% and 52.90%, respectively, highlighting the challenge of tracking unseen animals. To address the low-overlap association challenge revealed by our analysis, we propose Center-Distance-Augmented Association (CDA), a lightweight module that adaptively combines IoU with center similarity normalized by the boxes' own scales. Without additional training, CDA improves TrackTrack's HOTA by 1.58 and 1.31 percentage points under the two protocols, respectively. To facilitate further MAT research, we will publicly release our benchmark and code.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. HieDG: A Hierarchical Discrete Geometry-Guided Framework for Multi-Animal Tracking

    Jul 1, 2026Chenxun Deng, Zhongde Zhang, Ye Yuan +7Multi-Object TrackingSpherical Latent Space

  2. COVTrack++: Learning Open-Vocabulary Multi-Object Tracking from Continuous Videos via a Synergistic Paradigm

    Mar 25, 2026Zekun Qian, Wei Feng, Ruize Han +1Multi-Object TrackingOpen-Vocabulary Object Detection

  3. Promptable Animal Pose Tracking Across Species

    Aug 5, 2026Le Li, Daniela Ivanova, Nicolas PugeaultSpeciesKeypoint Detection