Mobile robots operating among pedestrians need trajectories that become available quickly, remain spatially credible through missed observations, preserve identity, and fit within an embedded computing budget. Aggregate tracking scores provide limited insight into when and how trajectories fail, while varying detector inputs can confound tracker and detector quality. We present a deployment-focused, tracker-only evaluation protocol that uses shared detections to isolate tracker behavior and directly evaluates initialization, detector-gap continuation, identity recovery, close-neighbor association, and load-dependent tracker-step runtime, while Higher Order Tracking Accuracy (HOTA) is retained as a complementary aggregate measure. We apply the protocol to the JackRabbot Dataset and Benchmark (JRDB) using six open-source trackers and our lightweight Pedestrian Reference Tracker (PedRefTrack), together with a GT-assisted variant that estimates the remaining tracker-side gap under idealized association and motion. Under fixed detections, the non-GT trackers span only 24.26%-29.67% HOTA yet exhibit markedly different capability profiles. After 1.0 s without detector support, no tracker without GT assistance maintains spatially correct, same-identity output in more than half of eligible cases, making missing-observation continuation the dominant limitation among the tested properties. Close-neighbor failures are smaller and increase mainly at the shortest separations. Tracker-step runtime on an NVIDIA Jetson Orin is heavy-tailed and load-sensitive, causing several trackers to fall below the 10 Hz real-time target in crowded frames. The protocol provides a reproducible way to characterize tracker behavior and deployment suitability in pedestrian-centric environments. Code and evaluation scripts are released at https://github.com/SCAI-Lab/tracker_eval.
Figures & tables
Fig. 1 : Trajectory partition for deployment-focused tests. An example pedestrian trajectory is divided into three complementary regions. Green marks the first 0.5 s after detector support begins, blue marks mostly continuous supported motion containing only brief detection gaps, and orange marks longer gaps bounded by supported detections. Together, these regions span the trajectory and define separate event sets for the tests in Sec. III .
Fig. 2 : GT-derived detector-input models. Dropout removes contiguous observations. Instability introduces perturbations to box center, orientation, and dimensions. Combined inputs apply matched dropout and instability severities.
Parameter
Mild
Moderate
Severe
Dropout start pdrop (%/frame)
0.5
1.0
2.0
Dropout duration (s)
0.05 – 0.50
0.05 – 1.00
0.05 – 2.00
Box-error hypotheses K
3
4
5
Hypothesis switch pswitch (%/frame)
5
10
20
XY σxybias/σxyjit (m)
0.03/0.01
0.06/0.02
0.10/0.03
Yaw σψbias/σψjit (rad)
0.50/0.10
0.80/0.15
1.10/0.20
TABLE I : Combined GT-derived perturbation severities. Each track draws K persistent box-error hypotheses and switches the active one with probability pswitch . Bias terms σbias define hypothesis-specific fixed offsets from GT; jitter terms σjit define independent zero-mean noise added each frame. L/W/H values are relative.
Fig. 3 : Pedestrian-centric setting and evaluation coverage. (a) Example bird’s-eye-view JRDB scene. (b) Counts supporting the detector-gap, nearest-neighbor, and runtime analyses.
Fig. 4 : Accuracy, runtime, and tracker-side headroom. (a) HOTA versus median tracker-step FPS. Whiskers show the 10th–90th percentile ranges of sequence HOTA (vertical) and frame-level FPS (horizontal). (b) Per-sequence HOTA deficit to GT-assisted PedRefTrack, with quantiles from paired same-sequence differences. GT-assisted PedRefTrack idealizes association and motion while retaining the remaining tracker pipeline.
Method
HOTA ↑ (%)
Detection (%)
Association (%)
LocA ↑ (%)
Runtime
DetA ↑
DetRe ↑
DetPr ↑
AssA ↑
AssRe ↑
AssPr ↑
Med. FPS ↑
P10–P90 FPS ↑
CPU cores used ↓
GT-assisted PedRefTrack (ours) †
33.33
29.34
36.16
43.20
38.62
42.54
57.51
66.26
62.26
[17.47,150.13]
1
PedRefTrack (ours, no GT)
29.67
27.18
33.79
41.28
33.13
37.80
51.43
65.74
74.96
[18.71,172.39]
1
Fast-Poly [ 24 ]
27.48
25.61
32.23
39.83
30.27
33.85
52.52
66.12
29.15
[7.77,70.55]
3
AB3DMOT [ 19 ]
26.83
26.84
33.65
40.99
27.72
31.05
52.93
65.95
9.59
[0.82,45.20]
7
SimpleTrack [ 21 ]
26.53
25.42
34.34
36.52
28.52
32.10
52.13
65.49
2.23
[0.32,9.54]
7
TABLE II : JRDB test results with shared detections. † marks the GT-assisted variant. FPS measures tracker-step throughput on Jetson Orin with brackets showing the 10th–90th quantiles. CPU cores reports the number used out of 12 available. Bold and underlining mark the best value including and excluding the GT-assisted variant, respectively. No tracker uses a GPU. Results use our corrected JRDB 3D-IoU evaluator and are not directly comparable to the official JRDB leaderboard.
Fig. 5 : Trajectory-level success tests on fixed detections. (a) Same-ID, spatially accepted output through detector gaps. (b) Pre-gap identity recovery when detector support returns, regardless of gap output. (c) Same-ID continuity failure over 1.0 s windows versus 10th-percentile GT nearest-neighbor distance. (d) Cumulative stable initialization within 0.5 s, requiring a subsequent second of uninterrupted correct output. Shading shows 95% bootstrap intervals from sequence resampling.
Fig. 6 : Tracker-step FPS versus input detections. Whiskers show the 10th–90th quantiles, and the dashed line marks the 10 Hz LiDAR rate. Detector inference is excluded.
Fig. 7 : GT-derived detector perturbations. The upper strip shows GT-assisted PedRefTrack HOTA. The main panel shows each tracker’s HOTA deficit to this reference. Mild, moderate, and severe correspond to Table I .
Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences but miss the physical artifacts that matter most, particularly unstable support and incorrect contacts such as foot skating and mistimed touch-downs. Meanwhile, widely used test suites are small and lack the diversity needed to stress contact-rich, long-horizon behaviors. We introduce HumanTracker to make humanoid tracking evaluation both perceptually aligned and scalable. The HumanTracker benchmark contains approximately 153 hours of optical motion trajectories from multiple professional performers, organized into four motion families with text labels for fine-grained diagnosis. We further propose HumanScore, a preference-aligned metric trained on 12K motion pairs containing 24K motions. Across representative state-of-the-art trackers, HumanScore better predicts human preferences and reveals contact and stability failures that kinematic metrics often miss.
Dairu Liu, Zekun Qi, Jiayu Zeng +11
1Nankai University · 3Galbot · 2Tsinghua University +3
In real-world applications, pedestrian trajectory prediction models rely on inputs from detection and tracking systems. Prior trajectory prediction benchmarks either contain relatively sparse pedestrian interactions, assume perfect tracking inputs, or rely on overhead viewpoints that minimize occlusion and perspective distortion, limiting evaluation in realistic dense-crowd scenarios. We present CrowdTraj, a benchmark for pedestrian trajectory prediction in natural dense crowd scenes. Unlike previous datasets, CrowdTraj supports end-to-end evaluation from detection through tracking to trajectory prediction under severe occlusion in CCTV views. It also captures diverse, natural pedestrian behaviours, including abrupt directional changes rarely observed in existing benchmarks. CrowdTraj includes five diverse scenes, with an average of 1,146 unique pedestrians per scene, maximum frame-level densities ranging from 114 to 372 pedestrians, and over 3.2 million annotated head bounding boxes. CrowdTraj provides pixel and real-world coordinates via per-scene homography matrices for physically meaningful analysis. Our experimental results show that tracking accuracy (IDF1) drops to 0.68 to 0.70 in the densest scenes, compared with approximately 0.90 in less crowded scenes. Trajectory prediction training also becomes substantially more computationally expensive in dense scenes, with training times increasing by up to 8 times. These findings show that CrowdTraj exposes limitations in current trajectory prediction pipelines that remain hidden on existing sparse-crowd benchmarks, particularly in robustness to tracking noise and computational scalability.
Antonius Bima Murti Wijaya, Paul Henderson, Marwa Mahmoud
School of Computing Science University of Glasgow University Avenue G12 8QQ, Glasgow, United Kingdom
Reducing the number of cameras reduces the deployment cost but removes views that correct BEV responses stretched away from true pedestrian positions by projection and short score drops that can split tracks} in Bird's-Eye View (BEV) tracking. We introduce GRACE, a camera-efficient multi-view tracker with three components. Volumetric-Guided Fusion combines homography-based BEV features with features lifted through 3D space. Ray Conditioning exposes each camera's viewing direction to the fusion network. Its tracking component, BEV Track Recovery (BTR), uses low-confidence detections only to continue existing tracks. The same detections cannot start new tracks. With two WildTrack cameras, GRACE improves MOTA from 83.54 for TrackTacular, our baseline, to 91.07.
Taigo Sakai, Kazuhiro Hotta, Hiroki Kouno +1
Meijo university · Chubu Electric Power Co., Inc. · Department of Science Technology 1-1 Higashishin-cho, Higashi-ku, 1-501, Shiogamaguchi, Nagoya 461-8680, Japan Tempaku, Nagoya 468-8501, Japan