Multi-camera pedestrian localization is useful for wide-area monitoring in public and commercial spaces. However, deploying these systems often requires considerable setup for each new environment. Existing methods typically require camera calibration, position annotations, or target-scene training. CAT-Free removes all three requirements. It uses synchronized RGB video as its only scene-specific input. Camera configuration is estimated directly from the video. Pedestrian locations are then estimated by combining observations from multiple cameras. Automatic camera estimation is not always accurate. This can produce unreliable pedestrian locations. CAT-Free therefore introduces two adaptive geometric filters. They remove unreliable position estimates. Their thresholds are estimated from each input sequence. CAT-Free achieves 82.5, 84.5, and 65.7 MODA on WildTrack, MultiviewX, and GMVD. It uses no supplied calibration, position annotations, or target-scene training. Published methods using such scene-specific information report 88.2--95.0 MODA on WildTrack and 83.9--96.5 on MultiviewX under their respective protocols. CAT-Free also transfers without retuning. It reaches 74.9 MODA on four additional sequences and 78.6 on an unseen 8-camera installation. Finally, localization uncertainty predicts MODA with r=−0.98. This provides a label-free estimate of localization reliability.
Figures & tables
Figure 1 : RGB video in, tracked ground positions out, with no supplied calibration, no position annotations from the target scene, and no target-scene training. (b) shows the recovered ground, cameras, and every trajectory in the WildTrack test sequence. (c) relates published WildTrack MODA to each method’s requirements. Table 9 in the supplement lists the protocol differences.
Figure 2 : CAT-Free pipeline. ❶ Pretrained models with fixed weights extract information from each camera view: VGGT estimates the camera configuration and ground surface, YOLO11x with pose provides pedestrian detections and 3D rays with reliability wi , and OSNet provides appearance features. ❷ Detections of the same pedestrian are matched across cameras, and their rays are combined to estimate candidate 3D pedestrian positions. ❸ The proposed filters remove positions that are imprecise ( σ>τσ ) or inconsistent with the ground ( g>τg ), followed by non-maximum suppression (NMS). ❹ The remaining positions are linked over time to form pedestrian trajectories on the ground. Both thresholds are estimated from the input sequence, without supplied calibration, position annotations, or target-scene training. The lower row shows the four stages on WildTrack frame 1900. No ground truth is used.
Localization
Tracking
Dataset
MODA
MODP
Precision
Recall
MOTA
MOTP ↓
IDF1
HOTA
WildTrack
82.46
70.94
92.52
89.71
81.41
0.156
71.36
64.72
MultiviewX
84.54
77.04
96.74
87.48
79.05
0.138
64.96
59.65
GMVD (s5/c1/seq1)
65.69
62.12
96.56
68.11
64.00
0.212
55.23
46.44
WildTrack, 360 unused frames
58.27
65.49
88.96
66.52
59.46
0.215
52.31
46.52
GMVD c1, seqs. 2–5, mean
74.90
61.22
94.69
79.37
74.36
0.216
58.03
50.19
Table 1 : Main results. The upper block contains the three development sequences. The lower block reports post-development transfer. The bold column is the primary localization metric. MODA and MODP use a 0.5 m matching threshold; MOTA, MOTP, and IDF1 use 1 m. HOTA is computed with TrackEval.
Method
Supplied calibration
Target annotations
Target training
WildTrack MODA
MultiviewX MODA
UMPD [ 11 ]
Yes
No
Yes
76.6
67.5
DCHM [ 13 ]
Yes
No
Yes
84.2
78.4
CAT-Free (ours)
No
No
No
82.5
84.5
MVDet [ 7 ]
Yes
Yes
Yes
88.2
83.9
MVDeTr [ 8 ]
Yes
Yes
Yes
91.5
93.7
EarlyBird [ 16 ]
Yes
Yes
Yes
91.2
94.2
Table 2 : Published MODA on WildTrack and MultiviewX together with the per-installation resources used by each reported result. “Supplied calibration” indicates that dataset camera calibration is given to the method, “Target annotations” indicates that annotations from the target environment are used for learning, and “Target training” indicates that model fitting or fine-tuning is performed on images from that environment. Published accuracy values follow the respective papers and may use different training and evaluation protocols; the table is therefore intended to contextualize accuracy against deployment requirements rather than provide a strict controlled ranking.
Variant
WT
MVX
GMVD
Mean
Δ Mean
Q (camera support)
80.99
82.06
59.93
74.33
–
Ground only
81.51
81.73
59.71
74.32
-0.01
Uncertainty, add
76.89
82.93
64.58
74.80
+0.47
Uncertainty, replace
78.47
83.07
64.68
75.41
+1.08
Ground → uncertainty, replace
80.46
83.94
63.14
75.85
+1.52
Ground → uncertainty, add
79.83
84.40
63.80
76.01
+1.68
Table 3 : Adaptive-filter ablation on the development sequences (MODA). Δ Mean is relative to Q. The best ordering applies localization uncertainty before ground-plane consistency.
Variant
WT
MVX
GMVD
(A) single-camera ground projection
32.88
15.66
42.84
(B) best two-ray estimate
79.10
75.23
60.59
CAT-Free
82.46
84.54
65.69
(C) CAT-Free with ground-truth calibration
79.94
81.26
65.17
Table 4 : Controlled localization baselines on the development sequences (MODA). WT denotes WildTrack and MVX denotes MultiviewX. All rows use the same pedestrian detections, estimated ground, and trajectory stage.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Stage
Fixed setting
Estimated per sequence
Detection
confidence ≥0.30
–
Rays
σθ=0.0206 rad
pose/non-pose ratio (spose/sother)2 ; median box height h~
Association
residual ≤2σθ ; appearance threshold 0.6; uniqueness ratio 0.3
–
Triangulation
Huber δ=0.01E ; outlier >3δ
–
Scoring
–
signal weights and threshold (Otsu); reconstruction confidence gain
Selection
residual ≤2σθ ; support radius σθ ; NMS r=0.018E ;
τσ , τg (log-Otsu)
Appendix
Table 5 : Downstream settings after detection and camera estimation. Fixed values are shared across datasets. E is the reconstructed scene extent, and σθ is the angular noise standard deviation.
Figure 3 : Ray reliability wi on WildTrack frame 1900. Each box is a detection and each dot is the bottom-center point that defines its back-projection ray. Colour is wi from Eq. ( 3 ): distant people project to short boxes, so their foot point carries a larger angular error and receives a smaller weight. Triangulation and view support use these weights; no ground truth is involved.
Figure 4 : The two geometric filters. (a) Localization uncertainty σ increases with viewing distance and poor intersection angles. Ellipses show the positional covariance Λ−1 of Eq. ( 6 ). (b) The point-to-plane distance g exposes an incorrect correspondence that ground-constrained triangulation would hide.
Recovered rig (RGB only)
GT calibration (diagnostic)
Frames
GT
People/frame
Cameras/person
MODA
Recall
Precision
MODA
Recall
Precision
0–199
989
24.7
3.68
33.37
46.01
78.45
40.34
49.14
84.82
200–399
703
17.6
3.74
48.22
64.72
79.68
52.49
65.29
83.61
400–599
898
22.4
3.75
50.33
62.69
83.53
54.68
63.47
87.83
600–799
931
23.3
4.16
62.84
67.56
93.46
63.05
67.78
93.48
800–999
1,264
31.6
4.35
60.05
66.06
91.66
63.05
66.06
95.65
Appendix
Table 6 : Evaluation on 360 WildTrack frames not used during development, grouped into 200-frame intervals. Camera parameters and dense reconstruction come from frames 1800–1995. Thresholds are estimated again on the 360 evaluation frames. The right block replaces the estimated camera parameters with ground-truth calibration as a diagnostic.
Seq.
GT
MODA
MODP
Prec.
Rec.
MOTA
MOTP ↓
IDF1
HOTA
Out
MODA out-FP
2
2,802
74.09
66.81
96.97
76.48
73.98
0.184
66.80
57.41
265
64.63
3
2,803
79.02
58.25
94.86
83.55
79.27
0.221
61.04
50.22
216
71.32
4
4,160
71.32
62.32
93.01
77.12
71.01
0.230
51.52
46.26
80
69.40
5
2,692
75.15
57.49
93.92
80.35
73.18
0.227
52.77
46.86
80
72.18
Mean
–
74.90±3.19
61.22±4.29
94.69±1.70
79.37±3.26
74.36±3.51
0.216±0.021
58.03±7.21
50.19±5.20
–
69.38±3.37
Appendix
Table 7 : Post-development evaluation on four additional GMVD scene-5/configuration-1 sequences. The method and sequence-1 camera parameters are fixed. MODA out-FP counts predictions outside the evaluation grid as false positives.
Seq.
GT
MODA
MODP
Prec.
Rec.
MOTA
MOTP ↓
IDF1
HOTA
Out
MODA out-FP
1
4,086
70.31
69.70
95.16
74.08
69.43
0.172
65.61
56.74
191
65.64
2
2,802
73.55
65.03
96.40
76.41
73.09
0.194
72.71
61.03
487
56.17
3
2,803
86.41
69.43
95.91
90.26
84.34
0.167
71.24
61.55
366
73.35
4
4,161
80.25
67.90
94.94
84.76
77.63
0.185
61.96
56.28
159
76.42
5
2,692
82.24
66.58
93.28
88.63
79.05
0.182
58.82
55.13
238
73.40
Mean
–
78.55±6.54
67.73±1.96
95.14±1.19
82.83±7.25
76.71±5.71
0.180±0.011
66.07±5.92
58.15±2.95
–
69.00±8.25
Appendix
Table 8 : Post-development evaluation on an unseen 8-camera installation, GMVD scene 5, configuration 2. Camera parameters are estimated from RGB with the fixed 25-iteration procedure. Iteration 20 is selected by triangulation residual. MODA out-FP counts predictions outside the evaluation grid as false positives.
Figure 5 : Qualitative results on the middle evaluation frame of each development dataset. Top: camera 1 and person detections. Bottom: bird’s-eye view after evaluation-only similarity alignment. Predictions and people are matched at the 0.5 m MODA threshold.
Method
Given calibration
Target labels
Target training
Eval. GT align
WildTrack
MultiviewX
MVDet [ 7 ]
✓
✓
✓
–
88.2
83.9
MVDeTr [ 8 ]
✓
✓
✓
–
91.5
93.7
EarlyBird [ 16 ]
✓
✓
✓
–
91.2
94.2
TrackTacular [ 17 ]
✓
✓
✓
–
93.2
96.5
CaMuViD [ 4 ]
–
✓
✓
–
95.0
96.5
UMPD [ 11 ]
✓
–
✓
–
76.6
67.5
Appendix
Table 9 : Resource and protocol context using reported MODA. This is not a controlled ranking because each method follows its authors’ protocol. CAT-Free uses development-set method selection, full-sequence processing, and evaluation-time similarity alignment.
Figure 6 : Localization uncertainty σ and point-to-plane distance g after angular consistency filtering, normalized by the NMS radius. Log-Otsu estimates τσ from all hypotheses and τg from hypotheses with σ≤τσ . The shaded region is accepted. Ground truth is used only to color points.
Cameras
Co-visibility
MODA
Recall
Precision
2
1.80
−8.51
5.9
29.0
3
2.76
47.58
49.4
96.5
4
2.95
56.20
62.5
90.8
5
3.68
66.81
75.0
90.2
6
4.66
81.41
87.9
93.1
7
5.28
82.46
89.7
92.5
Appendix
Table 10 : Camera-removal experiment on the WildTrack development frames. Frames, detections, camera parameters, ground plane, and fixed constants do not change. The first k of seven cameras provide the rays. Co-visibility is the mean number of those cameras that observe each annotated person.
Setting
Median σ (cm)
MODA
2 cameras, development frames
76.4
−8.51
3 cameras, development frames
31.9
47.58
4 cameras, development frames
31.2
56.20
5 cameras, development frames
20.3
66.81
6 cameras, development frames
18.2
81.41
7 cameras, development frames
15.2
82.46
Appendix
Table 11 : One curve, two ways of varying difficulty. Median σ from Eq. ( 6 ) at the annotated positions, against measured MODA, pooling the camera-removal sweep with two time windows evaluated at the full seven cameras. Over these eight settings logσ predicts MODA with Pearson r=−0.98 (Spearman −0.93 , p<0.001 ).
Dataset
Cameras
All-camera MODA
LOO mean
LOO min
LOO SD
Held-out center error
WildTrack
7
82.46
81.87
79.62
1.13
0.30 m
MultiviewX
6
84.54
84.56
84.27
0.16
0.19 m
GMVD (s5/c1/seq1)
6
65.69
64.64
62.48
1.14
0.40 m
Appendix
Table 12 : Leave-one-camera-out evaluation alignment. The similarity transform is refitted after withholding one camera center, and all predictions are rescored. The last column measures the distance between the withheld camera’s estimated and annotated centers after alignment.
+ ground-plane filter after the uncertainty filter
81.62
84.34
64.51
76.82
BS
+ inlier refit of the triangulated position
82.25
84.54
64.63
77.14
Appendix
Table 13 : Development-set MODA after cumulative changes. Method I is the initial controlled baseline. Variants were selected on these sequences.
Figure 7 : (a) Development MODA after each adopted change. (b) Outcome of every annotated person under the final pipeline. Outcomes are a true positive, a prediction 0.5–1 m away, removal by tracking, rejection by selection, or no triangulated hypothesis. FP is the number of false positives.
Variant
WildTrack
MultiviewX
GMVD
Mean
no trajectory stage (per-frame positions)
79.94
81.12
62.01
74.36
BS (all rules)
82.25
84.54
64.63
77.14
− short static-track removal (BT)
82.46
84.54
65.69
77.56
− gap interpolation
82.25
84.54
64.63
77.14
− track extension
80.46
78.92
59.03
72.80
− appearance in linking and extension
81.72
84.20
64.46
76.79
Appendix
Table 14 : Trajectory ablation on the peaks of BS (MODA). Each row removes one rule from BS.
Threshold rule
WildTrack
MultiviewX
GMVD
Mean
fixed quantile p50
81.83
55.69
41.56
59.69
fixed quantile p60
79.52
74.50
57.19
70.40
fixed quantile p70
76.26
84.27
67.38
75.97
log-Otsu (ours)
82.46
84.54
65.69
77.56
Appendix
Table 15 : Per-sequence versus fixed thresholds (MODA). Both filter thresholds are replaced by the same fixed quantile on every dataset. The best fixed quantile differs by dataset, while log-Otsu is within 0.6 points of the best fixed choice on WildTrack and MultiviewX.
Variant
seq1
seq2
seq3
seq4
seq5
Mean
weighted view support
64.83
67.31
80.73
74.86
77.04
72.95
uncertainty filter only
67.79
64.85
83.59
77.12
73.37
73.34
ground-plane filter first
67.16
69.38
81.63
78.23
79.83
75.25
uncertainty then ground-plane filter (ours)
70.31
73.55
86.41
80.25
82.24
78.55
Appendix
Table 16 : Filter ablation on the unseen 8-camera installation (GMVD configuration 2, MODA). The variants match Table 3 . The order selected on the development sequences performs best on every sequence.
Separation
Filter
ϕ<10∘
10 – 30∘
30 – 45∘
45 – 60∘
60 – 75∘
75 – 90∘
0.5 m
uncertainty
93.4
93.3
93.1
93.4
93.2
93.3
both
64.2
65.1
68.1
72.0
75.0
77.4
1.0 m
uncertainty
93.5
93.5
93.6
93.9
93.8
93.8
both
40.1
42.4
49.1
57.8
66.1
72.0
2.0 m
uncertainty
92.3
92.6
93.3
94.2
94.7
94.9
both
8.1
11.3
19.9
32.9
47.8
58.3
Appendix
Table 17 : Fraction of incorrect two-camera pairs that survive each filter, grouped by the angle ϕ between the directions joining the two people and the two cameras. Lower is better.
WildTrack
MultiviewX
GMVD
all triangulated hypotheses
92.2
93.4
74.6
+ angular consistency filter
91.5
93.4
74.3
+ uncertainty filter
89.0
90.3
67.4
+ ground-plane filter
87.2
89.1
65.1
BT: unmatched GT–prediction pairs within 0.5–1 m
12
11
46
BT: MODA if all were fixed
84.98 (+2.52)
86.01 (+1.47)
67.94 (+2.25)
Appendix
Table 18 : GT coverage (fraction of annotated people within 0.5 m of some output) after each stage of BK, and the MODA upper bound from fixing near misses of BT.
Reducing the number of cameras reduces the deployment cost but removes views that correct BEV responses stretched away from true pedestrian positions by projection and short score drops that can split tracks} in Bird's-Eye View (BEV) tracking. We introduce GRACE, a camera-efficient multi-view tracker with three components. Volumetric-Guided Fusion combines homography-based BEV features with features lifted through 3D space. Ray Conditioning exposes each camera's viewing direction to the fusion network. Its tracking component, BEV Track Recovery (BTR), uses low-confidence detections only to continue existing tracks. The same detections cannot start new tracks. With two WildTrack cameras, GRACE improves MOTA from 83.54 for TrackTacular, our baseline, to 91.07.
Taigo Sakai, Kazuhiro Hotta, Hiroki Kouno +1
Meijo university · Chubu Electric Power Co., Inc. · Department of Science Technology 1-1 Higashishin-cho, Higashi-ku, 1-501, Shiogamaguchi, Nagoya 461-8680, Japan Tempaku, Nagoya 468-8501, Japan
Visual localization -- estimating a camera pose within a pre-existing map -- is a fundamental problem in computer vision. Floorplans are an attractive map representation: they are readily available for most buildings, compact, and inherently invariant to visual appearance changes. However, bridging the severe domain gap between camera observations and floorplan geometry remains challenging. Existing methods address this gap through data-driven learning, yet they require large-scale training data and environment-specific retraining, limiting their practical deployment. We propose a zero-shot floorplan localization method that generalizes to novel environments without any retraining. Our key insight is that dominant geometric primitives -- lines and circles -- are ubiquitous in human-made environments and provide appearance-invariant structural constraints. We extract these primitives from a bird's-eye-view (BEV) projection of monocular 3D reconstructions and match them to the floorplan via dedicated minimal solvers within a robust estimation framework. Experiments on both simulated and real-world datasets show that our approach outperforms state-of-the-art learning-based methods on unseen environments, while using a single fixed set of hyperparameters across all experiments. The source code will be made publicly available.
Ayumi Umemura, Toshinori Kuwahara, Marc Pollefeys +1
Multi-Camera People Tracking (MCPT) traditionally relies on precise intrinsic and extrinsic camera calibration to project 2D detections into a unified 3D world coordinate system.However, manual calibration constitutes a major bottleneck in large-scale dataset generation from unconstrained video archives. This work proposes a unified calibration-free 3D MCPT framework that infers geometric structure directly from visual data using deep foundation models. The system integrates anchor-free detection (YOLOX), robust tracking (BoT-SORT), omni-scale appearance embedding (OsNet), pose estimation (HRNet via MMPose), and transformer-based geometric reconstruction using the Visual Geometry Grounded Transformer (VGGT). A pose-guided 3D lifting strategy projects head keypoints onto a reconstructed manifold, eliminating dependence on ground-plane homography. Global identity association is formulated as hierarchical agglomerative clustering under a joint appearance-geometry cost with strict velocity gating. Evaluation on the AI City Challenge 2024 demonstrates a HOTA score of 53.13% without access to ground-truth calibration matrices, establishing a strong baseline for purely vision-based 3D tracking.
Ponleur Veng, Dominique Vaufreydaz, Phutphalla Kong
CADT, M-PSI · Cambodia Academy of Digital Technology · Univ. Grenoble Alpes, CNRS, Grenoble INP, LIG, 38000 Grenoble, France +2