CAT-Free: Multi-View Pedestrian Localization without Calibration, Annotations, or Target-Scene Training via Adaptive Geometric Filtering
Organizations: Meijo University Department of Science Technology · Chubu Electric Power Co., Inc.
Abstract
Multi-camera pedestrian localization is useful for wide-area monitoring in public and commercial spaces. However, deploying these systems often requires considerable setup for each new environment. Existing methods typically require camera calibration, position annotations, or target-scene training. CAT-Free removes all three requirements. It uses synchronized RGB video as its only scene-specific input. Camera configuration is estimated directly from the video. Pedestrian locations are then estimated by combining observations from multiple cameras. Automatic camera estimation is not always accurate. This can produce unreliable pedestrian locations. CAT-Free therefore introduces two adaptive geometric filters. They remove unreliable position estimates. Their thresholds are estimated from each input sequence. CAT-Free achieves 82.5, 84.5, and 65.7 MODA on WildTrack, MultiviewX, and GMVD. It uses no supplied calibration, position annotations, or target-scene training. Published methods using such scene-specific information report 88.2--95.0 MODA on WildTrack and 83.9--96.5 on MultiviewX under their respective protocols. CAT-Free also transfers without retuning. It reaches 74.9 MODA on four additional sequences and 78.6 on an unseen 8-camera installation. Finally, localization uncertainty predicts MODA with . This provides a label-free estimate of localization reliability.
Figures & tables
| Localization | Tracking | |||||||
|---|---|---|---|---|---|---|---|---|
| Dataset | MODA | MODP | Precision | Recall | MOTA | MOTP | IDF1 | HOTA |
| WildTrack | 82.46 | 70.94 | 92.52 | 89.71 | 81.41 | 0.156 | 71.36 | 64.72 |
| MultiviewX | 84.54 | 77.04 | 96.74 | 87.48 | 79.05 | 0.138 | 64.96 | 59.65 |
| GMVD (s5/c1/seq1) | 65.69 | 62.12 | 96.56 | 68.11 | 64.00 | 0.212 | 55.23 | 46.44 |
| WildTrack, 360 unused frames | 58.27 | 65.49 | 88.96 | 66.52 | 59.46 | 0.215 | 52.31 | 46.52 |
| GMVD c1, seqs. 2–5, mean | 74.90 | 61.22 | 94.69 | 79.37 | 74.36 | 0.216 | 58.03 | 50.19 |
| Method | Supplied calibration | Target annotations | Target training | WildTrack MODA | MultiviewX MODA |
|---|---|---|---|---|---|
| UMPD [ 11 ] | Yes | No | Yes | 76.6 | 67.5 |
| DCHM [ 13 ] | Yes | No | Yes | 84.2 | 78.4 |
| CAT-Free (ours) | No | No | No | 82.5 | 84.5 |
| MVDet [ 7 ] | Yes | Yes | Yes | 88.2 | 83.9 |
| MVDeTr [ 8 ] | Yes | Yes | Yes | 91.5 | 93.7 |
| EarlyBird [ 16 ] | Yes | Yes | Yes | 91.2 | 94.2 |
| Variant | WT | MVX | GMVD | Mean | Mean |
| Q (camera support) | 80.99 | 82.06 | 59.93 | 74.33 | – |
| Ground only | 81.51 | 81.73 | 59.71 | 74.32 | -0.01 |
| Uncertainty, add | 76.89 | 82.93 | 64.58 | 74.80 | +0.47 |
| Uncertainty, replace | 78.47 | 83.07 | 64.68 | 75.41 | +1.08 |
| Ground uncertainty, replace | 80.46 | 83.94 | 63.14 | 75.85 | +1.52 |
| Ground uncertainty, add | 79.83 | 84.40 | 63.80 | 76.01 | +1.68 |
| Variant | WT | MVX | GMVD |
|---|---|---|---|
| (A) single-camera ground projection | 32.88 | 15.66 | 42.84 |
| (B) best two-ray estimate | 79.10 | 75.23 | 60.59 |
| CAT-Free | 82.46 | 84.54 | 65.69 |
| (C) CAT-Free with ground-truth calibration | 79.94 | 81.26 | 65.17 |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Stage | Fixed setting | Estimated per sequence |
|---|---|---|
| Detection | confidence | – |
| Rays | rad | pose/non-pose ratio ; median box height |
| Association | residual ; appearance threshold 0.6; uniqueness ratio 0.3 | – |
| Triangulation | Huber ; outlier | – |
| Scoring | – | signal weights and threshold (Otsu); reconstruction confidence gain |
| Selection | residual ; support radius ; NMS ; | , (log-Otsu) |
| Recovered rig (RGB only) | GT calibration (diagnostic) | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Frames | GT | People/frame | Cameras/person | MODA | Recall | Precision | MODA | Recall | Precision |
| 0–199 | 989 | 24.7 | 3.68 | 33.37 | 46.01 | 78.45 | 40.34 | 49.14 | 84.82 |
| 200–399 | 703 | 17.6 | 3.74 | 48.22 | 64.72 | 79.68 | 52.49 | 65.29 | 83.61 |
| 400–599 | 898 | 22.4 | 3.75 | 50.33 | 62.69 | 83.53 | 54.68 | 63.47 | 87.83 |
| 600–799 | 931 | 23.3 | 4.16 | 62.84 | 67.56 | 93.46 | 63.05 | 67.78 | 93.48 |
| 800–999 | 1,264 | 31.6 | 4.35 | 60.05 | 66.06 | 91.66 | 63.05 | 66.06 | 95.65 |
| Seq. | GT | MODA | MODP | Prec. | Rec. | MOTA | MOTP | IDF1 | HOTA | Out | MODA out-FP |
| 2 | 2,802 | 74.09 | 66.81 | 96.97 | 76.48 | 73.98 | 0.184 | 66.80 | 57.41 | 265 | 64.63 |
| 3 | 2,803 | 79.02 | 58.25 | 94.86 | 83.55 | 79.27 | 0.221 | 61.04 | 50.22 | 216 | 71.32 |
| 4 | 4,160 | 71.32 | 62.32 | 93.01 | 77.12 | 71.01 | 0.230 | 51.52 | 46.26 | 80 | 69.40 |
| 5 | 2,692 | 75.15 | 57.49 | 93.92 | 80.35 | 73.18 | 0.227 | 52.77 | 46.86 | 80 | 72.18 |
| Mean | – | – |
| Seq. | GT | MODA | MODP | Prec. | Rec. | MOTA | MOTP | IDF1 | HOTA | Out | MODA out-FP |
| 1 | 4,086 | 70.31 | 69.70 | 95.16 | 74.08 | 69.43 | 0.172 | 65.61 | 56.74 | 191 | 65.64 |
| 2 | 2,802 | 73.55 | 65.03 | 96.40 | 76.41 | 73.09 | 0.194 | 72.71 | 61.03 | 487 | 56.17 |
| 3 | 2,803 | 86.41 | 69.43 | 95.91 | 90.26 | 84.34 | 0.167 | 71.24 | 61.55 | 366 | 73.35 |
| 4 | 4,161 | 80.25 | 67.90 | 94.94 | 84.76 | 77.63 | 0.185 | 61.96 | 56.28 | 159 | 76.42 |
| 5 | 2,692 | 82.24 | 66.58 | 93.28 | 88.63 | 79.05 | 0.182 | 58.82 | 55.13 | 238 | 73.40 |
| Mean | – | – |
| Method | Given calibration | Target labels | Target training | Eval. GT align | WildTrack | MultiviewX |
|---|---|---|---|---|---|---|
| MVDet [ 7 ] | ✓ | ✓ | ✓ | – | 88.2 | 83.9 |
| MVDeTr [ 8 ] | ✓ | ✓ | ✓ | – | 91.5 | 93.7 |
| EarlyBird [ 16 ] | ✓ | ✓ | ✓ | – | 91.2 | 94.2 |
| TrackTacular [ 17 ] | ✓ | ✓ | ✓ | – | 93.2 | 96.5 |
| CaMuViD [ 4 ] | – | ✓ | ✓ | – | 95.0 | 96.5 |
| UMPD [ 11 ] | ✓ | – | ✓ | – | 76.6 | 67.5 |
| Cameras | Co-visibility | MODA | Recall | Precision |
|---|---|---|---|---|
| 2 | 1.80 | 5.9 | 29.0 | |
| 3 | 2.76 | 47.58 | 49.4 | 96.5 |
| 4 | 2.95 | 56.20 | 62.5 | 90.8 |
| 5 | 3.68 | 66.81 | 75.0 | 90.2 |
| 6 | 4.66 | 81.41 | 87.9 | 93.1 |
| 7 | 5.28 | 82.46 | 89.7 | 92.5 |
| Setting | Median (cm) | MODA |
|---|---|---|
| 2 cameras, development frames | 76.4 | |
| 3 cameras, development frames | 31.9 | 47.58 |
| 4 cameras, development frames | 31.2 | 56.20 |
| 5 cameras, development frames | 20.3 | 66.81 |
| 6 cameras, development frames | 18.2 | 81.41 |
| 7 cameras, development frames | 15.2 | 82.46 |
| Dataset | Cameras | All-camera MODA | LOO mean | LOO min | LOO SD | Held-out center error |
|---|---|---|---|---|---|---|
| WildTrack | 7 | 82.46 | 81.87 | 79.62 | 1.13 | 0.30 m |
| MultiviewX | 6 | 84.54 | 84.56 | 84.27 | 0.16 | 0.19 m |
| GMVD (s5/c1/seq1) | 6 | 65.69 | 64.64 | 62.48 | 1.14 | 0.40 m |
| Change | WildTrack | MultiviewX | GMVD | Mean | |
|---|---|---|---|---|---|
| I | earlier baseline (pose-confirmed detections only, unweighted rays) | 78.68 | 78.92 | 56.36 | 71.32 |
| K | all detections + ray reliability | 80.57 | 81.39 | 59.64 | 73.87 |
| Q | remove exclusive assignment after support recount | 80.99 | 82.06 | 59.93 | 74.33 |
| BF | weighted view support uncertainty filter | 78.47 | 83.07 | 64.68 | 75.41 |
| BK | + ground-plane filter after the uncertainty filter | 81.62 | 84.34 | 64.51 | 76.82 |
| BS | + inlier refit of the triangulated position | 82.25 | 84.54 | 64.63 | 77.14 |
| Variant | WildTrack | MultiviewX | GMVD | Mean |
|---|---|---|---|---|
| no trajectory stage (per-frame positions) | 79.94 | 81.12 | 62.01 | 74.36 |
| BS (all rules) | 82.25 | 84.54 | 64.63 | 77.14 |
| short static-track removal (BT) | 82.46 | 84.54 | 65.69 | 77.56 |
| gap interpolation | 82.25 | 84.54 | 64.63 | 77.14 |
| track extension | 80.46 | 78.92 | 59.03 | 72.80 |
| appearance in linking and extension | 81.72 | 84.20 | 64.46 | 76.79 |
| Threshold rule | WildTrack | MultiviewX | GMVD | Mean |
|---|---|---|---|---|
| fixed quantile p50 | 81.83 | 55.69 | 41.56 | 59.69 |
| fixed quantile p60 | 79.52 | 74.50 | 57.19 | 70.40 |
| fixed quantile p70 | 76.26 | 84.27 | 67.38 | 75.97 |
| log-Otsu (ours) | 82.46 | 84.54 | 65.69 | 77.56 |
| Variant | seq1 | seq2 | seq3 | seq4 | seq5 | Mean |
|---|---|---|---|---|---|---|
| weighted view support | 64.83 | 67.31 | 80.73 | 74.86 | 77.04 | 72.95 |
| uncertainty filter only | 67.79 | 64.85 | 83.59 | 77.12 | 73.37 | 73.34 |
| ground-plane filter first | 67.16 | 69.38 | 81.63 | 78.23 | 79.83 | 75.25 |
| uncertainty then ground-plane filter (ours) | 70.31 | 73.55 | 86.41 | 80.25 | 82.24 | 78.55 |
| Separation | Filter | – | – | – | – | – | |
|---|---|---|---|---|---|---|---|
| 0.5 m | uncertainty | 93.4 | 93.3 | 93.1 | 93.4 | 93.2 | 93.3 |
| both | 64.2 | 65.1 | 68.1 | 72.0 | 75.0 | 77.4 | |
| 1.0 m | uncertainty | 93.5 | 93.5 | 93.6 | 93.9 | 93.8 | 93.8 |
| both | 40.1 | 42.4 | 49.1 | 57.8 | 66.1 | 72.0 | |
| 2.0 m | uncertainty | 92.3 | 92.6 | 93.3 | 94.2 | 94.7 | 94.9 |
| both | 8.1 | 11.3 | 19.9 | 32.9 | 47.8 | 58.3 |
| WildTrack | MultiviewX | GMVD | |
| all triangulated hypotheses | 92.2 | 93.4 | 74.6 |
| + angular consistency filter | 91.5 | 93.4 | 74.3 |
| + uncertainty filter | 89.0 | 90.3 | 67.4 |
| + ground-plane filter | 87.2 | 89.1 | 65.1 |
| BT: unmatched GT–prediction pairs within 0.5–1 m | 12 | 11 | 46 |
| BT: MODA if all were fixed | 84.98 (+2.52) | 86.01 (+1.47) | 67.94 (+2.25) |