Native Association: Confidence-Aware Human Perception in the Wild with a Foundation VLM
Organizations: WSC Sports, Tel Aviv, Israel
Abstract
Extracting who is where, on which team, wearing which number from a broadcast frame is typically done by stitching a detector, an OCR engine, and classifiers together -- and the stitching step swaps identities under occlusion. We make association native instead: a 0.77B vision-language model (Florence-2) is fine-tuned to emit all per-person attributes as one grammar-constrained sequence, with each attribute generated inside its owner's block. Output is therefore schema-valid on every frame by construction, and no post-hoc binding step exists to attach a correctly read number to the wrong player: residual misassociation is pure perception error, rarer than zero-shot-prompted frontier APIs' (0.057 vs. 0.21-0.24). On a frozen multi-sport test set, this single pass reaches 0.95 detection F1 (APIs: 0.65-0.75). A single extra forward pass yields a per-field confidence that supports a reject option (jersey precision at half coverage) and routes a training-free zoom-and-re-read for small players. Surprisingly, once the grammar is learned, further parameter-efficient tuning yields no measurable gain under the adaptation configurations we test; the identical recipe on WIDER-Attribute reaches 93.1 mAP given-box, yields the first detection-coupled end-to-end results under its standard test protocol (84.5 mAP), and reproduces the same tuning result. In this regime, the gains live in the structure, not in added weights.
Figures & tables
| Ours | Ours | Gemini | Gemini | Specialist | Florence-2 | |
| Metric | (E1b) | +re-query | 2.5-pro | 3.1-pro | cascade † | cascade † |
| Detection F1 | 0.950 | 0.950 | 0.652 | 0.745 | GT † | GT † |
| Team acc. (exact) | 0.809 | 0.809 | 0.838 | 0.880 | 0.558 | 0.558 |
| Team purity (perm-inv) | 0.952 | 0.952 | 0.974 | 0.994 | 0.764 | 0.764 |
| Match coverage | 0.950 | 0.950 | 0.704 | 0.772 | GT † | GT † |
| Jersey F1 | 0.757 | 0.792 | 0.704 | 0.737 | 0.460 | 0.520 |
| Method | Metric | Score |
| mAP era (given-box) | ||
| R*CNN [ 9 ] | mAP | 80.5 |
| DHC [ 19 ] | mAP | 81.3 |
| SRN [ 33 ] | mAP | 86.2 |
| DIAA [ 23 ] | mAP | 86.4 |
| Da-HAR [ 29 ] | mAP | 87.3 |