Refereeing decisions in professional soccer remain inconsistent because referees cannot easily compare a contentious foul against similar past cases. We cast this as a retrieval problem and introduce SoccerNet-FoulRet, the first benchmark for semantic foul retrieval. Given a query foul, the task is to retrieve past fouls judged to be relevant precedents, regardless of camera angle, teams, or appearance. This differs from prior video-to-video retrieval, which matches clips by visual similarity or a shared event. Here, relevance is defined by refereeing interpretation. We build the benchmark from the SoccerNet-MVFoul dataset and evaluate retrieval ability of zero-shot video and vision-language embedders together with a task-specific fine-tuned baseline on 693 human-verified queries and category-relevance labels. Semantic foul retrieval remains challenging. The strongest zero-shot model achieves under 5% HitRate@10 on human-verified precedents, while category-supervised fine-tuning improves category relevance but transfers only modestly to precedent retrieval. We release SoccerNet-FoulRet to establish semantic foul retrieval as an open problem: https://github.com/SoccerNet/sn-foulret.
Figures & tables
Figure 1 : A memory bank for referees. We aim to augment referees rather than replace them. Automated classifiers and generative explainers emit a verdict directly, but can do so overconfidently and even hallucinate. Instead, our system takes a new foul (left) and retrieves semantically similar precedents together with the decisions that past referees reached on them (right). The referee reasons from these precedents and keeps the final call. We introduce a human-verified benchmark for this video-to-video retrieval task and find that off-the-shelf zero-shot video and vision-language retrievers fall well short, leaving it an open problem.
Dataset
#Videos
#Queries
Query Origin
Domain
Task
CC_WEB_VIDEO [ 23 ]
13K
24
Real
Web (YouTube)
NDVR
UQ_VIDEO [ 20 ]
170K
24
Real
Web (YouTube)
NDVR
SVD [ 9 ]
562K
1,206
Real
Web (Douyin)
NDVR
VCDB [ 10 ]
100K
528
Real
Web (YouTube)
Copy Detection
TRECVID-CBCD [ 13 ]
11,503
11,256
Synthetic
Broadcast TV
Copy Detection
FIVR-200K [ 12 ]
226K
100
Real
Web (YouTube)
Fine-grained IR
Table 1: Comparison of video-to-video retrieval datasets. Our benchmark is the first to target semantic similarity in Soccer.
Figure 2 : Two fouls with identical mechanics but vastly different sanctions depending on pitch location and context (penalty area vs. midfield).
Figure 3 : Manual Validation Interface. The annotator reviews a query action (left) alongside a candidate precedent (right), labeling each pair as Valid or Rejected .
Action class
Gal.
Qry.
Prec.
Offence / Severity
Gal.
Qry.
Prec.
Standing tackling
1,264
291
1,183
Offence
2,495
613
2,490
Tackling
448
124
532
No offence
324
53
162
Challenge
383
89
302
Between
96
27
99
Holding
361
93
373
Elbowing
178
39
155
Sev. 1
1,402
355
1,399
High leg
103
24
104
Sev. 2
403
91
367
Table 2: SoccerNet-FoulRet composition. Gallery clips (Gal.), query fouls (Qry.), and approved precedent pairs (Prec.) across the three attribute types. All are long-tailed: standing tackling is ∼44% of the gallery, 86% of clips are an offence , and severities 4 – 5 are rare ( <3% ). Counts are over each attribute’s labelled subset (full gallery 2,916 ).
HitRate@ K
Recall@ K
nDCG@ K
Method
Params
@1
@5
@10
@1
@5
@10
@5
@10
MRR10
Random
—
0.14
0.68
1.35
0.03
0.17
0.34
0.16
0.24
0.40
Text-mediated retrieval
VARS+XVARS
—
0.14 [-1pt] [0.00,0.43]
2.16 [-1pt] [1.15,3.32]
3.17 [-1pt] [1.88,4.47]
0.03 [-1pt] [0.00,0.09]
0.64 [-1pt] [0.32,1.00]
0.94 [-1pt] [0.56,1.38]
0.51 [-1pt] [0.26,0.81]
0.65 [-1pt] [0.38,0.96]
0.99 [-1pt] [0.54,1.51]
Video-native retrieval
InternVideo2-1B
1B
0.72 [-1pt] [0.14,1.44]
2.60 [-1pt] [1.44,3.90]
4.62 [-1pt] [3.03,6.20]
0.15 [-1pt] [0.03,0.30]
0.66 [-1pt] [0.36,1.00]
1.13 [-1pt] [0.74,1.55]
0.62 [-1pt] [0.34,0.94]
0.87 [-1pt] [0.56,1.21]
1.65 [-1pt] [0.96,2.44]
Table 3: Retrieval performance on the human-verified precedent pool, pooled over the 693 queries that have at least one verified precedent. HitRate@ K is the fraction of queries with at least one verified precedent in the top K ; Recall@ K is the fraction of verified precedents retrieved; nDCG@ K measures the discounted ranking of the available verified precedents using binary relevance; and MRR@10 is the mean reciprocal rank of the first verified precedent within the top 10. All values are percentages. Because the relevance pool is incomplete, unjudged retrieved clips may also be valid precedents and receive no relevance gain under these metrics. Qwen3-VL-Emb-2B-FT is fine-tuned exclusively on SoccerNet-MVFoul’s train set. All other visual embedders are evaluated zero-shot.
Action
+Offence
+Severity ∗
Model
Params
@1
@5
@10
@1
@5
@10
@1
@5
@10
Random
—
24.9
16.1
13.5
20.8
13.3
11.0
11.0
6.2
4.8
Text-mediated retrieval
VARS+XVARS †
—
38.5 [-1pt] [35.0,42.2]
29.4 [-1pt] [27.1,31.7]
29.0 [-1pt] [26.8,31.2]
32.7 [-1pt] [29.2,36.2]
24.5 [-1pt] [22.4,26.7]
24.0 [-1pt] [21.9,26.2]
20.6 [-1pt] [17.5,23.8]
13.5 [-1pt] [11.8,15.2]
10.7 [-1pt] [9.5,11.9]
Video-native retrieval
InternVideo2-1B
1B
37.5 [-1pt] [34.0,41.2]
26.3 [-1pt] [24.3,28.4]
23.1 [-1pt] [21.5,24.8]
29.9 [-1pt] [26.7,33.3]
20.9 [-1pt] [19.0,22.8]
18.3 [-1pt] [16.8,20.0]
17.1 [-1pt] [14.3,20.0]
9.5 [-1pt] [8.4,10.8]
7.4 [-1pt] [6.5,8.2]
Table 4: Category relevance, scored with the official SoccerNet-MVFoul attribute labels. A retrieved clip is relevant if it matches the query’s foul at one of three increasingly strict levels: Action, Action+Offence, or Action+Offence+Severity verdict. For each level, both the retrieval gallery and evaluated queries are restricted to clips carrying the required labels. Action and +Offence use 2,905 gallery clips and 709 queries; +Severity uses 2,561 gallery clips and 650 queries. Values are mean Average Precision ( mAP@ K , %). Qwen3-VL-Emb-2B-FT is fine-tuned exclusively on SoccerNet-MVFoul’s train set, whereas the other visual embedders are evaluated zero-shot. VARS+XVARS † uses VARS-predicted categories for candidate filtering and is therefore excluded from the best-model comparison.
Gold Precedents
Action
Model
Instruction
MRR@10
mAP@1
mAP@5
mAP@10
Qwen3-VL-Emb-2B
Generic
1.52
37.7
27.0
22.9
Domain
1.15
39.1
26.6
22.1
Detailed
0.82
36.1
25.8
21.7
Qwen3-VL-Emb-8B
Generic
1.08
35.3
23.7
20.5
Domain
0.66
30.3
19.7
16.6
Table 5: Instruction-sensitivity ablation. Query-side instruction specificity for the four instruction-following embedders, from Generic to Detailed . We report referee-gold MRR@10 and category mAP at the Action level (%); higher is better. Bold = best instruction per model and column. InternVideo2 is video-native and takes no instruction.
Figure 4 : A successful retrieval: a human-verified precedent is retrieved despite differences in camera zoom.
Figure 5 : An illustrative failure: VARS misclassifies an elbowing as a challenge , removing relevant precedents before semantic ranking.
Vision-language models (VLMs) have recently shown strong potential in soccer video understanding. However, given the high complexity of soccer videos due to large viewpoint variations, rapid shot transitions, and cluttered scenes, it remains unclear on whether VLMs rely on meaningful visual evidence or exploit spurious correlations and shortcut learning. Existing evaluation protocols focus primarily on classification accuracy and do not assess visual grounding. To address this limitation, we introduce SoccerLens, a benchmark for grounded soccer video understanding. The benchmark contains annotated video segments spanning 13 common soccer events, with structured visual cues organized into three levels of semantic relevance. We further extend the attribution method of Chefer [arXiv:2103.15679] to jointly model spatial and temporal attention, and introduce evaluation metrics that measure whether model attention aligns with annotated cues or drifts toward spurious regions. Our evaluation of state-of-the-art soccer VLMs shows that, despite strong classification accuracy, current models fail to exceed 50% grounding performance even under the loosest cue definitions and consistently underutilize temporal information. These results reveal a substantial gap between predictive performance and true visual grounding, highlighting the need for grounded evaluation in complex spatio-temporal domains such as soccer.
Ismael Elsharkawi, Ahmed Sait, Silvio Giancola +3
Department of Computer Science and Engineering, The American University in Cairo · Image And Visual Understanding Lab (IVUL), KAUST
While Multimodal Large Language Models (MLLMs) excel at generic video understanding, their ability to support specialized, rule-grounded decision-making remains insufficiently explored. In this paper, we introduce RefereeBench, the first large-scale benchmark for evaluating MLLMs as automatic sports referees. Spanning 11 sports with 925 curated videos and 6,475 QA pairs, RefereeBench evaluates five core officiating abilities: foul existence, foul and penalty classification, foul and penalty reasoning, entity perception, and temporal grounding. The benchmark is fully human-annotated to ensure high-quality annotations grounded in authentic officiating logic and multimodal evidence. Extensive evaluations of state-of-the-art MLLMs show that even the strongest models, such as Doubao-Seed-1.8 and Gemini-3-Pro, achieve only around 60% accuracy, while the strongest open-source model, Qwen3-VL, reaches only 47%. These results indicate that current models remain far from being reliable sports referees. Further analysis shows that while models can often identify incidents and involved entities, they struggle with rule application and temporal grounding, and frequently over-call fouls on normal clips. Our benchmark highlights the need for future MLLMs that better integrate domain knowledge and multimodal understanding, advancing trustworthy AI-assisted officiating and broader multimodal decision-making.
Yichen Xu, Yuanhang Liu, Chuhan Wang +5
Renmin University of China Beijing, China · Sichuan University Chengdu, China
Refereeing is vital in sports, where fair, accurate, and explainable decisions are fundamental. While intelligent assistant technologies are being widely adopted in soccer refereeing, current AI-assisted approaches remain preliminary. Existing research mostly focuses on isolated video perception tasks and lacks the ability to understand and reason about foul scenarios. To fill this gap, we propose SoccerRef-Agents, a holistic and explainable multi-agent decision-making framework for soccer refereeing. The main contributions are: (i) constructing the multimodal benchmark SoccerRefBench with over 1,200 referee theory questions and 600 foul video clips; (ii) building a vector-based knowledge base RefKnowledgeDB using the latest "Laws of the Game" and a classic case database for precise, knowledge-driven reasoning; (iii) designing a novel multi-agent architecture that collaborates via cross-modal RAG to bridge the semantic gap between visual content and regulatory texts. This work explores the technical capability of integrating MLLMs with refereeing expertise, and evaluations show our system significantly outperforms general-purpose MLLMs in decision accuracy and explanation quality. All databases, benchmarks, and code will be made available.
Zi Meng, Wanli Song, Yi Hu +2
University of Michigan, Michigan, USA · Shanghai Jiao Tong University, Shanghai, China