Refereeing decisions in professional soccer remain inconsistent because referees cannot easily compare a contentious foul against similar past cases. We cast this as a retrieval problem and introduce SoccerNet-FoulRet, the first benchmark for semantic foul retrieval. Given a query foul, the task is to retrieve past fouls judged to be relevant precedents, regardless of camera angle, teams, or appearance. This differs from prior video-to-video retrieval, which matches clips by visual similarity or a shared event. Here, relevance is defined by refereeing interpretation. We build the benchmark from the SoccerNet-MVFoul dataset and evaluate retrieval ability of zero-shot video and vision-language embedders together with a task-specific fine-tuned baseline on 693 human-verified queries and category-relevance labels. Semantic foul retrieval remains challenging. The strongest zero-shot model achieves under 5% HitRate@10 on human-verified precedents, while category-supervised fine-tuning improves category relevance but transfers only modestly to precedent retrieval. We release SoccerNet-FoulRet to establish semantic foul retrieval as an open problem: https://github.com/SoccerNet/sn-foulret.
Figures & tables
Figure 1 : A memory bank for referees. We aim to augment referees rather than replace them. Automated classifiers and generative explainers emit a verdict directly, but can do so overconfidently and even hallucinate. Instead, our system takes a new foul (left) and retrieves semantically similar precedents together with the decisions that past referees reached on them (right). The referee reasons from these precedents and keeps the final call. We introduce a human-verified benchmark for this video-to-video retrieval task and find that off-the-shelf zero-shot video and vision-language retrievers fall well short, leaving it an open problem.
Dataset
#Videos
#Queries
Query Origin
Domain
Task
CC_WEB_VIDEO [ 23 ]
13K
24
Real
Web (YouTube)
NDVR
UQ_VIDEO [ 20 ]
170K
24
Real
Web (YouTube)
NDVR
SVD [ 9 ]
562K
1,206
Real
Web (Douyin)
NDVR
VCDB [ 10 ]
100K
528
Real
Web (YouTube)
Copy Detection
TRECVID-CBCD [ 13 ]
11,503
11,256
Synthetic
Broadcast TV
Copy Detection
FIVR-200K [ 12 ]
226K
100
Real
Web (YouTube)
Fine-grained IR
Table 1: Comparison of video-to-video retrieval datasets. Our benchmark is the first to target semantic similarity in Soccer.
Figure 2 : Two fouls with identical mechanics but vastly different sanctions depending on pitch location and context (penalty area vs. midfield).
Figure 3 : Manual Validation Interface. The annotator reviews a query action (left) alongside a candidate precedent (right), labeling each pair as Valid or Rejected .
Action class
Gal.
Qry.
Prec.
Offence / Severity
Gal.
Qry.
Prec.
Standing tackling
1,264
291
1,183
Offence
2,495
613
2,490
Tackling
448
124
532
No offence
324
53
162
Challenge
383
89
302
Between
96
27
99
Holding
361
93
373
Elbowing
178
39
155
Sev. 1
1,402
355
1,399
High leg
103
24
104
Sev. 2
403
91
367
Table 2: SoccerNet-FoulRet composition. Gallery clips (Gal.), query fouls (Qry.), and approved precedent pairs (Prec.) across the three attribute types. All are long-tailed: standing tackling is ∼44% of the gallery, 86% of clips are an offence , and severities 4 – 5 are rare ( <3% ). Counts are over each attribute’s labelled subset (full gallery 2,916 ).
HitRate@ K
Recall@ K
nDCG@ K
Method
Params
@1
@5
@10
@1
@5
@10
@5
@10
MRR10
Random
—
0.14
0.68
1.35
0.03
0.17
0.34
0.16
0.24
0.40
Text-mediated retrieval
VARS+XVARS
—
0.14 [-1pt] [0.00,0.43]
2.16 [-1pt] [1.15,3.32]
3.17 [-1pt] [1.88,4.47]
0.03 [-1pt] [0.00,0.09]
0.64 [-1pt] [0.32,1.00]
0.94 [-1pt] [0.56,1.38]
0.51 [-1pt] [0.26,0.81]
0.65 [-1pt] [0.38,0.96]
0.99 [-1pt] [0.54,1.51]
Video-native retrieval
InternVideo2-1B
1B
0.72 [-1pt] [0.14,1.44]
2.60 [-1pt] [1.44,3.90]
4.62 [-1pt] [3.03,6.20]
0.15 [-1pt] [0.03,0.30]
0.66 [-1pt] [0.36,1.00]
1.13 [-1pt] [0.74,1.55]
0.62 [-1pt] [0.34,0.94]
0.87 [-1pt] [0.56,1.21]
1.65 [-1pt] [0.96,2.44]
Table 3: Retrieval performance on the human-verified precedent pool, pooled over the 693 queries that have at least one verified precedent. HitRate@ K is the fraction of queries with at least one verified precedent in the top K ; Recall@ K is the fraction of verified precedents retrieved; nDCG@ K measures the discounted ranking of the available verified precedents using binary relevance; and MRR@10 is the mean reciprocal rank of the first verified precedent within the top 10. All values are percentages. Because the relevance pool is incomplete, unjudged retrieved clips may also be valid precedents and receive no relevance gain under these metrics. Qwen3-VL-Emb-2B-FT is fine-tuned exclusively on SoccerNet-MVFoul’s train set. All other visual embedders are evaluated zero-shot.
Action
+Offence
+Severity ∗
Model
Params
@1
@5
@10
@1
@5
@10
@1
@5
@10
Random
—
24.9
16.1
13.5
20.8
13.3
11.0
11.0
6.2
4.8
Text-mediated retrieval
VARS+XVARS †
—
38.5 [-1pt] [35.0,42.2]
29.4 [-1pt] [27.1,31.7]
29.0 [-1pt] [26.8,31.2]
32.7 [-1pt] [29.2,36.2]
24.5 [-1pt] [22.4,26.7]
24.0 [-1pt] [21.9,26.2]
20.6 [-1pt] [17.5,23.8]
13.5 [-1pt] [11.8,15.2]
10.7 [-1pt] [9.5,11.9]
Video-native retrieval
InternVideo2-1B
1B
37.5 [-1pt] [34.0,41.2]
26.3 [-1pt] [24.3,28.4]
23.1 [-1pt] [21.5,24.8]
29.9 [-1pt] [26.7,33.3]
20.9 [-1pt] [19.0,22.8]
18.3 [-1pt] [16.8,20.0]
17.1 [-1pt] [14.3,20.0]
9.5 [-1pt] [8.4,10.8]
7.4 [-1pt] [6.5,8.2]
Table 4: Category relevance, scored with the official SoccerNet-MVFoul attribute labels. A retrieved clip is relevant if it matches the query’s foul at one of three increasingly strict levels: Action, Action+Offence, or Action+Offence+Severity verdict. For each level, both the retrieval gallery and evaluated queries are restricted to clips carrying the required labels. Action and +Offence use 2,905 gallery clips and 709 queries; +Severity uses 2,561 gallery clips and 650 queries. Values are mean Average Precision ( mAP@ K , %). Qwen3-VL-Emb-2B-FT is fine-tuned exclusively on SoccerNet-MVFoul’s train set, whereas the other visual embedders are evaluated zero-shot. VARS+XVARS † uses VARS-predicted categories for candidate filtering and is therefore excluded from the best-model comparison.
Gold Precedents
Action
Model
Instruction
MRR@10
mAP@1
mAP@5
mAP@10
Qwen3-VL-Emb-2B
Generic
1.52
37.7
27.0
22.9
Domain
1.15
39.1
26.6
22.1
Detailed
0.82
36.1
25.8
21.7
Qwen3-VL-Emb-8B
Generic
1.08
35.3
23.7
20.5
Domain
0.66
30.3
19.7
16.6
Table 5: Instruction-sensitivity ablation. Query-side instruction specificity for the four instruction-following embedders, from Generic to Detailed . We report referee-gold MRR@10 and category mAP at the Action level (%); higher is better. Bold = best instruction per model and column. InternVideo2 is video-native and takes no instruction.
Figure 4 : A successful retrieval: a human-verified precedent is retrieved despite differences in camera zoom.
Figure 5 : An illustrative failure: VARS misclassifies an elbowing as a challenge , removing relevant precedents before semantic ranking.