Soccer analytics draws on two kinds of information: spatio-temporal data describing where players and the ball are, and event data describing what they do. Public datasets offer them apart, or together only on broadcast footage that leaves players outside the frame unobserved. SoccerTrack v2 combines continuous full-pitch video, long player trajectories and actor-linked events in one resource: ten university-level matches, 932 minutes of fixed-camera 4K panoramic video, annotated per frame with metric pitch coordinates, jersey numbers and persistent identities, roles and team sides for all players, and with ball action events in twelve classes, linked to the acting players through the same identifiers used in the trajectories. We fix a match-level split and report baselines for two tasks. For game state reconstruction, we run a full pipeline over all twenty halves and find that GS-HOTA scores degrade as sequence length increases. For ball action spotting, we train a model on the player trajectories, with and without the ball track. The data, the split and the evaluation tooling are released so that both tasks can be developed and compared at match length on the same footage.
Figures & tables
Figure 1 : Example panoramic frames from SoccerTrack v2: the released field of view spans the pitch across venues and lighting conditions.
Figure 2 : One match at a glance: M2, first half. (a) A frame of the released panoramic video at the instant of a Throw In event, 24:40 into the half: all twenty-two players are annotated, drawn as markers carrying their jersey numbers, with the acting player named by the event (jersey 21) in red. (b) The game state at the same frame, drawn from the game state annotations: all twenty-two players in metric pitch coordinates, coloured by team side, the event’s actor at (−10.5,35.4) m, beyond the touchline to take the throw. (c) The ball action layer of the same half: all 1,138 events over 45 continuous minutes, one row per class, coloured by acting team, with the pictured instant marked. The events name their actors in the same identity space as the tracks, so the three views resolve to the same players at the same instants.
Figure 3 : Dataset statistics. (a) Ball action class distribution over the 21,428 -event benchmark (log scale): the two most frequent classes account for 82% of events and the support spans a factor of 300 . (b) Ball action events per match; asterisks mark the two matches whose recorded ball track is clamped to the pitch rectangle (Table 1 ). (c) Density of all 30.7 million annotated player positions of the game state layer in metric pitch coordinates, pooled over the ten matches (log colour scale). Positions are annotated on a 100×100 grid of the pitch, so each histogram cell measures 1.05×0.68 m; 0.2% of positions fall outside the displayed margin. (d) How long identities stay annotated. Each of the 522 ground-truth tracks follows one player through one half; the histogram counts tracks by their annotated duration (log count scale). The median track is annotated for 45 continuous minutes, while the dashed line marks the 30 -second clips on which existing game state benchmarks are defined Somers et al. (2024b) .
Match
Duration (min)
Events
M1
90
2,085
M2
92
2,252
M3
97
2,117
M4
97
2,129
M5
97
1,945
M6
94
2,106
Table 1 : The ten SoccerTrack v2 matches, labelled M1 to M10 throughout the paper. Duration is the annotated play time of each match at 25fps . Every frame carries game state annotations for all players on the pitch: metric pitch coordinates on a canonical 105×68m system, jersey number and persistent player identity, role, and team side. Events counts the twelve-class ball actions. A curated tracking subset, one four-minute clip per match with every player boxed in every frame ( 59,842 frames and 1.3 million boxes in total), is described in Section 4.5 . ∗ marks the two matches whose recorded ball track is clamped to the pitch rectangle and therefore never leaves play.
Dataset
Video released
Pitch state
Events
Length
SoccerNet Giancola et al. (2018) ; Deliege et al. (2021) ; Cioppa et al. (2022) ; Cioppa et al. (2021)
broadcast, full matches
estimated positions of visible persons, no identities; image-space boxes on 30 s clips
110,458, not actor-linked
500 matches, 764 h
SoccerNet-GSR Somers et al. (2024b)
broadcast, 30 s clips
visible persons only
none
200 clips, 100 min
SportsMOT Cui et al. (2023)
broadcast clips, three sports
no
none
240 clips, 150 k frames
SoccerNet-GAR Karki et al. (2025)
broadcast, 4.5 s event clips
all 22 and ball, refined from broadcast
87,939, group-level
64 matches, event windows only
FOOTPASS Ochin et al. (2026)
broadcast, full matches (NDA)
all 22, partly imputed (9 to 12 visible)
102,992, actor-linked
54 matches
Bassek et al. Bassek et al. (2025)
withheld
all 22, observed at 25 Hz
11,137, actor-linked
7 matches
Table 2 : Positioning of SoccerTrack v2. The first four groups are defined by where the game state comes from relative to the released video: inferred from the released broadcast, measured by an instrument separate from the released camera, released without video, or rendered. The final group is defined by footage instead: releases of full-pitch video, with this release last; whether its provider measured the positions on the released view is not documented (Section 4.1 ). Events counts released event annotations and states whether each names the acting player. Length is as each release reports it.
Opening 30 s
Full half (45 min)
Half
GS-HOTA
Attrs off
GS-HOTA
DetA
AssA
LocA
Attrs off
M1, 1st
23.10
27.37
6.68
4.29
10.50
71.77
8.46
M1, 2nd
6.49
31.31
3.60
1.15
11.25
72.06
7.44
M2, 1st
55.20
66.12
14.28
10.04
20.37
83.02
22.14
M2, 2nd
26.49
49.15
12.13
8.01
18.38
83.00
22.08
M3, 1st
36.49
61.75
16.59
10.78
25.73
82.73
26.09
Table 3 : Game state reconstruction on all twenty halves, each run and scored over the whole half and, independently, on its opening 30 seconds, the clip length of the SoccerNet-GSR benchmark: the game state reconstruction pipeline of Section 4.7 , configured by the authors on released footage that includes a test-split half and run end to end under one configuration, the per-detection team assignment. GS-HOTA is the official metric, which classes detections by (role, team, jersey); DetA, AssA and LocA are its detection, association and localisation components; Attrs off rescores the same predictions with attribute matching disabled, so that only geometry and association count (components in Supplementary Table 1 ). The test split of the released match-level split is marked † . Means are unweighted means of the per-half scores in each column, computed from unrounded scores.
Configuration
GS-HOTA
DetA
AssA
Marginal
Geometry and association only
26.04
51.21
13.41
ceiling
+ role
25.64
49.36
13.48
−0.40
+ role, team
22.36
35.12
14.37
−3.27
+ role, team, jersey (official)
18.09
14.67
22.32
−4.27
Table 4 : Marginal cost of each game state attribute, computed on the first half of match M7 (the best-performing half of the test split in Table 3 ). GS-HOTA classes detections by the triple (role, team, jersey), so a detection with any incorrect attribute cannot match ground truth at all; each row switches one further attribute on, and Marginal is the incremental cost of enabling that attribute after those in the preceding row, computed from unrounded scores, so it may differ by 0.01 from the difference of the displayed values. DetA and AssA are the detection and association components of the metric. The first row, with no attributes scored, is the ceiling available to the official metric: perfect attributes would score exactly this, so it carries no marginal cost.
Figure 4 : Game state reconstruction accuracy against the length scored. Filled markers: the whole-half predictions of every half (Table 3 ) scored over their first 30 seconds, 1, 2, 5 and 10 minutes and the whole half, mean over the twenty halves (Supplementary Table 2 ), under the official metric (black) and with attribute matching disabled (blue); the bands span the twenty halves. Hollow markers: the mean of the independent 30-second runs of Table 3 , in which the pipeline saw only the clip.
τ=1 s
τ=5 s
Input
Test fold
mAP
mAP w
mAP
mAP w
trajectory
M7, M9 (Challenge)
0.207
0.202
0.472
0.637
M1, M2
0.222
0.272
0.501
0.676
M3, M4
0.231
0.259
0.600
0.709
M5, M6
0.199
0.258
0.516
0.678
M8, M10
0.252
0.275
0.531
0.664
Table 5 : Five-fold cross-match evaluation of the spotting model on the supplied trajectories. Each match appears in the test set exactly once; fold 0 is the SoccerTrack Challenge 2025 pair. Validation for each fold is the following fold’s test pair, so no fold’s decoding or stopping epoch is chosen on its own test matches; the network settings shared by all folds were fixed once on M2 and M10 (Section 4.9 ). Each entry is a single training run. The last block gives the two matches of the Challenge fold separately: the ball track is unclamped in M7 and clamped to the pitch rectangle in M9, so the value of the ball differs sharply between them and the fold’s pooled row averages two regimes.
Class
n
chance
trajectory
+ball
Pass
9315
0.042
0.286
0.853
Drive
8255
0.036
0.227
0.830
High Pass
1157
0.007
0.320
0.745
Out
771
0.004
0.148
0.364
Cross
394
0.001
0.327
0.794
Throw In
385
0.006
0.254
0.910
Table 6 : Per-class average precision at τ=1 s, pooled over the five cross-match folds and therefore over all ten matches. n is the number of ground-truth events across the benchmark. Rows marked † have n<100 . Classes are ordered by support. Because AP is pooled over matches, the column means differ slightly from the fold-averaged means of Table 5 .
Figure 5 : From released pixels to metric pitch coordinates, match M2, first half. (a) A frame of the released panoramic video: the stitched view is distorted and the pitch lines are curved. (b) The same frame after distortion correction, with the FIFA-standard pitch template and the frame’s annotated player positions (coloured by team side) projected into the image through a homography fitted on the hand-annotated pitch keypoints, shown over the pitch region with the boxed region enlarged below. (c) The same instant in the metric coordinate system of the annotations: a 105×68 m pitch with the origin at the centre circle, the touchline nearer the camera at the bottom.
Figure 6 : The two baselines. (a) Game state reconstruction. White boxes are published components; highlighted elements mark our contributions: the calibration stage fitted from the hand-annotated keypoints, the relaxed jersey-region gate, the per-detection team-assignment variant compared with the original on nested prefixes (Section 4.7 ), and the four fixes that let the pipeline run at match length (Section 4.8 ). (b) Ball action spotting, every stage of which is ours (Section 4.9 ). Per-frame features computed from the supplied ground-truth player trajectories and sampled at 5Hz , 82 in four groups and optionally 19 more from the ball track, pass through a non-causal temporal convolutional network: a pointwise projection to 64 channels, six residual blocks with dilations 1 to 32 (receptive field ±25s ) and a pointwise projection to twelve per-frame logits. Independent sigmoids score each class, and per-class peak picking above a floor with within-class non-maximum suppression yields ranked spots per class.
Half
DetA
AssA
LocA
Tracklets
Identities
First 30 s, all twenty halves
M1, 1st
21.60
35.39
73.05
24
22
M1, 2nd
25.33
39.19
73.59
27
22
M2, 1st
60.06
73.36
85.98
26
22
M2, 2nd
41.67
58.57
82.69
33
22
M3, 1st
57.79
66.38
85.65
25
22
Supplementary Table 1 : Components of the attributes-off scores of Table 3 : the same predictions rescored with attribute matching disabled, so that DetA, AssA and LocA measure geometry and association alone; Tracklets is the number of predicted identities and Identities the number of annotated identities in the sequence, substitutes included. The upper block is the opening 30 seconds of every half; the lower block is the full half of the same halves. Configuration is as in that table; the released test split is marked † . Means are per column within each block, computed from unrounded scores.
Half
30 s
1 min
2 min
5 min
10 min
20 min
Whole half
Official GS-HOTA
M1, 1st
14.28
10.70
8.18
8.60
6.32
6.39
6.68
M1, 2nd
9.90
7.08
5.48
5.77
3.92
4.14
3.60
M2, 1st
29.15
21.02
17.24
17.44
16.87
17.01
14.28
M2, 2nd
25.45
25.33
23.47
26.12
20.70
16.99
12.13
M3, 1st
32.19
29.11
29.82
21.39
17.31
17.16
16.59
Supplementary Table 2 : GS-HOTA of the whole-half predictions of Table 3 scored over growing windows from kickoff: the first 30 seconds, 1, 2, 5, 10 and 20 minutes, and the whole half (Methods). The predictions are fixed and only the scored window changes, so the last column reproduces Table 3 . The upper block is the official metric, the lower block the same windows with attribute matching disabled. Means are per column, computed from unrounded scores; the released test split is marked † .
We present our submission to the SoccerNet 2026 Player-Centric Ball Action Spotting challenge, which uses a two-stage pipeline: a Track-Aware Action Detector (TAAD) produces per-player action logits from broadcast video, and a Denoising Sequence Transduction (DST) transformer converts game-state features and TAAD logits into structured event sequences. We improve the TAAD with a temporal transformer that adds cross-frame context, alongside several training fixes. For the DST stage, we introduce a two-stage per-player attention mechanism operating on game-state features, and show that a spatial-first attention ordering (cross-player attention before temporal attention) improves validation Macro-F1 by 1.87%. To exploit architectural diversity, we train four model variants and combine them with a Weighted Event Fusion ensemble that applies agreement filtering to suppress single-model false positives while preserving recall, plus a dedicated exception for the rare tackle class. Our final system improves the challenge Macro-F1 from a baseline of 48.6 to 58.94.
We describe our system for the SoccerNet 2026 Player-Centric Ball-Action Spotting Challenge, which requires predicting who performs which action and when, across eight classes in broadcast soccer. Building on the three FOOTPASS baselines [1] (TAAD, TAAD+GNN, and TAAD+DST), we contribute four extensions: (1) gradient check pointing to enable full-backbone fine-tuning on a single GPU; (2) fusion of GNN logits into the DST encoder, combining graph-based tactical context with per-player visual features; (3) square-root frequency class weighting to address the 213:1 pass-to-tackle imbalance in the training data; and (4) a post processing pipeline comprising per-class logit gating, temporal frame refinement, jersey re-assignment, and a two-model ensemble. Our system achieves 0.548 Macro F1 on the test set and 0.446 on the challenge set (server evaluation).
Player-centric ball action spotting requires temporally precise event detection together with actor attribution in crowded, partially observed multi-agent sports videos. Existing Denoising Sequence Transduction (DST) baselines treat the player-role dimension as part of a flattened frame-level representation, which weakens the inductive bias for modeling player-specific temporal evolution and inter-player interactions. To address this limitation, we propose Multi-Entity Denoising Sequence Transduction (ME-DST). ME-DST keeps the role-slot dimension throughout encoding. It uses temporal attention to model the history of each role slot, and spatial attention to exchange information across role slots at each frame. This factorized design gives the model a direct structure for separating within-player evolution from inter-player context. We also add learnable role embeddings, tracking-derived tactical features, and fused visual predictions from X3D-L and Swin3D-S. Experiments on the FOOTPASS dataset show that ME-DST reaches a Micro F1 of 0.778. This improves the strongest official TAAD+DST baseline by 10.3 percentage points. Controlled ablations show that preserving the entity axis and encoding role identity are central to this gain. These results suggest that explicit entity modeling is an effective inductive bias for player-centric sports event understanding.
Ruifeng Wang, Di Yang, Jiangtao Wang
School of Artificial Intelligence & Data Science, USTC, Hefei, China · Suzhou Institute for Advanced Research, USTC, Suzhou, China