Organizations: Graduate School of Informatics, Nagoya University, Nagoya, Japan. · Graduate School of Arts and Sciences, The University of Tokyo, Tokyo, Japan. · Department of Economics, The Hong Kong University of Science and Technology, Hong Kong, China. · Center for Advanced Intelligence Project, RIKEN, Osaka, Japan.
Valuing player actions in football requires accounting for strategic interactions among 22 players, including off-ball movements and defensive positioning. Existing reinforcement-learning-based methods commonly aggregate decisions at the team level or estimate player values independently, leaving strategic interdependence among players insufficiently represented. This study proposes an action valuation framework inspired by Markov perfect equilibrium (MPE) for all players. Each possession is modeled as a finite-horizon dynamic game, with each player represented as an autonomous agent whose policy depends on the current game state. MPE is used as a motivating solution concept rather than an exact equilibrium. To improve interpretability, we use Expandable Decision-Making States (EDMS) and decompose the Q-value into a successor-feature basis and a linear reward-weight vector. The value basis is estimated by linear TD initialization followed by nonlinear refinement. Using tracking and event data from 95 J1 League matches, we compare the proposed formulation with an independent reinforcement learning baseline. Because the two formulations define TD errors in different target spaces, TD MSE is used only for within-formulation consistency. With EDMS fixed, the independent baseline assigns the highest value to forward movement in 99.21% of evaluated off-ball states, whereas the most frequent direction under the proposed formulation accounts for 17.63%. Team-level average Q-values show a negative association with season-level expected goals for the baseline and a weakly positive association for the proposed formulation. Qualitative analyses illustrate context-dependent valuations of off-ball movements and defensive positioning. Overall, the proposed formulation produces more context-sensitive action rankings, although the comparison does not isolate the MPE-inspired component.
Table 1: Summary of the state representations used in the experiments.
Action category
No. of classes
Actions
On-ball attacking actions
5
Pass, through pass, shot, cross, and dribble
Defensive action
1
Tackle, interception, clearance, or block, aggregated into one defensive-action class
Off-ball actions
9
Movement in eight directions and idle
Total
15
Table 2: Summary of the feasible discrete action classes.
Player role
Feasible actions
Attacking ball carrier
Pass, through pass, shot, cross, and dribble
Attacking off-ball player
Movement in eight directions and idle
Defending player
Movement in eight directions, idle, and defensive action
Table 3: State-dependent action masking according to player role. Actions not listed as feasible are masked.
Learning formulation
TD target space
State representation
TD MSE
Independent RL [ 6 ]
Scalar Q-value
PVS
6.6947×10−4
Independent RL [ 9 ]
Scalar Q-value
EDMS
4.5057×10−4
MPE-inspired RL
Successor features ( d=3 )
PVS
1.1229×10−2
MPE-inspired RL
Successor features ( d=3 )
EDMS
1.0876×10−2
Table 4: TD MSE for the independent RL and MPE-inspired formulations using PVS and EDMS. The independent RL TD error is defined over a scalar Q-value, whereas the MPE-inspired TD error is defined over the successor-feature vector. Therefore, TD MSE values should be compared only within the same learning formulation, not across formulations.
Figure 1: Distribution of highest-valued off-ball actions obtained from post-training Q-value inference. The off-ball action space consists of eight directional movements and an idle action. The idle action was included when determining the highest-valued action but is omitted from the figure for visualization. The independent RL baseline assigns the highest Q-value to forward movement in almost all evaluated off-ball states, whereas the proposed MPE-inspired model distributes the highest-valued actions across multiple directions. “Up” corresponds to the attacking direction toward the opponent’s goal.
Figure 2: Exploratory association between team-level Q-value summaries computed from the held-out tracking-data sample and official 2024 J1 League full-season xG. Each point represents one of the n=20 matched teams. The independent RL baseline shows a negative Spearman rank correlation, whereas the proposed MPE-inspired model shows a weakly positive correlation.
Figure 3: Off-ball Q-value visualization. The left panel shows all players and the ball, with the white marker indicating the ball and the yellow-highlighted player indicating the target player whose Q-values are visualized. The right panel shows the target player’s name and a radar chart of Q-values for off-ball movements in eight directions. “Up” corresponds to the attacking direction toward the opponent’s goal. This example shows off-ball action valuation when open space is available ahead of the target player.
Figure 4: Example of off-ball action valuation in response to a teammate’s forward movement. The visualization layout follows Fig. 3 .
Figure 5: Example of off-ball action valuation during a scoring opportunity. The visualization layout follows Fig. 3 .
Figure 6: Example of defensive positioning evaluation. The visualization layout follows Fig. 3 .
Figure 7: Effect of reward weight adjustment on the highest-valued action. The visualization format follows Fig. 3 . The lower-right panel shows the reward-feature weights θ . Under the default reward weights, θ=(1.0,−0.5,0.2)⊤ , the highest-valued action is backward-right.
Figure 8: Effect of increasing the attacking-third weight. The visualization format is the same as in Fig. 7 . After changing the reward weights to θ=(1.0,−0.5,1.0)⊤ , the highest-valued action changes from backward-right to forward-right.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: This figure illustrates the procedure for calculating the space score. From left to right, it shows: the Voronoi diagram considering each player’s position and velocity; the importance of each area on the pitch; and the final space score distribution, which is the product of the two. The area importance is modeled based on proximity to the opponent’s goal and distance from the center of the pitch, and it is maximized in the central area in front of the opponent’s goal
Category
Subcategory
Main variables
Absolute State
–
Offside-line distance, formation
Relative State
Off-ball State
Ball distance, pressure, pass-line pressure, space score, pass score
The defensive performance of football players is commonly measured through a limited number of actions like tackles and interceptions while their continuous impact through positional behaviour has hardly been studied before. We formulate this problem as an attribution over multi-agent spatiotemporal trajectories without player-level ground truth labels, where event-level changes of expected threat are distributed among individuals. We propose a framework that performs this attribution using player involvement scores calculated from defensive pressure areas (DPAs). By computing role-conditioned baselines within automatically detected team structures, we can determine each defender's expected responsibility for threat created through arbitrary passes. The validity and robustness of this approach are evaluated on a uniquely extensive cross-gender and cross-competition data set, including positional and event data from 64 matches of the men's World Cup, 116 matches of the women's German Bundesliga and 336 matches of the men's German 3. Liga. In the absence of a ground truth, we propose an evaluation protocol that combines multiple relatively weak proxies into robust summary scores. We find a validity score that is improved by around 1 standard deviation compared to the best action-based metric and demonstrate that many popular measures show limited validity. The "blame" for conceding high-value actions shows especially strong correlations with external ratings and market values, making it the first published metric in football to reliably measure positioning errors. All code underlying this work is publicly available to support reproducibility and further research.
Jonas Bischofberger, Runqing Ma, Pascal Bauer +2
Centre for Sport Science and University Sports, University of Vienna, Vienna, Austria · Vienna Doctoral School of Pharmaceutical, Nutritional and Sport Sciences (VDS-PhaNuSpo), University of Vienna, Vienna, Austria · VfB Stuttgart 1893 AG, Stuttgart, Germany +2
Vision-language models (VLMs) can describe a scene, but can they act well within one? We study whether VLMs can make sound strategic decisions, using soccer as an objective testbed with quantifiably-valued actions. We introduce SportD, a dataset and evaluation consisting of 1421 decision scenarios across professional men's and women's soccer games, where a VLM must decide what action to take next. Models on average select the optimal action around 27% of the time, less often than the professional players, and capture markedly less of the value at stake. Furthermore, they exhibit a clear preference for safer actions, favoring lower-variance, lower-value choices that also make less physical progress toward goal. Frontier VLMs are better at estimating whether an action will succeed, placing the highest-success-probability action among their top choices in 72-85% of cases. Yet VLMs systematically conflate likelihood with value, assigning higher value to actions that are more likely to succeed (ρ=+0.30 to +0.52), despite no such relationship in the ground truth (ρ=−0.08). Modifying the deliberation instructions to encourage risk-taking brings the frontier models closer to the players' skill levels. SportD opens a new direction for rigorously evaluating physical strategic decision-making in VLMs, showing that careful decomposition of their choices can reveal the mechanisms underlying systematic biases such as risk aversion.
Jasin Cekinmez, Addison J. Wu, Haotian Xia +6
Princeton University · Rice University · UC Irvine +2
Player-centric ball action spotting requires temporally precise event detection together with actor attribution in crowded, partially observed multi-agent sports videos. Existing Denoising Sequence Transduction (DST) baselines treat the player-role dimension as part of a flattened frame-level representation, which weakens the inductive bias for modeling player-specific temporal evolution and inter-player interactions. To address this limitation, we propose Multi-Entity Denoising Sequence Transduction (ME-DST). ME-DST keeps the role-slot dimension throughout encoding. It uses temporal attention to model the history of each role slot, and spatial attention to exchange information across role slots at each frame. This factorized design gives the model a direct structure for separating within-player evolution from inter-player context. We also add learnable role embeddings, tracking-derived tactical features, and fused visual predictions from X3D-L and Swin3D-S. Experiments on the FOOTPASS dataset show that ME-DST reaches a Micro F1 of 0.778. This improves the strongest official TAAD+DST baseline by 10.3 percentage points. Controlled ablations show that preserving the entity axis and encoding role identity are central to this gain. These results suggest that explicit entity modeling is an effective inductive bias for player-centric sports event understanding.
Ruifeng Wang, Di Yang, Jiangtao Wang
School of Artificial Intelligence & Data Science, USTC, Hefei, China · Suzhou Institute for Advanced Research, USTC, Suzhou, China