Organizations: Graduate School of Informatics, Nagoya University, Nagoya, Japan. · Graduate School of Arts and Sciences, The University of Tokyo, Tokyo, Japan. · Department of Economics, The Hong Kong University of Science and Technology, Hong Kong, China. · Center for Advanced Intelligence Project, RIKEN, Osaka, Japan.
Valuing player actions in football requires accounting for strategic interactions among 22 players, including off-ball movements and defensive positioning. Existing reinforcement-learning-based methods commonly aggregate decisions at the team level or estimate player values independently, leaving strategic interdependence among players insufficiently represented. This study proposes an action valuation framework inspired by Markov perfect equilibrium (MPE) for all players. Each possession is modeled as a finite-horizon dynamic game, with each player represented as an autonomous agent whose policy depends on the current game state. MPE is used as a motivating solution concept rather than an exact equilibrium. To improve interpretability, we use Expandable Decision-Making States (EDMS) and decompose the Q-value into a successor-feature basis and a linear reward-weight vector. The value basis is estimated by linear TD initialization followed by nonlinear refinement. Using tracking and event data from 95 J1 League matches, we compare the proposed formulation with an independent reinforcement learning baseline. Because the two formulations define TD errors in different target spaces, TD MSE is used only for within-formulation consistency. With EDMS fixed, the independent baseline assigns the highest value to forward movement in 99.21% of evaluated off-ball states, whereas the most frequent direction under the proposed formulation accounts for 17.63%. Team-level average Q-values show a negative association with season-level expected goals for the baseline and a weakly positive association for the proposed formulation. Qualitative analyses illustrate context-dependent valuations of off-ball movements and defensive positioning. Overall, the proposed formulation produces more context-sensitive action rankings, although the comparison does not isolate the MPE-inspired component.
Table 1: Summary of the state representations used in the experiments.
Action category
No. of classes
Actions
On-ball attacking actions
5
Pass, through pass, shot, cross, and dribble
Defensive action
1
Tackle, interception, clearance, or block, aggregated into one defensive-action class
Off-ball actions
9
Movement in eight directions and idle
Total
15
Table 2: Summary of the feasible discrete action classes.
Player role
Feasible actions
Attacking ball carrier
Pass, through pass, shot, cross, and dribble
Attacking off-ball player
Movement in eight directions and idle
Defending player
Movement in eight directions, idle, and defensive action
Table 3: State-dependent action masking according to player role. Actions not listed as feasible are masked.
Learning formulation
TD target space
State representation
TD MSE
Independent RL [ 6 ]
Scalar Q-value
PVS
6.6947×10−4
Independent RL [ 9 ]
Scalar Q-value
EDMS
4.5057×10−4
MPE-inspired RL
Successor features ( d=3 )
PVS
1.1229×10−2
MPE-inspired RL
Successor features ( d=3 )
EDMS
1.0876×10−2
Table 4: TD MSE for the independent RL and MPE-inspired formulations using PVS and EDMS. The independent RL TD error is defined over a scalar Q-value, whereas the MPE-inspired TD error is defined over the successor-feature vector. Therefore, TD MSE values should be compared only within the same learning formulation, not across formulations.
Figure 1: Distribution of highest-valued off-ball actions obtained from post-training Q-value inference. The off-ball action space consists of eight directional movements and an idle action. The idle action was included when determining the highest-valued action but is omitted from the figure for visualization. The independent RL baseline assigns the highest Q-value to forward movement in almost all evaluated off-ball states, whereas the proposed MPE-inspired model distributes the highest-valued actions across multiple directions. “Up” corresponds to the attacking direction toward the opponent’s goal.
Figure 2: Exploratory association between team-level Q-value summaries computed from the held-out tracking-data sample and official 2024 J1 League full-season xG. Each point represents one of the n=20 matched teams. The independent RL baseline shows a negative Spearman rank correlation, whereas the proposed MPE-inspired model shows a weakly positive correlation.
Figure 3: Off-ball Q-value visualization. The left panel shows all players and the ball, with the white marker indicating the ball and the yellow-highlighted player indicating the target player whose Q-values are visualized. The right panel shows the target player’s name and a radar chart of Q-values for off-ball movements in eight directions. “Up” corresponds to the attacking direction toward the opponent’s goal. This example shows off-ball action valuation when open space is available ahead of the target player.
Figure 4: Example of off-ball action valuation in response to a teammate’s forward movement. The visualization layout follows Fig. 3 .
Figure 5: Example of off-ball action valuation during a scoring opportunity. The visualization layout follows Fig. 3 .
Figure 6: Example of defensive positioning evaluation. The visualization layout follows Fig. 3 .
Figure 7: Effect of reward weight adjustment on the highest-valued action. The visualization format follows Fig. 3 . The lower-right panel shows the reward-feature weights θ . Under the default reward weights, θ=(1.0,−0.5,0.2)⊤ , the highest-valued action is backward-right.
Figure 8: Effect of increasing the attacking-third weight. The visualization format is the same as in Fig. 7 . After changing the reward weights to θ=(1.0,−0.5,1.0)⊤ , the highest-valued action changes from backward-right to forward-right.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: This figure illustrates the procedure for calculating the space score. From left to right, it shows: the Voronoi diagram considering each player’s position and velocity; the importance of each area on the pitch; and the final space score distribution, which is the product of the two. The area importance is modeled based on proximity to the opponent’s goal and distance from the center of the pitch, and it is maximized in the central area in front of the opponent’s goal
Category
Subcategory
Main variables
Absolute State
–
Offside-line distance, formation
Relative State
Off-ball State
Ball distance, pressure, pass-line pressure, space score, pass score
Centre for Sport Science and University Sports, University of Vienna, Vienna, Austria · Vienna Doctoral School of Pharmaceutical, Nutritional and Sport Sciences (VDS-PhaNuSpo), University of Vienna, Vienna, Austria · VfB Stuttgart 1893 AG, Stuttgart, Germany +2