Long-horizon agent benchmarks typically report how far an agent progresses, but do not identify whether its performance comes from the foundation model, scaffold, responsibility scope, match-control granularity, or horizon. We introduce FromPitch2Board, a deterministic football-management benchmark that studies five configurable factors through controlled comparisons on a single simulator, using paired seeds and a frozen calibration. We evaluate four foundation models and four agent scaffolds. In the Model Track, Coach points Z-scores span 0.19, while Manager points Z-scores span 0.68, with GPT-5.6 showing a sharp rise in passivity under responsibility expansion. Its responsibility ladder rises from 46.1 to 58.1 points with recruitment, then falls to 46.8 under full management, localizing the regression to the final responsibility boundary. Across that boundary, its skipped-decision rate rises from 1.1% to 57.9%. Within the Flash-Pro pair crossed across every scaffold, scaffold choice changes Manager points Z-scores by up to 0.48 relative to the fixed stateless scaffold. The 3Y cohort shows a directional reversal in mean ranking between years one and three, while a selected Claude Code+Pro configuration peaks in year three and remains below that peak, showing that responsibility scope and horizon expose behavior changes that a single headline score conceals.
Figures & tables
Track
Scaffold
Model
Coach Z
Manager Z
Composition Gap
Model
Ours
DS Flash
+0.51
+0.67
−0.16
Ours
DS Pro
+0.59
+0.76
−0.17
Ours
Opus 5
+0.55
+0.61
−0.06
Ours
GPT-5.6
+0.70
+0.08
+0.62
Rule policy
Greedy reference
+0.08
−0.09
+0.17
Agent
Pi
DS Flash
+0.55
+1.08
−0.53
Table 1: Main leaderboard across model, scaffold, and responsibility. Rows average four scenarios with eight seeds each against the frozen calibration; bold marks per-track extrema.
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Horizon
Timescale
Scope axis
Harness × model
Det.
Multi-obj.
Same interface
FromPitch2Board
1–10y
min → career
Coach → Manager
Flash–Pro
yes
yes
yes
FM-Bench ( Wang et al., 2026 )
5–20y
season → career
no
no
yes
yes
yes
PokéAgent ( Karten et al., 2026a )
long
minute → run
no
partial
no
no
yes
CEO/YC ( Chen et al., 2026 , He et al., 2026a )
100s turns
day → year
no
no
NR
partial
yes
Enterprise ( Han et al., 2026 )
132mo
month → decade
no
no
NR
partial
yes
WebArena/OSWorld ( Zhou et al., 2024 , Xie et al., 2024 )
bounded
minutes
no
no
no
no
yes
Appendix
Table 2: Comparison with representative agent benchmarks. “NR” denotes a dimension not reported by the source paper.
Action
Coach
Recruiter
Manager
Continue (skipped decision)
∙
∙
∙
SetLineup
∙
∙
∙
SetTactics
∙
∙
∙
SetMatchPlan
∙
∙
∙
Substitute
∙
∙
∙
MatchTactics
∙
∙
∙
Appendix
Table 3: Action availability by responsibility scope. The scope is fixed for the whole episode and enforced by the runner: an out-of-scope action is rejected with an explanatory message and consumes the decision point. Substitute and MatchTactics are available to all three responsibility scopes only when in-match control is enabled.
Experiment
Factor varied (levels)
Run on
Main leaderboard, Model Track
foundation model (4)
Ours scaffold
Main leaderboard, Agent Track
scaffold (4)
Flash–Pro pair
Responsibility ladder
responsibility scope (3)
DS Pro stateless; Pi; GPT-5.6
Match-control granularity
in-match control (2)
DS Flash; DS Pro
Session continuity
session memory (2)
Claude Code + Flash
Thinking budget
reasoning cap (3)
DS Flash; DS Pro
Appendix
Table 4: Experiment matrix. The leaderboard crosses four scaffolds with the Flash–Pro model pair; the remaining axes are varied on the configurations listed.
Model / scope
crisis
moneyball
rebuild
title
DS Flash / Coach
+0.49
+0.61
+0.79
+0.16
DS Flash / Manager
+0.19
+0.27
+1.28
+0.95
DS Pro / Coach
+0.30
+0.62
+0.95
+0.48
DS Pro / Manager
+0.57
+0.97
+0.64
+0.87
Opus 5 / Coach
+0.38
+0.85
+0.77
+0.19
Opus 5 / Manager
+0.30
+0.25
+0.88
+1.03
Appendix
Table 5: Per-scenario Model-Track points Z-scores (8 seeds per entry).
Scaffold
Model
Coach Z ± SE
Manager Z ± SE
Ours
DS Flash
0.510±0.158
0.673±0.187
Ours
DS Pro
0.589±0.157
0.764±0.240
Ours
Opus 5
0.547±0.188
0.615±0.142
Ours
GPT-5.6
0.704±0.158
0.080±0.155
Pi
DS Flash
0.554±0.126
1.079±0.125
Pi
DS Pro
0.552±0.207
1.165±0.109
Appendix
Table 6: Seed-clustered uncertainty for the main points leaderboard. Each estimate uses four scenarios and eight seeds per scenario.
Configuration
Episodes
Manager Z
Input
Output
Cache read
Wall time
Cost (¥)
Ours + DS Flash
64
+0.67
11.8M
7.08M
1.67M
1,058s
25.5
Ours + DS Pro
64
+0.76
11.4M
4.88M
1.60M
1,017s
62.3
Claude Code + DS Flash (session continuity)
64
+1.02
15.8M
1.83M
467.6M
322s
28.9
Claude Code + DS Pro (session continuity)
64
+0.95
215.6M
2.34M
434.2M
661s
671.8
Pi + DS Flash
64
+1.08
13.8M
1.62M
374.1M
397s
24.5
Pi + DS Pro
64
+1.17
13.8M
2.11M
386.2M
532s
63.8
Appendix
Table 7 : Token, runtime, and normalized-cost accounting. Token and cost columns aggregate the episodes listed for each configuration; wall time is the mean per episode rather than a batch elapsed time. Dashes mark inapplicable scores.
Scope
Policy
Points Z
Coach
Random actions
−1.83
Coach
Best-XI heuristic
−0.30
Coach
Greedy reference
+0.08
Manager
Random manager
−2.83
Manager
Passive / no-op
−0.51
Manager
Greedy reference
−0.09
Appendix
Table 8: Environment baseline ordering. The expected weak-to-strong ordering holds on both responsibility scopes.
Responsibility
Overall
Crisis
Moneyball
Rebuild
Title
Steps
Coach only
45.0
35
50
50
44
38.0
+ Recruitment
58.4
41
65
57
71
47.0
+ Full management
56.8
38
65
55
69
47.9
Appendix
Table 9: DS Pro stateless responsibility ladder by scenario (eight seeds each; 96 episodes). Overall and scenario means are rounded independently.
Responsibility
Overall
Crisis
Moneyball
Rebuild
Title
Coach only
44.1
40.6
46.4
42.9
46.4
+ Recruitment
61.6
37.4
65.9
66.0
77.0
+ Full management
60.9
40.8
67.4
60.9
74.6
Appendix
Table 10: DS Pro responsibility ladder with Pi (eight seeds per scenario), where Coach and Manager correspond to the Agent-Track evaluations.
Responsibility
Overall
Crisis
Moneyball
Rebuild
Title
Coach only
46.1
36.9
53.3
47.5
46.9
+ Recruitment
58.1
41.9
55.1
62.8
72.6
+ Full management
46.8
32.0
49.3
48.8
57.3
Appendix
Table 11: GPT-5.6 responsibility ladder (eight seeds per scenario). Recruiter exceeds Manager by 11.28±2.28 paired points ( n=32 ).
Outcome
Recruiter
Manager
Manager − Recruiter
Paired SE
Points
58.09
46.81
−11.28
2.28
League position ↓
9.13
13.31
+4.19
1.00
Goal difference
8.75
−13.00
−21.75
4.05
Balance (£M)
61.53
67.32
+5.79
2.08
Net value (£M)
9.66
8.83
−0.83
2.07
Net spend (£M)
10.43
4.70
−5.73
2.08
Appendix
Table 12: GPT-5.6 Recruiter-to-Manager paired outcomes ( n=32 matched episodes). Full management improves liquidity but reduces sporting and squad-value outcomes.
Scaffold/model
Reasoning protocol
Coach Z
Manager Z
Ours / Flash
2048-token cap
+0.49
+0.87
Ours / Flash
non-binding
+0.51
+0.67
Ours / Pro
2048-token cap
+0.44
+0.46
Ours / Pro
non-binding
+0.59
+0.76
Pi / Flash
minimal
+0.66
+1.08
Pi / Flash
max
+0.55
+1.08
Appendix
Table 13: Protocol-alignment arms motivating the shared non-binding, max-reasoning protocol used by the main leaderboard.
Model
2048-token cap
8192-token cap
Non-binding
DS Pro
56.6
64.8
59.8
DS Flash
66.9
61.5
65.1
Appendix
Table 14: Thinking-budget ablation on title and rebuild (eight seeds each), reported as raw mean points.
Table 16: Match-control comparison on the eight paired episodes. Δ is control minus the default pre-match-only Coach score.
Configuration
Crisis Y1/Y2/Y3
Moneyball Y1/Y2/Y3
Rebuild Y1/Y2/Y3
Title Y1/Y2/Y3
Greedy
33/19/20
41/32/33
53/40/38
54/42/35
Claude Code + Flash (no memory)
47/44/60
60/54/72
57/69/81
71/77/82
Claude Code + Pro (no memory)
43/40/52
61/63/74
65/70/77
73/80/89
Ours + Flash
37/33/53
59/60/71
58/77/95
76/89/97
Ours + Pro
40/37/56
57/55/68
64/72/81
63/84/86
Appendix
Table 17 : Three-season scenario means (eight seeds per scenario; 32 trajectories per configuration). Figure 4 uses unrounded values.
Y1
Y2
Y3
Y4
Y5
Y6
Y7
Y8
Y9
Y10
Claude Code + Pro
70.1
75.8
85.4
74.2
65.4
47.1
48.3
46.2
48.6
41.3
Greedy
53.4
41.4
36.8
36.2
32.2
40.2
36.1
32.8
37.1
42.5
Claude Code + Pro, cumulative
70.1
145.9
231.3
305.5
370.9
418.0
466.3
512.5
561.1
602.4
Greedy, cumulative
53.4
94.8
131.6
167.8
200.0
240.2
276.3
309.1
346.2
388.7
Appendix
Table 18: Ten-season points, averaged over rebuild and title (eight seeds per scenario in every year). The agent peaks in year 3 and subsequently declines, finishing slightly below the greedy reference in year 10; across the full decade it nevertheless accumulates substantially more points than the reference ( 602.4 against 388.7 ), so the decline is relative to its own peak rather than a cumulative deficit. Cumulative rows sum the displayed annual means.
Scaffold/model
Scope
Crisis
Moneyball
Rebuild
Title
Pi / Flash
Coach
+0.38
+0.78
+0.91
+0.14
Pi / Flash
Manager
+0.17
+1.40
+1.31
+1.43
Pi / Pro
Coach
+0.59
+0.61
+0.63
+0.38
Pi / Pro
Manager
+0.58
+1.31
+0.90
+1.87
Claude Code / Flash
Coach
+0.36
+1.04
+0.61
+0.16
Claude Code / Flash
Manager
+0.67
+0.96
+0.69
+1.76
Appendix
Table 19: Per-scenario Agent-Track points Z-scores, computed against the frozen responsibility-scope- and scenario-specific calibration (eight seeds per entry).
Contrast
Δ
SE seed
SE pair
t
Manager points Z (32 paired cells, 8 seed clusters)
Pi − Ours (Flash)
+0.41
0.15
0.17
2.7
Claude Code − Ours (Flash)
+0.35
0.13
0.18
2.7
Codex − Ours (Flash)
+0.48
0.17
0.22
2.9
Pi − Ours (Pro)
+0.40
0.25
0.19
1.6
Claude Code − Ours (Pro)
+0.18
0.22
0.19
0.8
Appendix
Table 20: Paired contrasts. Each comparison matches configurations by scenario and seed. SEseed clusters observations by seed, while SEpair treats scenario–seed differences as independent. Three-year contrasts use 32 matched observations. The final block reports the year-three-minus-year-one interaction, that is how much each margin shifts between the two seasons, so a positive value is a widening lead rather than a single-season difference.
Configuration
Skipped decisions
Bid rate / step
Resolved-bid acceptance
Ours + DS Pro / Manager
16.9%
13.9%
33%
GPT-5.6 / Recruiter
1.1%
12.6%
35.0%
GPT-5.6 / Manager
57.9%
2.7%
38.7%
Appendix
Table 21: Action profiles. Skipped decisions counts decision points at which the agent issued no action and continued. GPT rates aggregate 32 paired episodes per responsibility rung; acceptance uses bids with explicit outcomes.
Scope
Trigger
Decision points
Skipped
Skipped (%)
Recruiter
Matchday
1,216
3
0.2
Recruiter
Market day
288
14
4.9
Manager
Matchday
1,216
705
58.0
Manager
Market day
310
179
57.7
Manager
Incoming offer
18
10
55.6
Appendix
Table 22: Where GPT-5.6 stops acting, by the event that triggered each decision point. Triggers are identified by deterministic replay of the recorded action sequences. Both rungs contain the same 1,216 matchday decision points.
Scenario
Recruiter
Manager
Δ (pp)
Positive pairs
Crisis
0.0%
57.2%
+57.2
8/8
Moneyball
1.0%
61.2%
+60.2
8/8
Rebuild
0.0%
58.6%
+58.6
8/8
Title
0.0%
54.9%
+54.9
8/8
All 32 pairs
median +57.9
32/32
Appendix
Table 23: Matchday skipped-decision rate per episode, paired across the two responsibility rungs. Every matched episode increases, so the aggregate change is not driven by a subset of trajectories. The two rungs share the same 1,216 matchday decision points; their state trajectories diverge through the agents’ own decisions.
Scenario / metric
Y1
Y3
Y4
Y6
Y9
Y10
Rebuild points
66.9
83.9
69.5
36.9
47.1
40.4
Rebuild net value (£M)
4.2
11.8
−10.7
7.3
51.7
48.5
Rebuild squad size
23.0
18.5
11.8
12.5
15.1
11.6
Title points
73.4
86.9
79.0
57.4
50.1
42.1
Title net value (£M)
−2.2
−0.6
−25.6
−19.0
19.1
30.2
Title squad size
25.4
20.0
13.9
14.2
14.2
14.6
Appendix
Table 24: Claude Code+Pro 10Y checkpoints (eight seeds per scenario in every reported year).
Season
Bought
Sold
Released
Squad at end
1
5.6
4.1
0.0
24.2
2
1.7
3.7
2.5
22.6
3
1.1
2.4
6.9
19.2
4
2.7
1.7
5.6
12.8
5
2.1
2.6
3.8
12.3
6
1.2
1.4
0.2
13.4
Appendix
Table 25: Claude Code+Pro roster flow over the first nine seasons, averaged over rebuild and title (eight seeds per scenario). Released denotes contract expiry. Squad size also moves with youth intake and retirements, which are not recorded as transfer movements, so the columns need not sum to the change in squad size.