Frontier coding performance is typically attained with large, costly proprietary models. We introduce ledger-based zero-shot self-orchestration (GVS5H), a training-free method in which fresh instances of one model decompose problems and coordinate through a shared file system. Across eleven open and closed-weight models on the 100 latest hard LiveCodeBench problems, the method yields as much as 25.6 points improvement, boosting several cheaper models to frontier-level performance. Orchestrated Qwen3.8 Flash Next scores 93.0% against Fable 5's 90.4% at 9% of the cost, while the smaller Qwen3.8-27B reaches 92.4%. Gains are not universal: some models are unchanged or worse. Transcript analysis attributes the gain to decomposition and persistent context. Inference-time organization can reach or exceed frontier coding accuracy at a fraction of the cost on self-hostable weights.
Figures & tables
Model
Single call
With manager
Δ vs single
p
Δ vs Fable 5
95% bound
p
Claude Fable 5
90.4±1.4
—
—
—
reference
—
—
GPT-5.6-Terra
80.8±1.0
88.0±1.2
+7.2
8.0×10−5
−2.4
−5.8
0.88
GPT-5.6-Luna
70.4±4.6
81.2±2.8
+10.8
<2.5×10−5
−9.2
−13.4
0.0013
Qwen3.8-27B
66.8±5.2
92.4±2.3
+25.6
<2.5×10−5
+2.0
−0.8
0.88
Qwen3.8-Flash-Next
84.2±1.8
93.0±1.2
+8.8
<2.5×10−5
+2.6
−0.1
0.56
DeepSeek-V4.1-Flash
88.4±2.6
92.0±1.8
+3.6
0.002
+1.6
−1.3
0.88
Table 1: Manager vs. single call, five passes. pass@1 on LCB-100 at a 128k cap, reasoning on; ± is a 95% t interval across the five passes (df = 4). Δ is in percentage points, against the model’s own single call and against Fable 5’s single call. Each p is a paired sign-flip permutation test, unit = problem ( n=100 ), Holm-corrected within its family of five; < marks the permutation floor. 95% bound is the one-sided lower bound on Δ vs Fable 5 over the same per-problem differences (t, df = 99): the largest deficit the data leave open.
Arm
Rate $/MTok in / out
In (MTok)
Out (MTok)
$/pass
$/solved
p vs single
p vs Fable 5
Qwen3.8-27B single
0.21/2.55
0.0753
7.4247
$18.95
$0.28
—
—
Qwen3.8-27B manager
0.21/2.55
1.5053
18.6277
$47.82
$0.52
1.5×10−4
6.1×10−4
Qwen3.8-Flash-Next single
0.15/0.47
0.0753
5.7751
$2.73
$0.032
—
—
Qwen3.8-Flash-Next manager
0.15/0.47
1.2930
11.8380
$5.76
$0.062
3.5×10−4
3.6×10−7
DeepSeek-V4.1-Flash single
0.30/1.20
0.0680
2.8626
$3.46
$0.039
—
—
DeepSeek-V4.1-Flash manager
0.30/1.20
1.4993
12.4873
$15.43
$0.17
1.6×10−4
2.2×10−9
Table 2: What one pass cost each arm. List rate × the tokens it consumed. The two p columns test each manager arm against its own single call and against Fable 5’s single call: Welch over the five passes, Holm-corrected across 11 comparisons. Qwen3.8-27B’s single arm is the 128k cap-matched one; at 250k it costs 26.95/pass.Retriedanddiscardedattemptsarecounted:theyweregeneratedandwouldbebilled.ThecheapestarmisGPT−5.6−Lunasingleat0.41 a pass and the most accurate is Qwen3.8-Flash-Next manager at 5.76—a14\times$ spread in price for +22.6 points.
Model
Cap hits single
Cap hits manager
No code s / m
of which refusals
pass@1 s / m
Emitted only s / m
Cap
GPT-5.6-Terra
0 / 500
0 / 3,392
0 / 0
0
80.8 / 88.0
80.8 / 88.0
128k
GPT-5.6-Luna
0 / 500
0 / 3,498
0 / 0
0
70.4 / 81.2
70.4 / 81.2
128k
Qwen3.8-27B
150 / 500
108 / 3,235
35 / 0
0
66.8 / 92.4
71.8 / 92.4
128k a
Qwen3.8-27B (250k)
124 / 500
—
21 / —
0
70.0 / —
73.1 / —
250k a
Qwen3.8-Flash-Next
51 / 500
92 / 3,141
20 / 0
0
84.2 / 93.0
87.7 / 93.0
128k
DeepSeek-V4.1-Flash
11 / 500
49 / 3,525
10 / 1
0
88.4 / 92.0
90.2 / 92.2
128k
Table 3: Cap hits and empty solutions per arm. A cap hit is a generation that stopped at the token limit ( finish_reason=length ). No code counts any run containing no code, due to a cap hit or other reason.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Single call
With manager
Model
As graded
Δ harness
Δ keys
Re- ported
As graded
Δ harness
Δ keys
Re- ported
Pinned backends, five passes
Claude Fable 5
87.4
+0.0
+3.0
90.4
—
—
—
—
GPT-5.6-Terra
69.8
+8.4
+2.6
80.8
72.6
+12.6
+2.8
88.0
GPT-5.6-Luna
60.4
+8.0
+2.0
70.4
70.4
+8.4
+2.4
81.2
Qwen3.8-27B
62.4
+2.6
+1.8
66.8
84.4
+5.4
+2.6
92.4
Appendix
Table 4: Every reported arm as the stock evaluator graded it, and after each correction (Appendix B ). As graded is LiveCodeBench’s own verdict on the stored generations. Harness is the move contributed by running each submission as a subprocess, which is what repairs the sys.stdin mock and sys.stdout.buffer ; keys is the move contributed by judging the four problems that accept more than one answer by the contest’s rule. Reported is the number used everywhere else in this paper. Upper block: five-pass means at a 128k cap, reasoning on (Table 1 ). Lower block: the 128k reasoning-on, one-pass OpenRouter arms (Table 6 ); Qwen3.5-9B is unusable in this condition and is omitted. Across the 48 arm-passes of the two remaining OpenRouter conditions the harness correction moves six by one point and the key correction moves each arm by 0 to 3 points. Qwen3.8-27B’s single call is the one row where as graded is not the stock evaluator on the code being reported: that arm is cap-matched by replaying its 250k generations at 128k (Appendix A ) and the truncated solutions were never put through the stock evaluator, so the column is the 250k grading and the two Δ columns carry the cap match as well as the corrections. Against that baseline the row gains 35 problem-passes and loses 13, eight of them generations whose 128k prefix contains no code at all.
Model
Manager only
Single call only
Problem-passes
Exact McNemar p
GPT-5.6-Terra
47
11
500
2.0×10−6
GPT-5.6-Luna
72
18
500
8.1×10−9
Qwen3.8-27B
134
6
500
1.4×10−32
Qwen3.8-Flash-Next
49
5
500
3.9×10−10
DeepSeek-V4.1-Flash
22
4
500
5.3×10−4
Appendix
Table 5: Problem-passes won by one arm alone. Of the 500 problem-passes per model (100 problems × 5 passes), the number the manager arm solved and the single call did not, against the number the single call solved and the manager did not. The gains are not compensating wins and losses. The test is exact McNemar on those discordant pairs (Section 2.2 ).
Model
Params
128k ⋅ ON (1 pass)
128k ⋅ OFF (1 pass)
16k ⋅ OFF ( × 5)
Opus-5
n/a
88 → 95 ( +7 )
–
–
Kimi-K3
∼2.8 T
86 → 86 (+0)
33 → 76 ( +43 )
33.4 → 65.0 ( +31.6 )
MiniMax-M3
428B
62 → 69 ( +7 )
26 → 40 ( +14 )
22.2 → 33.2 ( +11.0 )
Qwen3.6-35B
35B
26 → 44 ( +18 )
35 → 27 ( − 8)
28.2 → 27.4 ( − 0.8)
Qwen3.5-9B
9B
unusable
18 → 21 (+3)
15.0 → 22.6 ( +7.6 )
Appendix
Table 6: LCB-100 pass@1 (%), single → manager. For the OpenRouter-served models on the original scaffold. “–” = not run; Qwen3.5-9B reasoning-on returns reasoning-only replies and is unusable. Scores are from the corrected grader (Appendix B ): the four LCB-100 problems whose reference answer is not unique are judged by the original contest rule, and submissions run as real subprocesses so sys.stdout.buffer behaves as it does on the contest judge. A bolded delta clears the ∼4.5 point pass-to-pass band. The eleven pinned-backend, five-pass arms are reported separately in Section 3 and are not pooled here, because both the serving path and the scaffold version differ (Section 2.2 ).
Model (condition)
Problem
How the single call failed
What the scaffold added
Qwen3.5-9B (128k, off)
abc385_d
Timed out: a naïve per-segment scan of a path simulation.
The brainstorm flagged in writing that the check is O(N⋅M)≈4×1010 and prescribed a range-query structure before any code - the up-front planning a single pass skips (Section 4 ).
Qwen3.6-35B (on)
abc394_f
Cut off mid-reasoning at 32,768 tokens; returned nothing.
Solved in one worker cycle: the brainstorm fixed the structural insight (an “alkane” subtree needs a degree-4 center and ≥5 vertices); the worker used an iterative DFS to avoid a recursion-depth failure. The same 32k clamp was hit 117 times across this model’s manager calls, but with several attempts per problem one clamped call no longer costs the problem.
MiniMax-M3 (on)
3701
Printed nothing on cdcd : the lexicographic-reconstruction branch was broken.
Split the two coupled objectives across worker cycles - one worker on the forward cost DP over capped run-lengths, a separate worker on a suffix DP used purely for the lexicographic reconstruction.
Kimi-K3 (128k, off)
3687
Over-counted the longest unique-value path (9 and 3 where the answers were 6 and 2), never shrinking its window on a repeated value.
The brainstorm wrote the reduction into notes.md first - longest substring without repeating characters along each root-to-leaf path, window start via last-seen depths, O(n) DFS. Only then did a worker implement it.
GPT-5.6-Terra (on)
3688
4,615 characters implementing a Segment-Tree-Beats-shaped structure indexed by candidate value, and wrong.
The plan named the difficulty in advance - the state must be compressed rather than maintaining all distinct values per index. Nine calls later, the manager arm returned 1,860 characters over one standard segment tree. The contribution is subtraction . Manager-only win in four of five passes.
GPT-5.6-Luna (on)
abc397_e
Carried a set of unfinished path lengths per subtree across 4,092 characters, a general state it never closed.
The brainstorm established the bounding lemma (at most two unfinished paths at any vertex), after which a worker implemented the postorder scan directly in 2,166 characters over five calls. Manager-only win in two of five passes.
Appendix
Table 7: One transcript per model where the single call failed and the manager arm solved the problem (Section 4 ). LCB problem IDs as in release_v6; call counts are manager-arm model calls. Character counts are of the emitted program.