Corporate and investment banking teams use presentations to support credit decisions and advise clients on financing and transactions. Producing these decks requires reconciling financial data, tracing sources and turning analysis into a recommendation. We retrospectively study the development of an agentic harness combining a 27B language model, financial calculations, narrative templates and validation checks. LLM judges guide engineering changes and assess the resulting decks, raising the question of whether higher scores reflect better documents or changes in grading. In shared-session text-only grading with template markers removed, five judges score the complete system 20.4 to 33.6 points out of 95 above the same model generating directly from a short prompt. Every judge scores the system higher on all seventeen development deliverables. Margins against direct Opus generation from a short prompt range from -4.7 to +0.8 points. Judges agree on broad progress across development rounds but agree less on final-deck rankings than on pooled scores. Repeated grading also shifts scores on unchanged decks, making small improvements difficult to distinguish from judge variability.
Figures & tables
Figure 1: The hybrid generation system and its development loop. Blue boxes denote model roles; deterministic prose can bypass the conditional model role. The evaluated contrast starts from collected inputs. Collection services and the application’s banker and supervisor review loop lie outside this contrast.
Session
Deliverable
Sector
Coverage (3)
Company credit fact sheet
Bus. services
Coverage summary
TMT
Teaser and info. memorandum
TMT
Credit (5)
LBO participation review
Industrials
Project-finance review
Energy
Credit committee, green loan
Technology
Table 1: The seventeen deliverables. One judge session grades one group of three to five decks.
Opus / Opus, round 7
Opus / Opus, round 8
Opus / Sonnet, round 5
Campaign means
75.8 / 77.8
75.7 / 76.8
75.0 / 86.8
Mean shift [95% CI]
+2.0 [0.1, 3.8]
+1.1 [ − 0.8, 2.9]
+11.8 [8.7, 14.6]
Wilcoxon p / permutation p
0.033 / 0.063
0.31 / 0.31
< 0.001 / < 0.001
Shift by session (Cov, Cre, Fin, M&A)
+0.7, +5.6, +1.2, − 0.5
+4.7, +0.2, +1.0, − 0.3
+16.0, +6.0, +14.8, +12.0
sd( d ), MDC 95 per deck
3.97, 7.8
4.20, 8.2
6.28, 12.3
Spearman ρ
0.71
0.69
0.43
Table 2: Agreement between two judge passes on identical decks ( n=17 per column). d is the per-deck difference in total, second pass minus first. ICC(A,1) is the absolute-agreement, single-rater coefficient ( McGraw and Wong, 1996 ) ; intervals are bootstrap over decks. The Sonnet column is the development pass of 24 September.
Harness
27B without harness
Opus without harness
Judge
(27B, round 7)
short prompt
strong prompt a
short prompt
strong prompt
Claude Opus
76.9
+16.9 (17/17)
+20.4 (17/17)
− 2.1 (8/17)
− 3.4 (6/17)
Claude Sonnet
75.5
+13.6 (15/17)
+19.1 (16/17)
+1.1 (12/17)
− 1.4 (8/17)
DeepSeek V4.1 Flash
89.1
+12.9 (17/17)
+17.1 (16/17)
+3.2 (12/17)
+1.6 (10/16)
Qwen3.8-27B
93.4
+11.6 (16/17)
+4.3 (14/17)
Gemini 3.8 Flash
91.5
+10.7 (15/17)
+1.9 (10/17)
Table 3: Advantage of the harness deck (27B, round 7) over each no-harness condition, in points of the mean total, with the number of deliverables where the harness deck scores higher. a Graded in mixed sessions that also contain the harness decks; the difference is taken against the harness decks of those same sessions. Opus is graded here in the panel configuration; a blank cell means the judge did not grade that condition, and a count out of 16 means one deck’s score is missing.
Figure 3: Opus campaign mean over the 17 decks by round, with Sonnet passes marked. Open points repeat assessments on identical decks; labels show the observed mean shifts.
Rounds
Δ mean
Up
Down
Wilcoxon p
1 → 2
+6.7
10
7
0.169
2 → 3
+2.8
15
1
0.003
3 → 4
+3.9
14
3
0.003
4 → 5
− 1.8
4
11
0.064
5 → 7
+0.8
9
8
0.585
7 → 8
− 0.1
8
6
0.975
Table 4: Change in the Opus campaign mean between consecutive rounds (first passes), with decks up and down, alongside the exploratory replication scale of 3.2 points (Section 5.2 ; replications in Table 2 ). Round 6 has no Opus pass; values are rounded independently.
Figure 4: Mean score per judge for the 27B with the harness (filled) and for decks written without it (open). Panel (a) compares separate-session means for the harness, the direct 27B with a short prompt, and direct Opus with a short or strong prompt. Panel (b) compares the strong-prompt 27B and harness decks graded together in mixed sessions. Paired differences and deck counts are in Table 3 .
Judge
Round 1
Round 7
Gain
Up
ρcons
MDC
Claude Sonnet
65.4
75.5
+10.1
12
0.76
GPT-6 Luna Pro
71.8
79.8
+8.0
14
0.72
15.4
GLM 5.3 FlashX
80.4
87.8
+7.4
13
0.65
14.5
Claude Opus
63.3
75.8
+12.5
13
0.61
7.8
Gemini 3.8 Flash
80.3
91.5
+11.2
13
0.59
Qwen3.8-27B
84.8
93.4
+8.6
11
0.59
13.5
Table 5: Mean total in rounds 1 and 7, gain, decks that improved, Spearman ρ with the median of the other judges over the 34 deck versions, and per-deck MDC 95 between two passes. Opus values come from the development campaign, and its MDC from identical sessions; the other MDC values come from reshuffled sessions; a blank cell means no second pass (Gemini) or a second pass with an extraneous system instruction (Sonnet, Haiku).
Figure 5: Pairwise Spearman correlation among seven judges, excluding Haiku. Each line follows the same judge pair across round 1, round 7 and their pooled scores. Bars mark the median. Pooling includes the common improvement across rounds. Pairs share judges and decks and are not independent observations.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Criterion (weight)
27B short (7)
27B strong (5)
Opus short (7)
Reader’s problem (15)
+7%
+10%
0%
Synthesis (15)
+8%
+14%
+1%
Recommendation (15)
+6%
+6%
0%
Depth, scenarios (15)
+15%
+13%
− 1%
Consistency (20)
+19%
+32%
+3%
Storyline (10)
+8%
+13%
+1%
Appendix
Table 6: Mean advantage of the harness deck (27B, round 7) over each no-harness condition, per criterion, as a share of the criterion’s weight, averaged over the judges that graded the condition (number of judges in the column header). Short and strong refer to the prompt.
Harness minus no harness
Length
Judge
27B, short
Opus, short
ρ , Opus
Claude Sonnet
+33.6 (17/17) [30.1, 37.2]
+0.4 (10/17) [ − 3.4, 3.6]
0.46
Claude Opus
+29.8 (17/17) [24.8, 34.4]
− 4.7 (4/17) [ − 8.2, − 1.4]
0.54
DeepSeek V4.1 Flash
+23.8 (17/17) [19.3, 28.3]
− 1.1 (6/17) [ − 3.9, 1.1]
0.07
GPT-6 Luna Pro
+23.2 (17/17) [18.1, 28.6]
− 2.1 (8/17) [ − 4.6, 0.4]
0.31
GLM 5.3 FlashX
+20.4 (17/17) [17.0, 24.0]
+0.8 (11/17) [ − 1.5, 2.8]
0.02
Appendix
Table 7: Harness text (27B, round 7) minus text written without the harness in the cleaned text-only pass, out of 95 without visual quality. Cells give 95% bootstrap intervals and the number of deliverables where the harness text scores higher. All three versions of a deliverable were graded in the same session. The last column is the Spearman correlation, over deliverables, between the harness margin over the Opus text and the difference in word count. All totals and criterion scores are available in this pass.
Judge
Access
Effort
Slides
Opus
agent
default
full size
Sonnet
agent
default
full size
Haiku
agent
default
full size
DeepSeek V4.1 Flash
API
low
1600 px
GLM 5.3 FlashX
API
low
1600 px
GPT-6 Luna Pro
API
low
1600 px
Appendix
Table 8: Judge configurations. Effort refers to reasoning; sampling settings are provider defaults. All judges received extracted text. Text-only passes omitted slide images (Section 4 ).
Figure 6: Total of every development Opus judgment by verdict ( n=153 , vertical jitter added).
Figure 7: Per-deck totals from two passes on identical decks: Opus/Opus in rounds 7 and 8 (top, middle) and Opus/Sonnet on 24 September in round 5 (bottom). Open points changed verdict between the passes.
Figure 8: Projected dependability when averaging n judges, conditional on the round-7 variance decomposition. Curves show absolute ( Φ ) and relative ( Eρ2 ) coefficients. The residual includes interaction and error; extrapolation beyond this convenience panel is unvalidated.
Phase
Records
Scores
Development
187
187
Panel development decks
493
491
Short-prompt comparisons
238
238
Strong-prompt comparisons
238
236
First text-only pass
255
254
Cleaned text-only pass
255
255
Appendix
Table 9: Records and recoverable totals by evaluation phase. The strong-prompt phase includes the harness decks regraded in mixed sessions. Missing totals remain missing; none are imputed.