Better Deck or Different Judge? Evaluating Agentic Harness Gains in Corporate and Investment Banking
Organizations: AIData2Action · TW3 Partners
Abstract
Corporate and investment banking teams use presentations to support credit decisions and advise clients on financing and transactions. Producing these decks requires reconciling financial data, tracing sources and turning analysis into a recommendation. We retrospectively study the development of an agentic harness combining a 27B language model, financial calculations, narrative templates and validation checks. LLM judges guide engineering changes and assess the resulting decks, raising the question of whether higher scores reflect better documents or changes in grading. In shared-session text-only grading with template markers removed, five judges score the complete system 20.4 to 33.6 points out of 95 above the same model generating directly from a short prompt. Every judge scores the system higher on all seventeen development deliverables. Margins against direct Opus generation from a short prompt range from -4.7 to +0.8 points. Judges agree on broad progress across development rounds but agree less on final-deck rankings than on pooled scores. Repeated grading also shifts scores on unchanged decks, making small improvements difficult to distinguish from judge variability.
Figures & tables
| Session | Deliverable | Sector |
|---|---|---|
| Coverage (3) | Company credit fact sheet | Bus. services |
| Coverage summary | TMT | |
| Teaser and info. memorandum | TMT | |
| Credit (5) | LBO participation review | Industrials |
| Project-finance review | Energy | |
| Credit committee, green loan | Technology |
| Opus / Opus, round 7 | Opus / Opus, round 8 | Opus / Sonnet, round 5 | |
| Campaign means | 75.8 / 77.8 | 75.7 / 76.8 | 75.0 / 86.8 |
| Mean shift [95% CI] | +2.0 [0.1, 3.8] | +1.1 [ 0.8, 2.9] | +11.8 [8.7, 14.6] |
| Wilcoxon / permutation | 0.033 / 0.063 | 0.31 / 0.31 | 0.001 / 0.001 |
| Shift by session (Cov, Cre, Fin, M&A) | +0.7, +5.6, +1.2, 0.5 | +4.7, +0.2, +1.0, 0.3 | +16.0, +6.0, +14.8, +12.0 |
| sd( ), MDC 95 per deck | 3.97, 7.8 | 4.20, 8.2 | 6.28, 12.3 |
| Spearman | 0.71 | 0.69 | 0.43 |
| Harness | 27B without harness | Opus without harness | |||
|---|---|---|---|---|---|
| Judge | (27B, round 7) | short prompt | strong prompt a | short prompt | strong prompt |
| Claude Opus | 76.9 | +16.9 (17/17) | +20.4 (17/17) | 2.1 (8/17) | 3.4 (6/17) |
| Claude Sonnet | 75.5 | +13.6 (15/17) | +19.1 (16/17) | +1.1 (12/17) | 1.4 (8/17) |
| DeepSeek V4.1 Flash | 89.1 | +12.9 (17/17) | +17.1 (16/17) | +3.2 (12/17) | +1.6 (10/16) |
| Qwen3.8-27B | 93.4 | +11.6 (16/17) | +4.3 (14/17) | ||
| Gemini 3.8 Flash | 91.5 | +10.7 (15/17) | +1.9 (10/17) | ||
| Rounds | mean | Up | Down | Wilcoxon |
|---|---|---|---|---|
| 1 2 | +6.7 | 10 | 7 | 0.169 |
| 2 3 | +2.8 | 15 | 1 | 0.003 |
| 3 4 | +3.9 | 14 | 3 | 0.003 |
| 4 5 | 1.8 | 4 | 11 | 0.064 |
| 5 7 | +0.8 | 9 | 8 | 0.585 |
| 7 8 | 0.1 | 8 | 6 | 0.975 |
| Judge | Round 1 | Round 7 | Gain | Up | MDC | |
|---|---|---|---|---|---|---|
| Claude Sonnet | 65.4 | 75.5 | +10.1 | 12 | 0.76 | |
| GPT-6 Luna Pro | 71.8 | 79.8 | +8.0 | 14 | 0.72 | 15.4 |
| GLM 5.3 FlashX | 80.4 | 87.8 | +7.4 | 13 | 0.65 | 14.5 |
| Claude Opus | 63.3 | 75.8 | +12.5 | 13 | 0.61 | 7.8 |
| Gemini 3.8 Flash | 80.3 | 91.5 | +11.2 | 13 | 0.59 | |
| Qwen3.8-27B | 84.8 | 93.4 | +8.6 | 11 | 0.59 | 13.5 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Criterion (weight) | 27B short (7) | 27B strong (5) | Opus short (7) |
|---|---|---|---|
| Reader’s problem (15) | +7% | +10% | 0% |
| Synthesis (15) | +8% | +14% | +1% |
| Recommendation (15) | +6% | +6% | 0% |
| Depth, scenarios (15) | +15% | +13% | 1% |
| Consistency (20) | +19% | +32% | +3% |
| Storyline (10) | +8% | +13% | +1% |
| Harness minus no harness | Length | ||
|---|---|---|---|
| Judge | 27B, short | Opus, short | , Opus |
| Claude Sonnet | +33.6 (17/17) [30.1, 37.2] | +0.4 (10/17) [ 3.4, 3.6] | 0.46 |
| Claude Opus | +29.8 (17/17) [24.8, 34.4] | 4.7 (4/17) [ 8.2, 1.4] | 0.54 |
| DeepSeek V4.1 Flash | +23.8 (17/17) [19.3, 28.3] | 1.1 (6/17) [ 3.9, 1.1] | 0.07 |
| GPT-6 Luna Pro | +23.2 (17/17) [18.1, 28.6] | 2.1 (8/17) [ 4.6, 0.4] | 0.31 |
| GLM 5.3 FlashX | +20.4 (17/17) [17.0, 24.0] | +0.8 (11/17) [ 1.5, 2.8] | 0.02 |
| Judge | Access | Effort | Slides |
|---|---|---|---|
| Opus | agent | default | full size |
| Sonnet | agent | default | full size |
| Haiku | agent | default | full size |
| DeepSeek V4.1 Flash | API | low | 1600 px |
| GLM 5.3 FlashX | API | low | 1600 px |
| GPT-6 Luna Pro | API | low | 1600 px |
| Phase | Records | Scores |
| Development | 187 | 187 |
| Panel development decks | 493 | 491 |
| Short-prompt comparisons | 238 | 238 |
| Strong-prompt comparisons | 238 | 236 |
| First text-only pass | 255 | 254 |
| Cleaned text-only pass | 255 | 255 |