cs.AISep 30, 2026

Better Deck or Different Judge? Evaluating Agentic Harness Gains in Corporate and Investment Banking

Authors: Ludovic Gibert, Matis Despujols, Andre-Louis Rochet

Organizations: AIData2Action · TW3 Partners

Abstract

Corporate and investment banking teams use presentations to support credit decisions and advise clients on financing and transactions. Producing these decks requires reconciling financial data, tracing sources and turning analysis into a recommendation. We retrospectively study the development of an agentic harness combining a 27B language model, financial calculations, narrative templates and validation checks. LLM judges guide engineering changes and assess the resulting decks, raising the question of whether higher scores reflect better documents or changes in grading. In shared-session text-only grading with template markers removed, five judges score the complete system 20.4 to 33.6 points out of 95 above the same model generating directly from a short prompt. Every judge scores the system higher on all seventeen development deliverables. Margins against direct Opus generation from a short prompt range from -4.7 to +0.8 points. Judges agree on broad progress across development rounds but agree less on final-deck rankings than on pooled scores. Repeated grading also shifts scores on unchanged decks, making small improvements difficult to distinguish from judge variability.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places

    Sep 24, 2026Delip Rao, Chris Callison-BurchLarge Language Model Judges

  2. Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System

    Jun 23, 2026Tian Zheng, Kai-Tai HsuGradingData Science Agents

  3. Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why

    May 25, 2026Delip Rao, Chris Callison-BurchLarge Language Model JudgesJudgement