VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks
Organizations: University of Cambridge · Google Cloud AI Research
Abstract
As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repeated sampling yields multiple rollouts that can contain complementary correct claims, but we need a reliable verification mechanism to determine which claims to trust. We first find that disagreement often exposes correct alternatives, while consensus can conceal errors. These observations motivate VeriHarness, which turns the underlying LLM a generator uses into an agentic verifier by giving it a workspace, evidence tools, and reusable verification skills. A disagreement resolver checks competing claims against environmental evidence, while a consensus challenger tests shared claims and searches for omitted requirements. Their findings guide the selection and revision of the final artifact. Across five long-horizon workspace benchmarks and two frontier models, VeriHarness achieves the highest selection scores among the evaluated baselines. Evidence-backed revision further improves average performance, bringing gains over a single rollout to 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. We further show that verification skills can self-improve from failure feedback, demonstrating VeriHarness as a novel and critical approach for scaling long-horizon agentic verification. We release the full pool of approximately 26,000 rollouts across all five benchmarks and both models, produced at a cost of over $100,000, to support future research on agentic verification.
Figures & tables
| Benchmark | Tasks | Domain | Artifact | Grader |
|---|---|---|---|---|
| APEX-Agents (v1.0) | 480 | banking, law, consulting | memos, models, answers | expert rubric |
| Workspace-Bench Lite | 100 | file-heavy workspaces | documents and files | rubric judge |
| WorkBuddy Bench | 200 | code, office, web | patches and office files | composite verifier |
| SpreadsheetBench 2 | 321 | business spreadsheets | workbooks | recompute, compare |
| JobBench | 65 | 35 white-collar occupations | multi-file artifacts | rubric judge |
| Method | Avg | APEX | WSB | WorkBuddy | SB-2 | JobBench |
|---|---|---|---|---|---|---|
| Gemini 3.5 Flash | ||||||
| Single rollout | 47.2 | 48.1 | 56.5 | 73.2 | 27.3 | 30.9 |
| Majority voting | 47.5 | 47.3 | 57.0 | 74.0 | 27.0 | 32.4 |
| Best-of- with judge | 47.3 | 47.2 | 57.0 | 73.6 | 27.7 | 30.9 |
| Pairwise tournament | 49.1 | 49.2 | 59.6 | 74.8 | 30.4 | 31.4 |
| LLM-as-a-Verifier | 49.5 | 49.8 | 60.4 | 75.3 | 30.2 | 31.9 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Source | How it helps verification |
|---|---|
| Exposed alternatives | Multiple rollouts make competing values and interpretations explicit. The verifier can investigate a concrete discrepancy that an individual generation attempt never encountered. |
| Focused checking | Given completed artifacts, the verifier can devote its attention to testing particular claims. Separate resolver and challenger contexts reserve attention for both disputed and shared claims. |
| Environmental evidence | A targeted source lookup or recomputation can reveal evidence omitted from the rollouts. The generator has the same tools; verification directs their use toward a specific claim whose validity is now in question. |
| Accumulated skills | Human-authored checks and procedures learned from development failures suggest what to examine, including requirements that every candidate omitted. This experience can guide later verification without changing model weights. |
| Method | Avg | APEX | WSB | WorkBuddy | SB-2 | JobBench |
|---|---|---|---|---|---|---|
| Gemini 3.5 Flash | ||||||
| Single rollout | 47.2 | 48.1 | 56.5 | 73.2 | 27.3 | 30.9 |
| Resolver only | 51.5 | 53.1 | 63.2 | 77.0 | 30.0 | 34.0 |
| Challenger only | 50.0 | 50.9 | 60.3 | 75.2 | 29.7 | 33.7 |
| VeriHarness without skills | 51.3 | 50.9 | 63.5 | 77.1 | 30.6 | 34.4 |
| VeriHarness in a single context | 52.6 | 53.9 | 64.5 | 77.9 | 31.3 | 35.3 |
| Part | What it provides | Examples |
|---|---|---|
| Workspace and tools | Access to the objects under check | Read a file, recompute a table, trace a source |
| Verification protocol | Division of checking labor | Resolver, challenger, adjudication |
| Verification skills | Reusable knowledge of what to check | Period, unit, source version, required items |
| Skill | Deliverable | What it supplies | Scripts |
|---|---|---|---|
| evidence-xlsx | any | formulas beside cached values, workbook diffs, rendered sheets | xlsx_dump , xlsx_diff , xlsx_gaps , xlsx_forks , xlsx_render |
| evidence-pdf | any | tables with their columns, words with coordinates, rendered pages | pdf_render , pdf_tables , pdf_words |
| evidence-docx | any | text in reading order, tables, tracked changes and comments | docx_changes , docx_text |
| evidence-pptx | any | shapes in slide order, tables, charts, notes, rendered slides | pptx_render , pptx_text |
| evidence-bundle | file bundle | inventory of what each candidate delivered and states | bundle_inventory |
| evidence-patch | patch or page | build and test of every candidate patch; headless-browser probe of a delivered page | patchlab , pageprobe |
| Relation between the tested claim and the graded item | Share |
|---|---|
| The tested claim is an input or intermediate quantity; the graded quantity lies downstream | 70% |
| Same quantity, different definition, period, or source version | 14% |
| The challenger confirmed the wrong value of the graded quantity | 12% |
| Omission or mapping noise | 4% |