DAYJOB: A Benchmark for Long-Horizon Professional Work
Organizations: Surge AI
Abstract
Professional work often starts with a brief request that leaves the professional to work out what is needed, which documents matter, and whether the request's premise holds. We introduce DAYJOB, a benchmark of 130 tasks built by professionals in healthcare (50) and finance (80). The tasks are estimated to take a professional 13.6 hours on average in healthcare and 16.6 in finance. Each task is a containerized Harbor environment with an expert rubric of binary criteria (median 47.5 and 57.5 per task) that an agentic judge applies to the delivered files, and an attempt passes only if it meets every criterion. Across 30 model configurations from 13 developers, the strongest, Claude Opus 5.5, passes 24.7% of healthcare and 23.9% of finance attempts, and the median configuration passes 0.6% and 2.5%. In case studies, agents accept premises that the record contradicts and carry wrong inputs through otherwise consistent analyses. We release all healthcare tasks, 50 of the 80 finance tasks, the evaluation harness, and the leaderboard.
Figures & tables
| Request. “Caldervane Emerging Markets Fund is a new client’s first fund with us and we are delivering our first reports for it today for COB June 30th. The client wants the initial reports as soon as they are available, so they can review these BAU deliverables in detail. Go through the attached pack and let me know if there are any issues stopping me sending the reports to the client as they are. Write it up as a short .md memo.” |
|---|
| Criteria |
| Identifies that the Bloomberg screen quotes the Nerovexa price of 261.200 in South African cents rather than rand. |
| States that the Nerovexa position is overstated by about 25.5m–$25.9m). |
| Recommends not sending the reports until the Nerovexa position is corrected; a hold that lifts once the quotation basis is confirmed does not satisfy the criterion. |
| Identifies that the Kingdom of Saudi Arabia 4.500% 2046 bond was priced from BVAL when ICE supplied a current evaluation. |
| Does not call the Dravena position quantity or its 1,242.000 CZK price incorrect, and does not treat the stock split as a reason to withhold the reports. |
| DAYJOB: Healthcare | DAYJOB: Finance | GDPval gold subset | |
|---|---|---|---|
| Tasks | 50 | 80 | 220 |
| Request length, words (mean / median) | 48.9 / 42.5 | 80.7 / 62.5 | 337.2 / 307 |
| Input files per task (mean / median) | 19.8 / 16.5 | 25.7 / 20.5 | 1.19 / 1 |
| Input files per task (range) | 8–50 | 8–122 | 0–17 |
| Rubric criteria per task (mean / median) | 55.7 / 47.5 | 66.1 / 57.5 | 47.5 / 47 |
| Estimated hours per task (mean) ‡ | 13.6 | 16.6 § | 6.7 |
| DAYJOB: Healthcare | DAYJOB: Finance | ||||
|---|---|---|---|---|---|
| Configuration | Developer | Strict | Rank | Strict | Rank |
| Claude Opus 5.5 (adaptive/max) | Anthropic | 24.7 | 1 | 23.9 | 1 |
| GPT-6 Astra (max) | OpenAI | 11.6 | 2 | 21.5 | 2 |
| Claude Fable 5.1 (adaptive/max) | Anthropic | 9.6 | 3 | 19.8 | 3 |
| Grok 4.7 (xhigh) | xAI | 8.4 | 4 | 14.5 | 5 |
| Muse Spark 1.3 (max) | Meta | 7.8 | 6 | 14.8 | 4 |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.