AgentTime: Can Agents Estimate and Control Their Own Runtime?
Organizations: MATS · ELLIS Institute Tübingen, Max Planck Institute for Intelligent Systems, Tübingen AI Center
Abstract
An essential control of AI agents is their ability to manage runtime. This ability requires a sense of time-awareness, to predict and estimate wall-clock time and to control their own actions. Prior work has focused on time-awareness, but duration-following and control in native agent harnesses remain unexplored. We present AgentTime, a benchmark for testing whether agents can work for a requested duration, predict their runtime, and estimate elapsed time afterward. It comprises 222 tasks from 18 sources spanning coding, computer use, agentic work, and automated research. Duration-following experiments append a single instruction specifying how long to work, with requests ranging from about a minute to multiple days. Accuracy on these instructions varies substantially: Fable 5.1 in Claude Code deviates from requested runtimes by a typical factor of 2.9, compared with only 1.2 for GPT-6 Astra in Codex. However, matching the requested runtime does not, by itself, establish continued work on the task. Among 158 reviewed Astra runs with classifiable transcripts, 14 explicitly slept after appearing to finish. In forecasting experiments, predictions tend to overestimate natural runtimes. In retrospective experiments, removing temporal information more than doubles deviation for Sol and Astra and nearly doubles it for Fable. An agent's ability to complete a task does not guarantee that it can control its own time or work for the whole requested duration. For agents to run reliably, safely, and autonomously over long horizons, we require the evaluation of both.
Figures & tables
| Paper | Acts under budget | Measures time/budget | Predicts own time | Multi-hour wall clock | Native CLI harness | Duration following |
|---|---|---|---|---|---|---|
| Can LLMs Perceive Time? ( Garikaparthi, 2026 ) | ✓ | ✓ | ||||
| Temporally Blind (TicToc) ( Cheng et al., 2026 ) | ||||||
| Discrete Minds ( Wang et al., 2025 ) | ✓ | ✓ | ||||
| Real-Time Deadlines ( Sehgal et al., 2026 ) | ✓ | ✓ | ||||
| BATS ( Liu et al., 2026b ) | ✓ | ✓ | ||||
| BAGEN ( Lin et al., 2026 ) | ✓ | ✓ |
| Benchmark | Tasks | Requests | Benchmark | Tasks | Requests |
| GPQA Diamond ( Rein et al., 2024 ) | 28 | 1.25–20 min | Humanity’s Last Exam ( Phan et al., 2025 ) | 24 | 1.25–20 min |
| AppWorld ( Trivedi et al., 2024 ) | 18 | 1.5–40 min | AssistantBench ( Yoran et al., 2024 ) | 18 | 1.25–20 min |
| OSWorld 2.0 ( Yuan et al., 2026 ) | 16 | 15 min–4.2 h | Terminal-Bench 4.0 ( Merrill et al., 2026 ) | 16 | 4 min–8.3 h |
| ProgramBench ( Yang et al., 2026 ) | 14 | 8 min–8.3 h | Agents’ Last Exam ( Sun et al., 2026 ) | 12 | 4 min–2.5 h |
| CORE-Bench v1.1 ( Nadgir et al., 2026 ) | 12 | 2.5 min–3.3 h | DeepSWE v1.1 ( Huang et al., 2026 ) | 12 | 4 min–2.5 h |
| TUA-Bench ( Chen et al., 2026a ) | 12 | 1.25 min–1.7 h | PPTArena ( Ofengenden et al., 2026 ) | 10 | 1.5–25 min |
| Agent (harness) | Runs | On time | Early | Late | Deviation | Slope (ideal 1) |
| Fable 5.1 (Claude Code) | 659 | 4% | 55% | 41% | 2.86 [2.68, 2.99] | 0.59 [0.54, 0.64] |
| GPT-5.6 Sol (Codex) | 666 | 39% | 13% | 48% | 1.77 [1.59, 1.95] | 0.71 [0.66, 0.75] |
| GPT-6 Astra (Codex) | 666 | 63% | 3% | 34% | 1.18 [1.11, 1.25] | 0.94 [0.92, 0.96] |
| Harness swap (Appendix B ): 18 questions and 16 agentic tasks, within-task slope | ||||||
| Fable 5.1 (Claude Code) | 98 | 5% | 50% | 45% | 3.14 [2.73, 3.60] | 0.07 [ 0.02, 0.17] |
| Fable 5.1 (Codex) † | 98 | 1% | 34% | 65% | 2.99 [2.64, 3.41] | 0.23 [0.12, 0.34] |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Benchmark | Native metric | Success rule | Percentage |
|---|---|---|---|
| GPQA Diamond | Accuracy: the answer letter matches the key (0/1) | Correct letter | 0 or 100 |
| Humanity’s Last Exam | Accuracy (0/1) from HLE’s own judge, on text-only multiple-choice items | The judge marks the answer correct | 0 or 100 |
| AppWorld | Task Goal Completion (0/1) | Every final-state test passes | 0 or 100 |
| AssistantBench | Answer accuracy with partial credit, 0 to 1 | Exact match with the reference answer | score |
| OSWorld 2.0 | Checkpoint-weighted partial score, 0 to 1 | Binary completion: a score of 1.00, every checkpoint met | score |
| Terminal-Bench 4.0 | Verifier reward (0/1) | The task’s tests pass | 0 or 100 |
| Benchmark | Unit | Fable 5.1 | GPT-5.6 Sol | GPT-6 Astra |
|---|---|---|---|---|
| GPQA Diamond | % | 85.7 (84/84) | 94.0 (84/84) | 96.4 (84/84) |
| Humanity’s Last Exam | % | 87.5 (72/72) | 75.0 (72/72) | 77.8 (72/72) |
| AppWorld | % | 75.9 (54/54) | 70.4 (54/54) | 61.1 (54/54) |
| AssistantBench | % | 4.8 (54/54) | 10.6 (54/54) | 3.3 (54/54) |
| OSWorld 2.0 | % | 67.8 (48/48) | 63.3 (48/48) | 73.7 (48/48) |
| Terminal-Bench 4.0 | % | 64.6 (48/48) | 50.0 (48/48) | 58.3 (48/48) |
| Model | Harness | Runs | On time | Early | Late | Deviation | Slope (ideal 1) |
| 18 GPQA and HLE questions, each asked for 1.25, 5 and 20 minutes | |||||||
| Fable 5.1 | Claude Code | 54 | 1 | 31 | 22 | 3.51 | 0.03 [ 0.11, 0.19] |
| Codex † | 54 | 1 | 24 | 29 | 2.94 | 0.19 [0.03, 0.34] | |
| GPT-5.6 Sol | Codex | 54 | 23 | 1 | 30 | 1.28 | 0.78 [0.59, 0.92] |
| GPT-6 Astra | Codex | 54 | 35 | 0 | 19 | 1.06 | 0.96 [0.94, 0.96] |
| Claude Code † | 54 | 12 | 0 | 42 | 1.26 | 0.85 [0.80, 0.89] | |
| Fable 5.1 | GPT-5.6 Sol | GPT-6 Astra | ||||
| Question | Asked | Claude Code | Codex † | Codex | Codex | Claude Code † |
| GPQA, picked where Fable in Claude Code had missed a request | ||||||
| Question 1 | 1.25 | 1.5 1.23 | 1.6 1.24 | 1.4 1.13 | 1.4 1.10 | 1.5 1.21 |
| 5 | 0.54 0.11 | 5.6 1.12 | 5.2 1.05 | 5.2 1.04 | 5.3 1.06 | |
| 20 | 0.53 0.03 | 0.45 0.02 | 20.4 1.02 | 20.2 1.01 | 20.2 1.01 | |
| Question 2 | 1.25 | 1.5 1.20 | 1.1 0.85 | 1.5 1.24 | 1.4 1.10 | 1.5 1.21 |
| Fable 5.1 | GPT-5.6 Sol | GPT-6 Astra | ||||
| Task | Asked | Claude Code | Codex † | Codex | Codex | Claude Code † |
| Terminal-Bench | ||||||
| cad-model | 5 | 13.4 2.69 | 51.2 10 ∗ | 10.0 2.00 | 5.2 1.04 | 5.5 1.10 ∗ |
| 20 | 15.6 0.78 | 66.8 3.34 | 20.3 1.02 | 20.6 1.03 | 21.3 1.06 ∗ | |
| 80 | 20.4 0.26 | 114.7 1.43 | 80.8 1.01 | 80.8 1.01 | 81.8 1.02 ∗ | |
| fin-saccr-rwa | 4 | 12.1 3.02 | 21.2 5.30 | 21.8 5.44 | 8.4 2.09 § | 13.2 3.29 ∗ |
| Context | Tools | |
|---|---|---|
| Oracle | Session fork | Ordinary + clock |
| Native | Session fork | Ordinary |
| Context-only | Session fork | All disabled |
| Replay | Reconstructed, API | None |
| Scrubbed | Reconstructed, no timestamps, API | None |
| P1 – Text | P2 – Docs | P3 – Probe |
|---|---|---|
| Task prompt | P1 + documentation if exists | P2 + sandbox with tools |