Humans naturally possess the ability to work persistently toward long-term goals. Given a challenging task, humans can continuously work for months or even years to accomplish a specific objective. In this paper, we propose Marathoner, an autonomous agentic model possessing the ability of ultra-long-horizon execution. Specifically, we propose a comprehensive post-training pipeline to instill this critical capability into base model. For Ultra-Long-Horizon Task Synthesis, we leverage major release PRs containing 1000+ lines of new code from diverse GitHub repositories as the primary source for synthesizing challenging task-level data. Additionally, we introduce Multi-Task Chaining, which chains multiple generated tasks into a single more challenging task, enabling the synthesis of tasks with frontier-level difficulty. For rejection sampling finetuning, we combine strong teacher model with diverse harnesses to generate trajectories on our synthesized tasks and conduct supervised finetuning on base model with rejection sampled trajectories. For reinforcement learning, cold-started model performs real-world execution through harnesses in independent sandboxes during rollout process, effectively facilitating the acquisition of genuine ultra-long-horizon execution capability. We further propose a novel reward strategy, Later Stage Bonus Reward, which explicitly encourages model to perform meaningful maneuvers during later stages of execution. Through extensive evaluation on 5 benchmarks containing ultra-long-horizon tasks, Marathoner achieves consistent and substantial performance improvements over base model and even surpasses performance of strong proprietary model. Further analysis shows that Marathoner can consistently work for 10+ hours and conduct 1000+ tool calls on highly challenging tasks.
Figure 2: The proposed post-training pipeline for building ultra-long-horizon execution ability.
Models
Harness
FrontierSWE
NL2Repo
SWE-Marathon
Terminal Bench 2.0
SWE-Bench Verified
Proprietary Models with Harness
GLM-5.2
Claude Code
71.3
47.2
13.6
82.5
87.5
Qwen-3.7-Max
Claude Code
58.2
44.6
10.3
65.8
79.2
Kimi K3
Claude Code
79.8
48.9
35.9
86.4
86.2
GPT-6-Astra
Codex
89.1
78.2
48.7
87.4
90.7
Claude Fable 5.1
Claude Code
86.2
73.2
42.8
85.9
91.8
Table 1: We conduct evaluations across a wide range of benchmarks to enable comprehensive assessment and analysis of the ultra-long-horizon execution capability of our models.
Benchmark
Avg. Execution Time
Avg. Steps
Avg. Tool Calls
FrontierSWE
3.56h
426.2
648.3
Terminal Bench 2.0
0.58h
104.8
239.4
Table 2: Statistics of the Marathoner execution process on FrontierSWE and Terminal Bench 2.0.
Setting
FrontierSWE
Terminal Bench 2.0
Marathoner (vanilla)
22.7
51.3
Marathoner (vanilla) w. Multi-Task Chaining
24.9
54.8
Marathoner (vanilla) w. Diverse-Harness Trajectory Generation
24.1
53.7
Marathoner (vanilla) w. Later Stage Bonus Reward
24.6
54.4
Table 3: Ablation of our key designs.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
# Direct Synthesized Task
# Multi-Task Chaining Task
# Chained Task
FrontierSWE
Terminal Bench 2.0
5000
3000
2
23.8
54.1
5000
3000
3
24.4
54.7
5000
3000
5
26.4
57.2
5000
3000
7
25.2
56.3
5000
3000
9
24.8
55.9
Appendix
Table 4: Ablation of the number of tasks chained in Multi-Task Chaining. # Direct Synthesized Task denotes the number of tasks from direct data synthesis during RL. # Multi-Task Chaining Task denotes the number of tasks from Multi-Task Chaining during RL. # Chained Task denotes the number of atomic tasks combined to construct a larger task in Multi-Task Chaining.
Stage
Bonus Reward
FrontierSWE
Terminal Bench 2.0
50%-100%
0.2
24.7
55.8
50%-100%
0.5
26.4
57.2
50%-100%
0.8
24.1
55.3
30%-100%
0.5
25.2
56.3
70%-100%
0.5
24.9
56.1
Appendix
Table 5: Ablation of the design choices for the Later Stage Bonus Reward. We analyze the effects of the bonus stage and bonus reward value on agent performance. We find that assigning a bonus reward of 0.5 to valuable actions performed during the 50%–100% stage yields the best results.
Model
Harness
FrontierSWE
NL2Repo
SWE-Marathon
Terminal Bench 2.0
SWE-Bench Verified
Qwen3.5-9B
Claude Code
10.2
17.9
0
27.3
43.8
Marathoner-9B-RFT
Claude Code
18.2
23.8
2.9
49.7
58.4
Marathoner-9B
Claude Code
26.4
34.7
8.2
57.2
77.5
Appendix
Table 6: Comparison of performance between raw base model, RFT model, and final model after RL.
Scientific and engineering progress is fundamentally a long-horizon iterative process: proposing changes, running experiments, measuring outcomes, and continuously refining artifacts. Yet existing benchmarks for frontier models primarily evaluate either single-turn responses or short-horizon agent trajectories, failing to capture the challenges of sustained iterative improvement over extended time horizons. To address this gap, we introduce AutoLab, a new benchmark for ultra long-horizon closed-loop optimization. AutoLab consists of 36 realistic, expert-curated tasks spanning four diverse domains: system optimization, puzzle & challenge, model development, and CUDA kernel optimization. Each task begins with a correct but deliberately suboptimal baseline and challenges agents to improve it within a strict wall-clock budget. Evaluating 17 state-of-the-art models reveals the dominant predictor of success is not the quality of an agent's initial attempt, but its persistence in repeatedly benchmarking, editing, and incorporating empirical feedback. While claude-opus-4.6 exhibits strong long-horizon optimization capabilities, most frontier models, including several proprietary ones, either terminate prematurely or exhaust their budgets with minimal progress. These results underscore the importance of time awareness and persistent iteration in autonomous agents. We open-source the full benchmark, evaluation harness, and task artifacts, to accelerate research toward truly capable long-horizon agents.
Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interdependent steps. However, existing agent harnesses maintain task execution, task state, and completion assessment within a growing context, making the state difficult to track and allowing incorrect self-assessments to propagate into later decisions. We reformulate long-horizon execution as a task-state management problem and propose LongHorizon-Harness, which maintains the task state explicitly outside execution and updates it only with facts independently verified from the environment. Its Manage-Execute-Audit(MEA) loop uses a manager to maintain the task state and determine the next subtask, a fresh-context executor to perform it, and a read-only auditor to verify the resulting environment state before the next round. A lightweight AgentAdapter supports interchangeable model and harness backends without modifying their native agent loops. LongHorizon-Harness improves Qwen3.7-Plus from 51.8% to 80.7% on WeaveBench, from 69.7% to 77.2% on Terminal-Bench2.1, and from 2.8% to 8.3% on OSWorld2.0. It also raises Claude Opus4.7 from 20.0% to 34.3% on an OSWorld2.0 subset, demonstrating consistent gains across models, harnesses, and interaction domains.
Long-horizon household tasks demand robust high-level planning and sustained reasoning capabilities, which are largely overlooked by existing embodied AI benchmarks that emphasize short-horizon navigation or manipulation and rely on fixed task categories. We introduce LongAct, a benchmark designed to evaluate planning-level autonomy in long-horizon household tasks specified through free-form instructions. By abstracting away embodiment-specific low-level control, LongAct isolates high-level cognitive capabilities such as instruction understanding, dependency management, memory maintenance, and adaptive planning. We further propose HoloMind, a VLM-driven agent with a DAG-based long-horizon hierarchical planner, a Multimodal Spatial Memory for persistent world modeling, an Episodic Memory for experience reuse, and a global Critic for reflective supervision. Experiments with GPT-5 and Qwen3-VL models show that HoloMind substantially improves long-horizon performance while reducing reliance on model scale. Even top models achieve only 59% goal completion and 16% full-task success, underscoring the difficulty of LongAct and the need for stronger long-horizon planning in embodied agents.
Zilin Zhu, Longteng Guo, Yanghong Mei +5
Institute of Automation, Chinese Academy of Sciences, Beijing, China · Zhongguancun Academy, Beijing, China · University of Chinese Academy of Sciences, Beijing, China