Adapting general-purpose large language models to specific tasks requires substantial human effort in designing data and training strategies. Sustaining improvement is especially challenging because model updates change the error distribution, requiring strategies to be continually refined. We introduce ImproveAnyTask, an autonomous post-training harness that improves task performance under a limited compute budget. Drawing inspiration from gradient-based parameter optimization, the harness organizes adaptation into error attribution, update-direction selection, and executable model updates. It combines metric-level and case-level analysis to identify a focal problem, then investigates research-backed strategies and compares their reported gains and reproduction difficulty. The selected strategy is translated into training data and a training configuration, with small-scale execution checks preceding full post-training. Subsequent evaluation guides model selection and further adaptation, while validated strategies and scripts are retained for reuse. Across 11 tasks, ImproveAnyTask achieves mean gains of 18.29 and 11.97 percentage points on the Base and Instruct models, respectively, with a maximum gain of 41.96 points, under a 24-hour budget with resources equivalent to eight H20 GPUs.
Figures & tables
Figure 1: Task-specific improvement of Qwen3.5-4B-Base. Scores across 11 tasks in five domains under resources equivalent to eight H20 GPUs for up to 24 hours or three rounds. ImproveAnyTask yields a mean gain of 18.29 points and a maximum of 41.96 points on LiveCodeBench v6. BFCL excludes Web Search and renormalizes the remaining weights.
Figure 2: Overview of ImproveAnyTask. The harness treats task adaptation as iterative optimization: evaluation and dual-level error attribution identify a focal problem; research-grounded strategy exploration selects an update direction; data preparation and post-training produce a candidate model; and re-evaluation determines the next checkpoint and records reusable assets.
Qwen3.5-4B-Base
Qwen3.5-4B-Instruct
Benchmark
Original
Codex
Claude Code
Ours
Original
Codex
Claude Code
Ours
Knowledge & STEM
MMLU-Pro
52.41
53.53 ( ↑ 1.12)
52.27 ( ↓ 0.13)
59.85 ( ↑ 7.44)
58.66
53.41 ( ↓ 5.25)
57.27 ( ↓ 1.38)
69.33 ( ↑ 10.66)
SuperGPQA
32.10
33.97 ( ↑ 1.86)
38.79 ( ↑ 6.68)
47.58 ( ↑ 15.48)
49.78
49.54 ( ↓ 0.23)
47.04 ( ↓ 2.73)
47.77 ( ↓ 2.00)
GPQA-Diamond
47.02
45.23 ( ↓ 1.78)
51.78 ( ↑ 4.76)
69.64 ( ↑ 22.61)
75.59
73.21 ( ↓ 2.38)
78.57 ( ↑ 2.97)
80.35 ( ↑ 4.76)
Mathematical & Verifiable Reasoning
Table 1: Mean results over two independent optimization runs across 11 tasks. Each run selects its post-trained candidate by validation score, excluding Original. BFCL uses the renormalized no-Web metric. All values are percentages.
Figure 3: IFEval optimization over 100 hours. The dashed line marks 24 hours. Checkpoints are equally spaced; changes compare adjacent displayed checkpoints. C5 resumes from C3.
Benchmark
Original
w/o EA
w/o UD
w/o ER
Full Harness
MMLU-Pro
58.66
61.20 ( ↑ 2.53)
60.30 ( ↑ 1.63)
52.51 ( ↓ 6.15)
69.33 ( ↑ 10.66)
MATH-500
51.29
56.00 ( ↑ 4.70)
56.70 ( ↑ 5.41)
54.11 ( ↑ 2.82)
63.52 ( ↑ 12.23)
LiveCodeBench v6
11.49
24.77 ( ↑ 13.28)
33.03 ( ↑ 21.54)
28.57 ( ↑ 17.07)
40.62 ( ↑ 29.12)
TAU2-Bench
36.68
39.21 ( ↑ 2.52)
41.36 ( ↑ 4.68)
38.09 ( ↑ 1.41)
44.94 ( ↑ 8.25)
IFEval
82.78
82.35 ( ↓ 0.43)
83.22 ( ↑ 0.43)
80.17 ( ↓ 2.61)
83.66 ( ↑ 0.87)
Overall Average
48.18
52.70 ( ↑ 4.52)
54.92 ( ↑ 6.74)
50.69 ( ↑ 2.50)
60.41 ( ↑ 12.23)
Table 2: Ablation results for Qwen3.5-4B-Instruct, averaged over two independent optimization runs. Each component is replaced with a basic alternative while preserving the complete optimization loop. All values are percentages.
Figure 4: Optimization trajectories for Qwen3.5-4B-Instruct on MMLU-Pro and MATH-500. Initial and final scores match Table 2 . Curves connect measured checkpoints for visualization.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Validation
Test
Primary metric
MMLU-Pro
1,805
10,227
Accuracy
SuperGPQA
3,980
22,549
Accuracy
GPQA-Diamond
30
168
Accuracy
MATH-500
75
425
Accuracy
AIME 2025
5
25
Accuracy
IFEval
82
459
Strict prompt-level accuracy
Appendix
Table 3: Validation and test sizes used in all comparisons. TAU2 is split by domain and BFCL by category.
Ckpt.
Parent
Strategy and data construction
Examples
C1
Initial Base
Persona-conditioned constraint tuning. Use released instruction–response pairs covering keyword, formatting, length, and capitalization constraints. Preserve message structure, validate formatting, deduplicate, and audit exact overlap with evaluation queries.
29,900
C2
C1
Rubric-guided response tuning. Combine persona-conditioned examples with approximately 26K released responses selected by best-of-six rubric scoring. Convert the selected pairs to the training message format before merging and filtering.
55,998
C3
C2
Multi-source instruction mixture tuning. Combine approximately 30K persona-conditioned, 24K rubric-selected, and 16K Nemotron examples. Retain non-thinking responses from the additional source, remove visual markers, and filter by length.
69,900
C4
C3
Execution-verified targeted constraint tuning. Mix approximately 23.4K verified synthetic examples with 13K replay examples. Target character and keyword frequencies, punctuation, capitalization, and sentence counts identified by the preceding error analysis.
36,299
C5
C3
Replay-preserving additive constraint tuning. Retain the full 69,900-example C3 mixture and add approximately 23,750 execution-verified examples, then deduplicate and filter. Unlike C4, this trial adds targeted supervision without replacing the broader mixture.
93,000
C6
C5
Residual-error-guided constraint augmentation. Extend the 93K-example mixture with 14,524 verified examples covering remaining lexical, counting, formatting, length, and punctuation errors.
107,000
Appendix
Table 4: Data preparation and training ancestry for the extended IFEval Base experiment. Counts in the last column are final exported examples.
Checkpoints
Learning rate
Sequence limit
C1–C3
1×10−5
4,096
C4
5×10−6
4,096
C5–C6
1×10−5
4,096
C7–C8
1×10−5
8,192
Appendix
Table 5: Training settings for the extended IFEval run. The sequence limit is measured in tokens.
Variant
Procedure omitted
Basic alternative retained
w/o EA
Joint metric- and case-level attribution with explicit focal-problem selection.
An unstructured summary of the evaluation results guides the next action.
w/o UD
Required candidate comparison, reproduction-difficulty assessment, resource verification, and outcome-aware reuse.
Ordinary research followed by adoption of a feasible method. The agent may still inspect resources or reconsider its choice.
w/o ER
Required sample-level checks, short training runs, and the execution-guided refinement procedure before full execution.
Ordinary execution and debugging within the same model-update stage. Repairs remain possible after problems are noticed.
Appendix
Table 6: Basic alternatives used in the ablations. ER denotes Execution-Guided Refinement within Executable Model Updates.
Variant
Step 1
Step 2
Step 3
MMLU-Pro
Full
Train on 20K long mathematical reasoning examples.
Train on 20K concise final-answer examples; repair end-of-sequence handling and premature termination.
Introduce cross-domain supervision using complete, verified teacher responses.
w/o EA
Train on 10K auxiliary knowledge questions with short rationales.
Expand to 18K examples and add output-format anchors.
Directly train on long mathematical reasoning traces.
Switch to final-answer supervision while termination and template issues remain unresolved.
Train on 215K examples without checking teacher-response completeness.
MATH-500
Appendix
Table 7: Recorded strategy and execution paths for the five Instruct-model ablation variants. These records provide context for the comparisons in Table 2 .
Training language models (LMs) remains a highly human-intensive process, even as frontier language model agents become increasingly capable at software engineering and other long-horizon tasks. A central challenge is that autonomous post-training is not just a coding problem: it requires the agent to repeatedly plan iterations, construct benchmark-aligned data, run stable training jobs, evaluate checkpoints, and preserve experiment state across many hours of interaction. We present AutoTrainess, a LM agent that exposes these operations as a repository of agent-computer interfaces for planning, data preparation, training, evaluation, and logging. Rather than leaving the agent to operate in a raw CLI environment with an underspecified action space, AutoTrainess externalizes prior human experience as explicit workflows, rules, and execution constraints that guide the agent toward effective and reliable training behavior. On PostTrainBench, AutoTrainess consistently outperforms CLI-only baselines, achieving 26.94 average score with GPT-5.4 (Codex) versus 23.21 for CLI-only. It also generalizes across models and harnesses, improving DeepSeek-V4-Flash (OpenCode) from 12.13 to 19.58.
Zhaojian Yu, Penghao Yin, Shuzheng Gao +3
1Tsinghua University · 2The Chinese University of Hong Kong · 3Simple Agent Lab
Post-training a frontier model is normally weeks of human work: proposing data and recipe changes, launching runs, reading evals, deciding what to keep. We report an autonomous system that runs this loop with no human in the loop, post-training a 30B Nemotron across four rounds over multiple weeks. The autonomously produced model reaches a held-out score of 0.86 against the top human submission's 0.87 on the public NVIDIA Nemotron-Reasoning Challenge leaderboard, placing 8th of ~4000 at the time of writing. More striking than the number: the loop detected that its own dev metric had stopped tracking external performance on the weakest domain -- candidates drove dev to record highs without moving the external target -- and revised its own search policy, no longer maximizing dev but seeking interventions that lowered the now-misleading proxy while improving the external target. We treat this as direct, auditable evidence that a scaled autonomous loop can produce discovery, not only optimization: it detected that its measurement frame had become misleading and changed what counted as evidence. We take the operational view that any system worth the "recursive self-improvement" label must eventually perform end-to-end post-training of a frontier-class model; this is one datapoint of that bar being cleared. We do not claim a "first autonomous match" of human researchers. The claim we make is narrower and auditable: to our knowledge, this is the first publicly reported autonomous post-training run at this scale, where prior public autonomous-ML-research demonstrations sit at GPT-2-class (~124M) budgets. The same system also post-trains the 120B and 550B Nemotron; with no public human baseline there, this shows only that the loop closes at that scale, not that its output is competitive -- infrastructure evidence, with the effectiveness claim deferred until a comparable human anchor exists.
Autonomous research seeks sustained model improvements through iterative experimentation and feedback. LLM agents show promise in automating machine learning and language-model post-training, but their ability to sustain multimodal improvement remains unclear. We introduce MMPostTrainBench, a benchmark spanning eight tasks in image, audio, video, and joint audio-video understanding and image-grounded software repair. Agents operate from a common base model within fixed budgets, using development feedback before independent evaluation of their submitted models. Evaluation covers target and non-target model outcomes, iterative model improvement and selection, and research integrity. Across all eight tasks, 52.1% of model--task means fall below the base, and evaluated submissions also exhibit non-target regressions. Model performance does not consistently improve across research iterations, and agents do not reliably select the best evaluated candidate for submission; final submissions trail that candidate by up to 5.38 percentage points. Extending autonomous research from text-only to multimodal tasks introduces additional sources of error in perception, cross-modal alignment, and temporal grounding. The observed regressions and selection gaps highlight the need to balance targeted improvements with non-target capability preservation and to retain gains across research iterations. These requirements motivate MMResearch, a multimodal research framework that connects media-grounded evidence to hypotheses and interventions, carries findings across rounds through hierarchical memory, and retains candidates using development evaluation. Added to existing code-agent runtimes, it improves submitted-model accuracy by up to 7.75 percentage points for Claude Opus 4.8 with Claude Code and 2.33 points for GPT-5.6-sol with Codex.
Yuxin Liu, Yuxuan Wang, Zhenxin Lei +9
University of Science and Technology of China · Alibaba Token Hub, Alibaba Group · University of the Chinese Academy of Sciences +2