Organizations: University of Science and Technology of China · Alibaba Token Hub, Alibaba Group · University of the Chinese Academy of Sciences · Tsinghua University · Shanghai Jiao Tong University
Autonomous research seeks sustained model improvements through iterative experimentation and feedback. LLM agents show promise in automating machine learning and language-model post-training, but their ability to sustain multimodal improvement remains unclear. We introduce MMPostTrainBench, a benchmark spanning eight tasks in image, audio, video, and joint audio-video understanding and image-grounded software repair. Agents operate from a common base model within fixed budgets, using development feedback before independent evaluation of their submitted models. Evaluation covers target and non-target model outcomes, iterative model improvement and selection, and research integrity. Across all eight tasks, 52.1% of model--task means fall below the base, and evaluated submissions also exhibit non-target regressions. Model performance does not consistently improve across research iterations, and agents do not reliably select the best evaluated candidate for submission; final submissions trail that candidate by up to 5.38 percentage points. Extending autonomous research from text-only to multimodal tasks introduces additional sources of error in perception, cross-modal alignment, and temporal grounding. The observed regressions and selection gaps highlight the need to balance targeted improvements with non-target capability preservation and to retain gains across research iterations. These requirements motivate MMResearch, a multimodal research framework that connects media-grounded evidence to hypotheses and interventions, carries findings across rounds through hierarchical memory, and retains candidates using development evaluation. Added to existing code-agent runtimes, it improves submitted-model accuracy by up to 7.75 percentage points for Claude Opus 4.8 with Claude Code and 2.33 points for GPT-5.6-sol with Codex.
Figures & tables
Figure 1: MMPostTrainBench at a glance. (a) Final model outcomes: post-training effectiveness measured by mean task score across all eight tasks relative to the common base model (Table 2 ). (b) Iterative model improvement: best-so-far held-out candidate accuracy on MMMU-Pro; labels report the best candidate score.
Figure 2: MMPostTrainBench benchmark overview. Agents conduct post-training from a common base model on eight multimodal tasks within fixed budgets. To limit evaluation leakage, execution and verification run in separate containers, research feedback is restricted to authorized development data, and final tests remain sealed. Submitted models, candidate trajectories, and audits support evaluation of model outcomes, iterative model improvement, and research integrity.
Task
Domain
Model Ability
MMMU-Pro ( Yue et al., 2025 )
Image
Image-grounded multidisciplinary reasoning
MMAU ( Sakshi et al., 2025 )
Audio
Audio understanding and reasoning
MMAR ( Ma et al., 2025 )
Audio
Audio reasoning
Video-MMMU ( Hu et al., 2025 )
Video
Video-grounded multidisciplinary reasoning
VideoMME-v2 ( Fu et al., 2026 )
Video
Video understanding
JointAVBench ( Chao et al., 2026 )
Omni
Joint audio-video understanding
Table 1: Eight post-training tasks, their domains, and target model abilities.
Figure 3: MMResearch framework. MMResearch coordinates a seven-step research loop (b), using hierarchical memory (a) to guide experimentation and multimodal evidence (c) to connect observations with experimental decisions, then freezes the retained candidate for delivery (d).
Category
Benchmark
Base model
Claude Opus 5
GPT-5.6 sol
Qwen 3.8 Omni Flash
GPT-5.6 terra
Gemini 3.8 Flash
Claude Opus 4.8
Claude Code
Codex
Qwen Code
Codex
Gemini CLI
Claude Code
Image
MMMU-Pro
41.91
43.26 +1.35
41.88 -0.03
44.71 +2.80
43.00 +1.09
41.88 -0.03
41.72 -0.19
Audio
MMAU
48.60
50.25 +1.65
50.30 +1.70
50.00 +1.40
49.40 +0.80
50.00 +1.40
44.01 -4.59
MMAR
72.40
72.16 -0.24
72.31 -0.09
72.01 -0.39
73.00 +0.60
73.50 +1.10
66.77 -5.63
Video
Video-MMMU
60.00
60.89 +0.89
60.40 +0.40
53.96 -6.04
58.42 -1.58
49.01 -10.99
56.44 -3.56
VideoMME-v2
24.81
23.40 -1.41
24.35 -0.46
25.81 +1.00
22.49 -2.32
23.21 -1.60
22.89 -1.92
Table 2: Final model outcomes on MMPostTrainBench. Three-loop mean task score (%) and change from base (pp). Harnesses appear below model names; bold marks the best result per row. The aggregate weights all eight tasks equally.
Target benchmark
Non-target modalities
Claude Opus 5
GPT-5.6 sol
GPT-5.6 terra
Gemini 3.8 Flash
Qwen 3.8 Omni Flash
Claude Opus 4.8
MMMU-Pro
Audio, Omni
+4.08
+0.84
-0.51
+2.78
-4.40
-8.45
MMAR
Image, Omni
-0.98
-0.79
-0.57
-0.15
+0.43
+1.35
MMAU
Image, Omni
+1.71
-1.06
+1.26
-1.43
+0.04
+4.75
Video-MMMU
Image, Audio
+1.18
+1.01
+0.39
+1.86
-1.17
-2.95
VideoMME-v2
Image, Audio
+0.82
+0.76
-0.42
-1.08
-1.50
-0.08
JointAVBench
Image, Audio
-0.85
+0.90
+0.34
-15.31
-3.37
-5.24
Table 3: Cross-modal capability retention. Cells show mean accuracy changes over the two indicated non-target probes (pp): Image/MMMU-Pro, Audio/MMAR, and Omni/JointAVBench. Underlining indicates the best result in each row.
Figure 4: Mean time and tokens.
Figure 8
Figure 6: Recorded integrity flags.
Research model
Harness
MMMU-Pro
MMAR
VideoMME-v2
OmniVideoBench
Image
Audio
Video
Omni
Claude Opus 4.8
Claude Code
41.72
66.77
22.89
31.86
+ MMResearch
43.74 (+2.02)
72.40 (+5.63)
24.81 (+1.92)
39.61 (+7.75)
GPT-5.6-sol
Codex
41.88
72.31
24.35
42.99
+ MMResearch
44.21 (+2.33)
72.90 (+0.60)
24.81 (+0.46)
42.61 (-0.38)
Base model (main-table reference)
41.91
72.40
24.81
39.61
Table 5: Effect of MMResearch on submission quality. Accuracy (%) measures final model performance; parentheses report gains over the same harness without MMResearch (pp).
Figure 7: Development feedback budget on MMMU-Pro.
Variant
Gain ↑
Full
+1.83
w/o Hierarchical Memory
+0.19
w/o Multimodal Evidence
+0.92
Table 6: MMResearch component ablations on MMMU-Pro. Gain is the accuracy change from the base (pp).
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Domain
Dev.
Test
Measure
MMMU-Pro
Image
558
1,172
Answer accuracy
MMAU
Audio
332
668
Answer accuracy
MMAR
Audio
332
668
Answer accuracy
Video-MMMU
Video
98
202
Answer accuracy
VideoMME-v2
Video
1,003
2,197
Answer accuracy
JointAVBench
Omni
905
1,948
Answer accuracy
Appendix
Table 7: Task categories, development and test set sizes, and evaluation measures.
Target domain
Probe 1
Probe 2
Image
MMAR
JointAVBench
Audio
MMMU-Pro
JointAVBench
Video
MMAR
MMMU-Pro
Omni
MMAR
MMMU-Pro
Code
MMMU-Pro
MMAR
Appendix
Table 8: Non-target probe assignments.
Probe setting
Probe
Test items
Base accuracy (%)
Gemini / Qwen
MMMU-Pro
1,172
41.7235
MMAR
668
72.9042
JointAVBench
1,948
60.0616
SWE-bench Multimodal
MMMU-Pro
1,730
41.91
MMAR
1,000
73.60
Appendix
Table 9: Matched baselines for non-target probes. SWE-bench Multimodal probes use the same submitted checkpoints as its target evaluation.
Target
Probe
Claude Opus 5
GPT-5.6 sol
Qwen 3.8 Omni Flash
GPT-5.6 terra
Gemini 3.8 Flash
Claude Opus 4.8
MMMU-Pro
MMAR
+1.25
+1.70
−9.88
+0.50
−0.60
−8.03
JointAVBench
+6.90
−0.03
+1.08
−1.52
+6.16
−8.86
MMAU
MMMU-Pro
−2.66
−1.98
−2.65
−0.53
−1.88
+1.26
JointAVBench
+6.08
−0.14
+2.72
+3.05
−0.98
+8.23
MMAR
MMMU-Pro
−0.53
−0.78
+0.34
−0.44
−0.09
−1.13
JointAVBench
−1.42
−0.80
+0.51
−0.70
−0.21
+3.82
Appendix
Table 10: Complete non-target accuracy changes (pp). Columns follow the research configurations in the main table. Matched probe baselines are listed in Table 9 .
Research model
Runtime
Mean time
Mean tokens
GPT-5.6-sol
Codex
21.7
103.9
Claude Opus 5
Claude Code
23.1
71.8
Qwen 3.8 Omni Flash
Qwen Code
17.8
69.0
GPT-5.6-terra
Codex
23.7
41.5
Gemini 3.8 Flash
Gemini CLI
17.4
35.5
Claude Opus 4.8
Claude Code
9.1
23.0
Appendix
Table 11: Recorded resource summaries. Time is in hours and agent tokens in millions.
Research model
Flagged
Trajectory
Contamination
Workspace
Split
GPT-5.6-sol
2
8/8
8/8
8/8
8/8
Claude Opus 5
2
8/8
8/8
8/8
8/8
GPT-5.6-terra
2
8/8
8/8
8/8
8/8
Claude Opus 4.8
1
8/8
8/8
8/8
8/8
Gemini 3.8 Flash
2
8/8
8/8
8/8
8/8
Qwen 3.8 Omni Flash
1
8/8
8/8
8/8
8/8
Appendix
Table 12: Recorded integrity flags and audit coverage. Counts cover flags and completed checks.
Research model
Target
Raw run
Base reset
Three-run mean
GPT-5.6-sol
MMMU-Pro
41.38
41.91
41.88
GPT-5.6-sol
VideoMME-v2
24.35
24.81
24.35
Claude Opus 5
MMAU
50.45
48.60
50.25
Claude Opus 5
OmniVideoBench
42.38
39.61
42.40
GPT-5.6-terra
MMAR
73.05
72.40
73.00
GPT-5.6-terra
JointAVBench
67.92
59.94
59.97
Appendix
Table 13: Audit-triggering records and updated target means (%). Raw and base scores refer to the flagged historical run; the last column averages three runs after per-loop audit adjustment.
Training language models (LMs) remains a highly human-intensive process, even as frontier language model agents become increasingly capable at software engineering and other long-horizon tasks. A central challenge is that autonomous post-training is not just a coding problem: it requires the agent to repeatedly plan iterations, construct benchmark-aligned data, run stable training jobs, evaluate checkpoints, and preserve experiment state across many hours of interaction. We present AutoTrainess, a LM agent that exposes these operations as a repository of agent-computer interfaces for planning, data preparation, training, evaluation, and logging. Rather than leaving the agent to operate in a raw CLI environment with an underspecified action space, AutoTrainess externalizes prior human experience as explicit workflows, rules, and execution constraints that guide the agent toward effective and reliable training behavior. On PostTrainBench, AutoTrainess consistently outperforms CLI-only baselines, achieving 26.94 average score with GPT-5.4 (Codex) versus 23.21 for CLI-only. It also generalizes across models and harnesses, improving DeepSeek-V4-Flash (OpenCode) from 12.13 to 19.58.
Zhaojian Yu, Penghao Yin, Shuzheng Gao +3
1Tsinghua University · 2The Chinese University of Hong Kong · 3Simple Agent Lab
Post-training a frontier model is normally weeks of human work: proposing data and recipe changes, launching runs, reading evals, deciding what to keep. We report an autonomous system that runs this loop with no human in the loop, post-training a 30B Nemotron across four rounds over multiple weeks. The autonomously produced model reaches a held-out score of 0.86 against the top human submission's 0.87 on the public NVIDIA Nemotron-Reasoning Challenge leaderboard, placing 8th of ~4000 at the time of writing. More striking than the number: the loop detected that its own dev metric had stopped tracking external performance on the weakest domain -- candidates drove dev to record highs without moving the external target -- and revised its own search policy, no longer maximizing dev but seeking interventions that lowered the now-misleading proxy while improving the external target. We treat this as direct, auditable evidence that a scaled autonomous loop can produce discovery, not only optimization: it detected that its measurement frame had become misleading and changed what counted as evidence. We take the operational view that any system worth the "recursive self-improvement" label must eventually perform end-to-end post-training of a frontier-class model; this is one datapoint of that bar being cleared. We do not claim a "first autonomous match" of human researchers. The claim we make is narrower and auditable: to our knowledge, this is the first publicly reported autonomous post-training run at this scale, where prior public autonomous-ML-research demonstrations sit at GPT-2-class (~124M) budgets. The same system also post-trains the 120B and 550B Nemotron; with no public human baseline there, this shows only that the loop closes at that scale, not that its output is competitive -- infrastructure evidence, with the effectiveness claim deferred until a comparable human anchor exists.
Adapting general-purpose large language models to specific tasks requires substantial human effort in designing data and training strategies. Sustaining improvement is especially challenging because model updates change the error distribution, requiring strategies to be continually refined. We introduce ImproveAnyTask, an autonomous post-training harness that improves task performance under a limited compute budget. Drawing inspiration from gradient-based parameter optimization, the harness organizes adaptation into error attribution, update-direction selection, and executable model updates. It combines metric-level and case-level analysis to identify a focal problem, then investigates research-backed strategies and compares their reported gains and reproduction difficulty. The selected strategy is translated into training data and a training configuration, with small-scale execution checks preceding full post-training. Subsequent evaluation guides model selection and further adaptation, while validated strategies and scripts are retained for reuse. Across 11 tasks, ImproveAnyTask achieves mean gains of 18.29 and 11.97 percentage points on the Base and Instruct models, respectively, with a maximum gain of 41.96 points, under a 24-hour budget with resources equivalent to eight H20 GPUs.