Organizations: University of Science and Technology of China · Alibaba Token Hub, Alibaba Group · University of the Chinese Academy of Sciences · Tsinghua University · Shanghai Jiao Tong University
Autonomous research seeks sustained model improvements through iterative experimentation and feedback. LLM agents show promise in automating machine learning and language-model post-training, but their ability to sustain multimodal improvement remains unclear. We introduce MMPostTrainBench, a benchmark spanning eight tasks in image, audio, video, and joint audio-video understanding and image-grounded software repair. Agents operate from a common base model within fixed budgets, using development feedback before independent evaluation of their submitted models. Evaluation covers target and non-target model outcomes, iterative model improvement and selection, and research integrity. Across all eight tasks, 52.1% of model--task means fall below the base, and evaluated submissions also exhibit non-target regressions. Model performance does not consistently improve across research iterations, and agents do not reliably select the best evaluated candidate for submission; final submissions trail that candidate by up to 5.38 percentage points. Extending autonomous research from text-only to multimodal tasks introduces additional sources of error in perception, cross-modal alignment, and temporal grounding. The observed regressions and selection gaps highlight the need to balance targeted improvements with non-target capability preservation and to retain gains across research iterations. These requirements motivate MMResearch, a multimodal research framework that connects media-grounded evidence to hypotheses and interventions, carries findings across rounds through hierarchical memory, and retains candidates using development evaluation. Added to existing code-agent runtimes, it improves submitted-model accuracy by up to 7.75 percentage points for Claude Opus 4.8 with Claude Code and 2.33 points for GPT-5.6-sol with Codex.
Figures & tables
Figure 1: MMPostTrainBench at a glance. (a) Final model outcomes: post-training effectiveness measured by mean task score across all eight tasks relative to the common base model (Table 2 ). (b) Iterative model improvement: best-so-far held-out candidate accuracy on MMMU-Pro; labels report the best candidate score.
Figure 2: MMPostTrainBench benchmark overview. Agents conduct post-training from a common base model on eight multimodal tasks within fixed budgets. To limit evaluation leakage, execution and verification run in separate containers, research feedback is restricted to authorized development data, and final tests remain sealed. Submitted models, candidate trajectories, and audits support evaluation of model outcomes, iterative model improvement, and research integrity.
Task
Domain
Model Ability
MMMU-Pro ( Yue et al., 2025 )
Image
Image-grounded multidisciplinary reasoning
MMAU ( Sakshi et al., 2025 )
Audio
Audio understanding and reasoning
MMAR ( Ma et al., 2025 )
Audio
Audio reasoning
Video-MMMU ( Hu et al., 2025 )
Video
Video-grounded multidisciplinary reasoning
VideoMME-v2 ( Fu et al., 2026 )
Video
Video understanding
JointAVBench ( Chao et al., 2026 )
Omni
Joint audio-video understanding
Table 1: Eight post-training tasks, their domains, and target model abilities.
Figure 3: MMResearch framework. MMResearch coordinates a seven-step research loop (b), using hierarchical memory (a) to guide experimentation and multimodal evidence (c) to connect observations with experimental decisions, then freezes the retained candidate for delivery (d).
Category
Benchmark
Base model
Claude Opus 5
GPT-5.6 sol
Qwen 3.8 Omni Flash
GPT-5.6 terra
Gemini 3.8 Flash
Claude Opus 4.8
Claude Code
Codex
Qwen Code
Codex
Gemini CLI
Claude Code
Image
MMMU-Pro
41.91
43.26 +1.35
41.88 -0.03
44.71 +2.80
43.00 +1.09
41.88 -0.03
41.72 -0.19
Audio
MMAU
48.60
50.25 +1.65
50.30 +1.70
50.00 +1.40
49.40 +0.80
50.00 +1.40
44.01 -4.59
MMAR
72.40
72.16 -0.24
72.31 -0.09
72.01 -0.39
73.00 +0.60
73.50 +1.10
66.77 -5.63
Video
Video-MMMU
60.00
60.89 +0.89
60.40 +0.40
53.96 -6.04
58.42 -1.58
49.01 -10.99
56.44 -3.56
VideoMME-v2
24.81
23.40 -1.41
24.35 -0.46
25.81 +1.00
22.49 -2.32
23.21 -1.60
22.89 -1.92
Table 2: Final model outcomes on MMPostTrainBench. Three-loop mean task score (%) and change from base (pp). Harnesses appear below model names; bold marks the best result per row. The aggregate weights all eight tasks equally.
Target benchmark
Non-target modalities
Claude Opus 5
GPT-5.6 sol
GPT-5.6 terra
Gemini 3.8 Flash
Qwen 3.8 Omni Flash
Claude Opus 4.8
MMMU-Pro
Audio, Omni
+4.08
+0.84
-0.51
+2.78
-4.40
-8.45
MMAR
Image, Omni
-0.98
-0.79
-0.57
-0.15
+0.43
+1.35
MMAU
Image, Omni
+1.71
-1.06
+1.26
-1.43
+0.04
+4.75
Video-MMMU
Image, Audio
+1.18
+1.01
+0.39
+1.86
-1.17
-2.95
VideoMME-v2
Image, Audio
+0.82
+0.76
-0.42
-1.08
-1.50
-0.08
JointAVBench
Image, Audio
-0.85
+0.90
+0.34
-15.31
-3.37
-5.24
Table 3: Cross-modal capability retention. Cells show mean accuracy changes over the two indicated non-target probes (pp): Image/MMMU-Pro, Audio/MMAR, and Omni/JointAVBench. Underlining indicates the best result in each row.
Figure 4: Mean time and tokens.
Figure 8
Figure 6: Recorded integrity flags.
Research model
Harness
MMMU-Pro
MMAR
VideoMME-v2
OmniVideoBench
Image
Audio
Video
Omni
Claude Opus 4.8
Claude Code
41.72
66.77
22.89
31.86
+ MMResearch
43.74 (+2.02)
72.40 (+5.63)
24.81 (+1.92)
39.61 (+7.75)
GPT-5.6-sol
Codex
41.88
72.31
24.35
42.99
+ MMResearch
44.21 (+2.33)
72.90 (+0.60)
24.81 (+0.46)
42.61 (-0.38)
Base model (main-table reference)
41.91
72.40
24.81
39.61
Table 5: Effect of MMResearch on submission quality. Accuracy (%) measures final model performance; parentheses report gains over the same harness without MMResearch (pp).
Figure 7: Development feedback budget on MMMU-Pro.
Variant
Gain ↑
Full
+1.83
w/o Hierarchical Memory
+0.19
w/o Multimodal Evidence
+0.92
Table 6: MMResearch component ablations on MMMU-Pro. Gain is the accuracy change from the base (pp).
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Domain
Dev.
Test
Measure
MMMU-Pro
Image
558
1,172
Answer accuracy
MMAU
Audio
332
668
Answer accuracy
MMAR
Audio
332
668
Answer accuracy
Video-MMMU
Video
98
202
Answer accuracy
VideoMME-v2
Video
1,003
2,197
Answer accuracy
JointAVBench
Omni
905
1,948
Answer accuracy
Appendix
Table 7: Task categories, development and test set sizes, and evaluation measures.
Target domain
Probe 1
Probe 2
Image
MMAR
JointAVBench
Audio
MMMU-Pro
JointAVBench
Video
MMAR
MMMU-Pro
Omni
MMAR
MMMU-Pro
Code
MMMU-Pro
MMAR
Appendix
Table 8: Non-target probe assignments.
Probe setting
Probe
Test items
Base accuracy (%)
Gemini / Qwen
MMMU-Pro
1,172
41.7235
MMAR
668
72.9042
JointAVBench
1,948
60.0616
SWE-bench Multimodal
MMMU-Pro
1,730
41.91
MMAR
1,000
73.60
Appendix
Table 9: Matched baselines for non-target probes. SWE-bench Multimodal probes use the same submitted checkpoints as its target evaluation.
Target
Probe
Claude Opus 5
GPT-5.6 sol
Qwen 3.8 Omni Flash
GPT-5.6 terra
Gemini 3.8 Flash
Claude Opus 4.8
MMMU-Pro
MMAR
+1.25
+1.70
−9.88
+0.50
−0.60
−8.03
JointAVBench
+6.90
−0.03
+1.08
−1.52
+6.16
−8.86
MMAU
MMMU-Pro
−2.66
−1.98
−2.65
−0.53
−1.88
+1.26
JointAVBench
+6.08
−0.14
+2.72
+3.05
−0.98
+8.23
MMAR
MMMU-Pro
−0.53
−0.78
+0.34
−0.44
−0.09
−1.13
JointAVBench
−1.42
−0.80
+0.51
−0.70
−0.21
+3.82
Appendix
Table 10: Complete non-target accuracy changes (pp). Columns follow the research configurations in the main table. Matched probe baselines are listed in Table 9 .
Research model
Runtime
Mean time
Mean tokens
GPT-5.6-sol
Codex
21.7
103.9
Claude Opus 5
Claude Code
23.1
71.8
Qwen 3.8 Omni Flash
Qwen Code
17.8
69.0
GPT-5.6-terra
Codex
23.7
41.5
Gemini 3.8 Flash
Gemini CLI
17.4
35.5
Claude Opus 4.8
Claude Code
9.1
23.0
Appendix
Table 11: Recorded resource summaries. Time is in hours and agent tokens in millions.
Research model
Flagged
Trajectory
Contamination
Workspace
Split
GPT-5.6-sol
2
8/8
8/8
8/8
8/8
Claude Opus 5
2
8/8
8/8
8/8
8/8
GPT-5.6-terra
2
8/8
8/8
8/8
8/8
Claude Opus 4.8
1
8/8
8/8
8/8
8/8
Gemini 3.8 Flash
2
8/8
8/8
8/8
8/8
Qwen 3.8 Omni Flash
1
8/8
8/8
8/8
8/8
Appendix
Table 12: Recorded integrity flags and audit coverage. Counts cover flags and completed checks.
Research model
Target
Raw run
Base reset
Three-run mean
GPT-5.6-sol
MMMU-Pro
41.38
41.91
41.88
GPT-5.6-sol
VideoMME-v2
24.35
24.81
24.35
Claude Opus 5
MMAU
50.45
48.60
50.25
Claude Opus 5
OmniVideoBench
42.38
39.61
42.40
GPT-5.6-terra
MMAR
73.05
72.40
73.00
GPT-5.6-terra
JointAVBench
67.92
59.94
59.97
Appendix
Table 13: Audit-triggering records and updated target means (%). Raw and base scores refer to the flagged historical run; the last column averages three runs after per-loop audit adjustment.