cs.AIOct 4, 2026

MMPostTrainBench: Benchmarking Autonomous Research for Multimodal Post-Training

Authors: Yuxin Liu, Yuxuan Wang, Zhenxin Lei, Lingchen Meng, Yuchong Sun, Junming Lin, Hongcheng Liu, Yunfei Chu, +4 more

Organizations: University of Science and Technology of China · Alibaba Token Hub, Alibaba Group · University of the Chinese Academy of Sciences · Tsinghua University · Shanghai Jiao Tong University

Abstract

Autonomous research seeks sustained model improvements through iterative experimentation and feedback. LLM agents show promise in automating machine learning and language-model post-training, but their ability to sustain multimodal improvement remains unclear. We introduce MMPostTrainBench, a benchmark spanning eight tasks in image, audio, video, and joint audio-video understanding and image-grounded software repair. Agents operate from a common base model within fixed budgets, using development feedback before independent evaluation of their submitted models. Evaluation covers target and non-target model outcomes, iterative model improvement and selection, and research integrity. Across all eight tasks, 52.1% of model--task means fall below the base, and evaluated submissions also exhibit non-target regressions. Model performance does not consistently improve across research iterations, and agents do not reliably select the best evaluated candidate for submission; final submissions trail that candidate by up to 5.38 percentage points. Extending autonomous research from text-only to multimodal tasks introduces additional sources of error in perception, cross-modal alignment, and temporal grounding. The observed regressions and selection gaps highlight the need to balance targeted improvements with non-target capability preservation and to retain gains across research iterations. These requirements motivate MMResearch, a multimodal research framework that connects media-grounded evidence to hypotheses and interventions, carries findings across rounds through hierarchical memory, and retains candidates using development evaluation. Added to existing code-agent runtimes, it improves submitted-model accuracy by up to 7.75 percentage points for Claude Opus 4.8 with Claude Code and 2.33 points for GPT-5.6-sol with Codex.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. AutoTrainess: Teaching Language Models to Improve Language Models Autonomously

    Jun 30, 2026Zhaojian Yu, Penghao Yin, Shuzheng Gao +3Large Language Model TrainingLanguage Modeling

  2. ImproveAnyTask: An Autonomous Post-Training Harness for Iterative Model Self-Improvement

    Oct 5, 2026Xingbo Yao, Xiaoman Wang, Zhengwu Lei +11