cs.LGSep 30, 2026

EvoGen-Harness: Learning Where and How to Evolve Image-Generation Harnesses

Authors: Jiabin Luo, Yinan Liu, Chunlei Meng, Yufei Guo

Organizations: Peking University · Beijing University of Technology · Fudan University

Abstract

Modern text-to-image (T2I) systems can be improved without modifying generator parameters by adapting the external system around frozen generators. However, existing approaches typically optimize a predefined dimension, such as prompts, routing, or workflows, restricting the space in which generation failures can be corrected. Allowing multiple generator-external responsibilities to evolve provides a broader adaptation space, but introduces a new challenge: visual feedback reveals what failed, but not where persistent evolution should occur or how this space should be explored efficiently. We introduce EvoGen-Harness, a generator-agnostic framework for multi-responsibility image-generation harness evolution, together with Trace (Trajectory-Relative Attribution and Coordinated Evolution). Trace aggregates evidence across stochastic executions, uses failure attribution as a search prior to focus candidate updates, and progressively re-attributes residual failures to coordinate evolution across responsibilities, while No-Patch and held-out validation prevent unnecessary or harmful updates. Across GenEval2, T2I-CompBench++, and WISE, EvoGen-Harness improves over the strongest evaluated baselines by +0.2633, +0.0720, and +0.0752, respectively, while achieving 87.9-91.4% attribution recall, 94.8% No-Patch accuracy, and only 1.9% regression. These results demonstrate that attribution-guided multi-responsibility evolution can substantially enhance frozen T2I systems beyond single-dimension adaptation.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. EPIC: Efficient Predicate-Guided Inference-Time Control for Compositional Text-to-Image Generation

    May 12, 2026Sunung Mun, Sunghyun Cho, Jungseul OkStrongest Prior Asynchronous BaselineInference-Time Steering

  2. Generation Navigator: A State-Aware Agentic Framework for Image Generation

    May 18, 2026Jinming Liu, Ruoyu Feng, Yuqi Wang +2Text-To-ImageImage Generation

  3. MemoGen: Can Past Experience Improve Future Text-to-Image Generation?

    Jun 2, 2026Wenshuo Chen, Kuimou Yu, Bowen Tian +10Text-To-ImageFrozen Backbone