EvoGen-Harness: Learning Where and How to Evolve Image-Generation Harnesses
Organizations: Peking University · Beijing University of Technology · Fudan University
Abstract
Modern text-to-image (T2I) systems can be improved without modifying generator parameters by adapting the external system around frozen generators. However, existing approaches typically optimize a predefined dimension, such as prompts, routing, or workflows, restricting the space in which generation failures can be corrected. Allowing multiple generator-external responsibilities to evolve provides a broader adaptation space, but introduces a new challenge: visual feedback reveals what failed, but not where persistent evolution should occur or how this space should be explored efficiently. We introduce EvoGen-Harness, a generator-agnostic framework for multi-responsibility image-generation harness evolution, together with Trace (Trajectory-Relative Attribution and Coordinated Evolution). Trace aggregates evidence across stochastic executions, uses failure attribution as a search prior to focus candidate updates, and progressively re-attributes residual failures to coordinate evolution across responsibilities, while No-Patch and held-out validation prevent unnecessary or harmful updates. Across GenEval2, T2I-CompBench++, and WISE, EvoGen-Harness improves over the strongest evaluated baselines by +0.2633, +0.0720, and +0.0752, respectively, while achieving 87.9-91.4% attribution recall, 94.8% No-Patch accuracy, and only 1.9% regression. These results demonstrate that attribution-guided multi-responsibility evolution can substantially enhance frozen T2I systems beyond single-dimension adaptation.
Figures & tables
| Type | Method | Object | Attribute | Count | Position | Verb | Overall |
|---|---|---|---|---|---|---|---|
| T2I Models | FLUX.1-dev | 0.8754 | 0.6734 | 0.5311 | 0.3928 | 0.2150 | 0.2115 |
| Janus-Pro | 0.9102 | 0.7213 | 0.4834 | 0.5217 | 0.2056 | 0.1878 | |
| Flow-GRPO | 0.9175 | 0.7438 | 0.6032 | 0.5233 | 0.2541 | 0.2198 | |
| Qwen-Image | 0.9730 | 0.8412 | 0.6857 | 0.6083 | 0.3752 | 0.3483 | |
| Z-Image | 0.9755 | 0.7851 | 0.6232 | 0.6041 | 0.2419 | 0.3034 | |
| Nano Banana | 0.9722 | 0.8981 | 0.6830 | 0.7044 | 0.4532 | 0.4456 |
| Method | Color | Shape | Texture | 2D-Spatial | 3D-Spatial | Numeracy | Non-Spatial | Complex | Average |
|---|---|---|---|---|---|---|---|---|---|
| T2I Models | |||||||||
| FLUX.1-dev | 0.7572 | 0.5066 | 0.6300 | 0.2700 | 0.3992 | 0.6165 | 0.3065 | 0.3628 | 0.4811 |
| Janus-Pro | 0.5145 | 0.3323 | 0.4069 | 0.1566 | 0.2753 | 0.4406 | 0.3137 | 0.3806 | 0.3526 |
| Flow-GRPO | 0.8379 | 0.6130 | 0.7236 | 0.5447 | 0.4471 | 0.6752 | 0.3195 | 0.3741 | 0.5669 |
| Qwen-Image | 0.8392 | 0.5941 | 0.7568 | 0.4512 | 0.4631 | 0.7705 | 0.3126 | 0.3944 | 0.5727 |
| Z-Image | 0.8413 | 0.6078 | 0.7512 | 0.3822 | 0.4412 | 0.6676 | 0.3181 | 0.4201 | 0.5537 |
| Method | Culture | Time | Space | Biology | Physics | Chemistry | Overall |
|---|---|---|---|---|---|---|---|
| Qwen-Image | 0.6275 | 0.5250 | 0.5583 | 0.3417 | 0.4833 | 0.2500 | 0.5100 |
| Nano Banana | 0.7000 | 0.5900 | 0.6700 | 0.4500 | 0.5800 | 0.4000 | 0.6028 |
| VisualPrompter | 0.5800 | 0.4500 | 0.5900 | 0.2300 | 0.4300 | 0.2900 | 0.4708 |
| OctoT2I | 0.5350 | 0.4160 | 0.5090 | 0.2620 | 0.5130 | 0.3140 | 0.4557 |
| Ours | 0.7589 | 0.6712 | 0.7160 | 0.5889 | 0.6651 | 0.4788 | 0.6780 |
| Method | Object | Attribute | Count | Position | Verb | Overall |
|---|---|---|---|---|---|---|
| FLUX.1-dev | 0.8754 | 0.6734 | 0.5311 | 0.3928 | 0.2150 | 0.2115 |
| + Promptist | 0.8533 | 0.6510 | 0.5122 | 0.3819 | 0.2134 | 0.2056 |
| + VisualPrompter | 0.9321 | 0.7223 | 0.5614 | 0.4350 | 0.2412 | 0.2552 |
| + Ours | 0.9799 | 0.9469 | 0.9053 | 0.6418 | 0.3851 | 0.7089 |
| over VisualPrompter | +0.0478 | +0.2246 | +0.3439 | +0.2068 | +0.1439 | +0.4537 |
| Qwen-Image | 0.9730 | 0.8412 | 0.6857 | 0.6083 | 0.3752 | 0.3483 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Responsibility | Persistent representation | Editable contents |
|---|---|---|
| Policy | Visual requirement and constraint rules | Requirement interpretation, constraint priority, and criteria used to represent explicit visual requirements such as count, attribute, relation, and action constraints. |
| Tools | Capability registry for available generation and auxiliary tools | Capability and limitation descriptions, applicable scenarios, and tool-specific knowledge used by the harness. Tool implementations and model parameters are never modified. |
| Skills | Reusable procedural templates | Multi-step generation routines, prompt templates, decomposition strategies, and reusable task-solving procedures. |
| Middleware | Run-time orchestration configuration | Routing, verification, retry, termination, and tool-invocation rules. The deterministic base controller itself remains fixed. |
| Memory | Persistent cross-task experience records | Reusable summaries of successful and failed executions, recurring failure patterns, and previously validated corrective experience. |
| Search Strategy | GenEval2 | Candidate Edits | Gen. Calls | Evol. Time (s) |
|---|---|---|---|---|
| Full-Space Search | 0.7112 | 20.0 | 80.0 | 2386 |
| Hard Top-1 | 0.6815 | 4.0 | 16.0 | 486 |
| Trace Top- | 0.7089 | 8.0 | 32.0 | 964 |
| Compound Causes | Cause Recall@3 | Both Recovered | Repair Success |
|---|---|---|---|
| Policy + Tools | 96.1% | 92.4% | 86.8% |
| Policy + Skills | 94.8% | 90.1% | 85.7% |
| Policy + Middleware | 94.6% | 89.5% | 84.3% |
| Policy + Memory | 95.5% | 91.3% | 85.1% |
| Tools + Skills | 94.1% | 88.7% | 83.6% |
| Tools + Middleware | 93.8% | 87.9% | 82.4% |
| Proposal LLM | GenEval2 | Held-out Gain | Commit Rate | Regression |
|---|---|---|---|---|
| Qwen3.5-27B | 0.6996 | +0.133 | 70.6% | 2.3% |
| GPT-5.5 | 0.7164 | +0.146 | 73.2% | 1.7% |
| GPT-4.1 (Default) | 0.7089 | +0.140 | 72.0% | 1.9% |