cs.AIOct 5, 2026

Internalizing Agent Experience into Diffusion Model Weights via On-Policy Context Distillation

Authors: Wenxuan Wang, Zekai Liu, Weinan Zhang, Yu Cheng, Yang Yang

Abstract

Wrapping an image generation model in an agentic harness can effectively boost Text-to-Image task performance: the harness can leverage memory, skills, workflow orchestration, result verification, and iterative refinement to continually construct and revise prompts, thereby eliciting better images. These gains, however, remain external to the diffusion model and are realized only while the full harness runs. We propose Diffusion On-Policy Context Distillation (D-OPCD), which treats the agent-improved prompt as privileged context and distills the knowledge encoded in the agent harness into the weights of the diffusion model, so that the model retains part of the harness's benefit when conditioned on the original query alone. Using a Text-to-Image agent equipped with our proposed Auto Skill Evolver (ASE), we show that D-OPCD can internalize harness capabilities into the generator's weights, raising the average direct-generation score from 60.52 to 65.09 across four benchmarks. With this knowledge absorbed into the weights, the harness can shed its saturated skills and resume evolving: a second ASE round on the updated generator improves on a skill-free harness by additional 1.83 points, pointing toward text-to-image systems in which harness and model keep improving each other through continual co-evolution.

Explore similar work

CardsList
  1. D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models

    May 6, 2026Dengyang Jiang, Xin Jin, Dongyang Liu +9Supervised Fine-TuningDiffusion Model Distillation

  2. PE-OPSD: Internalizing Prompt Enhancement into Flow-matching Models via On-Policy Self-Distillation

    Sep 29, 2026Mingfeng Lin, Chengfei Cai, Lin Xu +3On-Policy Self-DistillationT2I Generation

  3. DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models

    Date pendingQuanhao Li, Junqiu Yu, Kaixun Jiang +7Diffusion Model DistillationMulti-Task RL