cs.AIOct 6, 2026

How Much Evidence Should a Coding Agent's Self-Correction Carry? Adaptive Dirichlet Evidence for Self-Distillation

Authors: Yunbo Long, Guangya Hao, Yuhan Liu, Yiting Duan, Longyan Tan, Yunchen Long, Hao Wu

Abstract

Execution feedback lets coding agents revise programs and learn from their own corrections. A correction's learning weight should reflect both the transitions supported by its executions and the amount of evidence behind that support. We introduce Effective-Evidence Self-Distillation (EESD), which represents these quantities separately. Normalized execution relevance determines relative transition support and an effective pseudo-count mass; a Dirichlet posterior then produces an uncertainty-penalized weight for KL-anchored correction learning. Under a symmetric prior, changing mass preserves category ordering, and effective mass yields a supervised coefficient bounded by its matched fixed-mass counterpart. Across four model-domain history sweeps, increasing visible observations from one to eight reduces future-outcome NLL by 55.0-59.3%. At eight observations, effective mass achieves lower NLL than fixed mass in all four comparisons. In the primary matched DeepSeek/RunBugRun study, argmax predictions agree on all 3,000 examples, with the largest NLL gain under concentrated relevance. After one correction-learning round, DeepSeek/CodeARC all-tests Pass@1 increases from 15.0% to 20.4%, with a paired 95% source-bootstrap interval of [+2.8, +8.0] percentage points. The twelve-setting downstream evaluation establishes the model-domain scope of this update. These results show how separating evidence support from evidence mass changes probability estimation and correction learning in coding agents.

Explore similar work

CardsList
  1. Self Improvement via Fast Tree-search

    Sep 17, 2026Xinghong Fu, Aravinth Kulanthaivelu, Yutaro YamadaInference-Time SearchAI Coding Agents

  2. EviSD: Evidence-Conditioned Self-Distillation for Search-Augmented Agents

    Aug 2, 2026Jianan Xie, Xin Sun, Zhongqi Chen +4Agentic SearchCredit Assignment in RL

  3. Latent Programming Horizons in Coding Agents

    Jul 6, 2026André Silva, Han Tu, Martin MonperrusLLM InterpretabilityCoding Agents