cs.CLOct 7, 2026

Learning to Act with Task Progress: Distilling Small Agents from Compact Teacher Supervision

Authors: Wenxi Gan

Organizations: City University of Hong Kong Shenzhen Loop Area Institute

Abstract

Learning from large-model demonstrations offers a way to train small agents that can complete recurring tasks without calling a large model at every step. A central design choice is what to retain from teacher trajectories that contain reasoning, actions, and information about task progress. We introduce Task-Progress Distillation (TPD), an offline approach that pairs each demonstrated action with a short label describing the current task stage. The student learns these compact targets and selects actions by jointly scoring admissible stage--action pairs, which a deterministic harness executes in the environment. On ALFWorld, a 1.7B student trained with 404 demonstrations achieves 72.4% mean unseen task success with either TPD or action-only supervision, compared with 48.3% for a reasoning-trained student using constrained action selection. Explicit stages provide an additional benefit at 200 demonstrations, improving success from 48.0% to 67.7% over action-only supervision. With more demonstrations, the action-only student closes the gap, and both approaches reach 76.9% at 808 demonstrations. Shared-history analyses link part of TPD's local advantage to better decisions when moving between subgoals, particularly from object acquisition to processing. These results show that compact supervision can train effective small task agents, while explicit task progress provides additional guidance at an intermediate demonstration budget.

Figures & tables

Explore similar work

CardsList
  1. On-Policy Distillation with Curriculum Turn-level Guidance for Multi-turn Agents

    Jun 14, 2026Gengsheng Li, Mao Zheng, Mingyang Song +8TeacherCurriculum

  2. Look Ahead Before You Distill: Future Trajectory Validation of Teacher Guidance for Agentic On-Policy Distillation

    Aug 3, 2026Chishui Chen, Yaoyou Fan, Te Sun +11Efficient On-Policy DistillationTeacher