Learning to Act with Task Progress: Distilling Small Agents from Compact Teacher Supervision
Organizations: City University of Hong Kong Shenzhen Loop Area Institute
Abstract
Learning from large-model demonstrations offers a way to train small agents that can complete recurring tasks without calling a large model at every step. A central design choice is what to retain from teacher trajectories that contain reasoning, actions, and information about task progress. We introduce Task-Progress Distillation (TPD), an offline approach that pairs each demonstrated action with a short label describing the current task stage. The student learns these compact targets and selects actions by jointly scoring admissible stage--action pairs, which a deterministic harness executes in the environment. On ALFWorld, a 1.7B student trained with 404 demonstrations achieves 72.4% mean unseen task success with either TPD or action-only supervision, compared with 48.3% for a reasoning-trained student using constrained action selection. Explicit stages provide an additional benefit at 200 demonstrations, improving success from 48.0% to 67.7% over action-only supervision. With more demonstrations, the action-only student closes the gap, and both approaches reach 76.9% at 808 demonstrations. Shared-history analyses link part of TPD's local advantage to better decisions when moving between subgoals, particularly from object acquisition to processing. These results show that compact supervision can train effective small task agents, while explicit task progress provides additional guidance at an intermediate demonstration budget.
Figures & tables
| Arm | Supervised target | Execution interface |
|---|---|---|
| A | Teacher reasoning and action | Free generation followed by action parsing |
| Same checkpoint as A | Generate reasoning, then score admissible actions | |
| B / TPD | Prefix-aligned stage and action | Joint scoring over stage–action pairs |
| C | Action only | Score action candidates |
| Constant phase field and action | Score fixed-stage action candidates | |
| Shuffled stage labels and unchanged actions | Same candidate structure as B |
| Number of training demonstrations | ||||
| Configuration | 98 | 200 | 404 | 808 |
| A: reasoning–action | ||||
| C: action only | ||||
| B: stage–action ( TPD ) | ||||
| Arm | Success (%) | Steps / episode | Output tokens / episode |
| A | 5.0 | 47.9 | 14,635 |
| 48.3 | 33.6 | 9,995 | |
| B / TPD | 72.4 | 22.0 | 299 |
| C | 72.4 | 22.2 | 213 |
| Demonstrations | Both succeed | B only | C only | Both fail |
|---|---|---|---|---|
| 200 | 183 | 89 | 10 | 120 |
| 808 | 284 | 25 | 25 | 68 |
| Arm | Seed 42 | Seed 1234 | Seed 2026 | Mean |
|---|---|---|---|---|
| B / TPD | 55.2 | 53.4 | 56.9 | 55.2 |
| C | 35.3 | 37.9 | 50.9 | 41.4 |
| 31.9 | 49.1 | 50.0 | 43.7 | |
| 31.0 | 42.2 | 39.7 | 37.6 |
| Task | Paired records | B processes successfully | C processes successfully |
|---|---|---|---|
| Heat | 52 | 49/52 (94.2%) | 28/52 (53.8%) |
| Clean | 60 | 59/60 (98.3%) | 51/60 (85.0%) |
| Cool | 49 | 44/49 (89.8%) | 45/49 (91.8%) |
| Criterion or coverage | Result |
|---|---|
| Both original policies choose progress | 81/114 |
| Only original B chooses progress | 15/114 |
| Only original C chooses progress | 1/114 |
| Neither original policy chooses progress | 17/114 |
| B-only cases with useful matching stage and sensitivity to a mismatch | 11/15 |
| Also require original B to select reference stage and matching action | 11/15 |