GraphOPD: Graph-Augmented On-Policy Distillation for LLM Agents
Organizations: University of Science and Technology of China · Xiaohongshu · The Hong Kong University of Science and Technology (HKUST-GZ)
Abstract
On-policy distillation post-trains large language model agents by supplying dense, step-level guidance from a teacher policy when the reinforcement-learning reward is sparse and arrives only once per trajectory. Existing instantiations allocate this guidance by the size of the teacher-student divergence at each step, on the single-turn intuition that a large disagreement marks a mistake worth correcting. Once decisions chain over many turns, that rule misfires, since an early drift enters every later context both policies condition on, leaving the teacher consistent with the drifted trajectory instead of flagging its cause, while interchangeable steps register large but outcome-irrelevant divergences. We demonstrate this on an agentic benchmark, where distilling the highest-divergence steps brings no consistent benefit over random selection. To this end, we introduce GraphOPD, the first method to bring graph-based structural augmentation into on-policy distillation for agent capabilities. It reads which steps enabled which later ones from the environment's own record of state changes, immune to the drift that corrupts the teacher-student gap, organizes them into a dependency graph, scores each step by a random-walk stationary distribution over it, and fuses that structural credit with the divergence signal into a trajectory-relative mask concentrating supervision on each rollout's highest-aptitude steps. Across three model scales and eleven baselines on ALFWorld, WebShop, and SearchQA, GraphOPD shows competitive performance throughout, improving over the strongest baseline by up to +5.8 pp. An executed-replay audit further shows that this structural credit score tracks true causal impact far above chance, that both fused signals are independently necessary, and that the same signal transfers to out-of-domain tool-integrated reasoning.
Figures & tables
| ALFWorld | SearchQA | WebShop | |||||||||||||||
| Method | Pick | Look | Clean | Heat | Cool | Pick2 | Avg | NarQA | Triv | Pop | Hotp | 2Wk | MuS | Bam | Avg | Score | Succ. |
| Qwen2.5-3B-Instruct | |||||||||||||||||
| Vanilla | 44.4 | 11.1 | 6.2 | 15.4 | 28.6 | 12.5 | 21.9 | 24.6 | 48.1 | 31.0 | 26.3 | 25.3 | 7.2 | 59.7 | 31.7 | 6.7 | 0.8 |
| GRPO | 91.2 | 62.5 | 96.2 | 61.9 | 65.0 | 47.4 | 75.0 | 39.3 | 60.6 | 41.1 | 37.4 | 34.6 | 15.4 | 26.4 | 36.4 | 79.8 | 63.3 |
| Skill-Prompt | 51.7 | 66.7 | 48.4 | 0.0 | 4.3 | 10.0 | 28.9 | 23.7 | 46.2 | 30.6 | 24.4 | 22.1 | 7.5 | 12.5 | 23.9 | 0.2 | 0.8 |
| OPSD | 48.8 | 41.7 | 16.7 | 0.0 | 15.8 | 16.7 | 28.1 | 0.1 | 0.1 | 0.1 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 11.3 | 3.1 |
| Qwen3-1.7B | Qwen2.5-3B | Qwen2.5-7B | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Method | WebShop | SearchQA | ALFWorld | WebShop | SearchQA | ALFWorld | WebShop | SearchQA | ALFWorld |
| GraphOPD | 73.4 | 42.9 | 63.3 | 78.9 | 46.1 | 88.6 | 83.6 | 47.3 | 91.9 |
| Soft-Pruning | 71.1 | 37.4 | 62.5 | 77.3 | 44.2 | 85.9 | 82.0 | 45.3 | 90.6 |
| Divergence-only | 64.8 | 32.4 | 59.4 | 74.2 | 40.8 | 84.4 | 79.6 | 42.3 | 87.5 |
| Structural-only | 69.5 | 36.4 | 60.2 | 75.0 | 43.3 | 84.4 | 81.3 | 43.6 | 89.1 |
| Shuffled Centrality | 68.8 | 40.1 | 59.4 | 72.7 | 41.2 | 82.8 | 78.1 | 41.9 | 85.9 |
| 0 | 0.2 | 0.5 | 0.7 | 1.0 | |
|---|---|---|---|---|---|
| ALFWorld | 89.1 | 90.6 | 91.9 | 87.5 | 87.5 |
| WebShop | 81.3 | 79.7 | 83.6 | 76.6 | 79.6 |
| Signal | Hit@10% | |
|---|---|---|
| 0.58 | 0.48 | |
| 0.34 | 0.41 | |
| Random | 0.13 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Environment | Key action types | Footprint key format |
|---|---|---|
| ALFWorld | go to, take, put, heat/cool/clean, open/close | location:<loc> , visible_at:<loc> , holding:<obj> , placed:<obj>@<loc> , object_state:<obj>:<prop> , container_state:<container> |
| WebShop | search[q], click[product_id], click[attr], click[buy] | query_token:<word> , one per query word, selected_product:<id> , selected_attr:<attr> , purchase_action |
| SearchQA | <search> , <think> , <answer> | query_token:<word> , retrieved_passages , reasoning_state , answer_candidate , passage_entity:<entity> |
| Benchmark | Graph construction (s/step) | Total step time (s) | Share |
|---|---|---|---|
| ALFWorld | 0.6659 | 342.0936 | 0.19% |
| WebShop | 0.4535 | 144.2885 | 0.31% |
| SearchQA | 0.4144 | 322.2142 | 0.13% |
| Qwen3-1.7B | Qwen2.5-3B | Qwen2.5-7B | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Method | ALFWorld | SearchQA | WebShop | ALFWorld | SearchQA | WebShop | ALFWorld | SearchQA | WebShop |
| GraphOPD | 63.3 | 42.9 | 73.4 | 88.6 | 46.1 | 78.9 | 91.9 | 47.3 | 83.6 |
| SDAR | 51.3 | 40.7 | 59.9 | 81.8 | 43.4 | 69.3 | 84.4 | 47.5 | 82.0 |
| OPID | 58.2 | 39.6 | 66.4 | 82.8 | 44.6 | 74.5 | 89.1 | 48.2 | 80.0 |
| GEPO | 58.1 | 40.1 | 67.9 | 80.7 | 44.6 | 75.3 | 85.9 | 45.5 | 80.0 |
| Source within the turn | Footprint key format |
|---|---|
| Code block and reasoning text | symbol:<identifier> , one per name bound by an assignment, a function definition, or a comprehension, call:<function> , one per interpreter function invoked, import:<module> , one per module imported, answer_candidate:<value> , for the value placed inside \boxed{} or submitted as the final program |
| Interpreter response | result:<value> , one per value printed or returned, test_result:<example_id> , one per public example checked on LiveCodeBench-v5 |
| Method | AIME24 | AIME25 | LCB-v5 | GPQA-D | Avg |
|---|---|---|---|---|---|
| Vanilla | 11.9 | 8.8 | 14.2 | 21.7 | 14.2 |
| GRPO | 15.1 | 13.3 | 15.1 | 28.3 | 18.0 |
| OPSD | 19.2 | 12.2 | 16.4 | 31.5 | 19.8 |
| SDAR | 21.4 | 16.3 | 15.8 | 26.3 | 20.0 |
| GraphOPD | 29.8 | 21.4 | 28.2 | 36.4 | 29.0 |
| Steps | Late checkpoint, step 150 (succeeds) | Early checkpoint, step 20 (fails) |
|---|---|---|
| 0–2 | Action: go to sinkbasin 1 Action: go to countertop 1 Action: take soapbottle 1 from countertop 1 | Action: go to cabinet 1 Action: open cabinet 1 Action: examine cabinet 2 |
| 3–5 / 3–49 | Action: go to cabinet 1 Action: open cabinet 1 Action: move soapbottle 1 to cabinet 1 Result: reward , episode ends at step 5 | Actions (47 steps): examine cabinet , look , open cabinet , go to cabinet , plus 8 other revisit actions, never mentioning the soapbottle Raw KL: ranges 0.16 to 0.65, mean 0.32 Result: reward , times out at step 49 |
| The step-150 checkpoint carries the soapbottle from step 2 onward, so every later action follows directly from the object it already holds. | The step-20 checkpoint repeatedly opens and examines cabinets without ever picking up the soapbottle, so its failure is an entity it never engaged with rather than a late misstep. |
| Step | Action | Raw KL | |
|---|---|---|---|
| 3 | examine countertop 1 | 0.33 | |
| 4 | go to cabinet 1 | 0.36 | |
| 5 | open cabinet 1 | 0.29 | |
| 6 | go to countertop 1 | 0.26 | |
| 7 | examine countertop 1 | 0.17 | |
| 8 | go to cabinet 1 | 0.29 | 0.31 |
| Step | Action | Raw KL | |
|---|---|---|---|
| 0 | go to fridge 1 | 0.31 | |
| 1 | examine fridge 1 | 0.44 | |
| 2 | go to countertop 1 | 0.28 | |
| 3 | take potato 1 from countertop 1 | 0.22 | |
| 4 | go to microwave 1 | 0.24 | 0.95 |
| 5 | open microwave 1 | 0.26 |