Dependency-Aware Reward Shaping for Agentic Reinforcement Learning
Organizations: University of Illinois Urbana-Champaign · National University of Singapore · Zhejiang University · University of Rochester
Abstract
When training large language models with reinforcement learning, terminal rewards provide little guidance about which steps matter. Common methods for assigning step credit overlook that work built on uncorrected mistakes is wasted while independent work remains valid. With only a final success/failure reward, every step in a failed episode has zero total future reward, even when it made progress. We propose Dependency-Aware Reward Shaping (DARS), which represents task progress as predicates linked by prerequisite relations and assigns step-level credit over the dependency graph. An annotator marks which predicates each step verifies, invalidates, or repairs. Verified predicates are discounted according to graph distance from the nearest broken prerequisite, while independent predicates are unaffected. Repairs update these weights based on any errors that remain; invalidated predicates need re-verification to regain credit. A fixed potential converts these annotations into signed per-step rewards. A common reward and annotation interface allows DARS to integrate with a range of reasoning and agentic training methods, such as GiGPO and ARPO/AEPO, without changing their rollout strategies or optimizers. Across five task families and models from 1.5B to 8B, DARS improves success by up to 10 points over GiGPO trained with the same budget and harness (ALFWorld), raises the WebShop task score and Search-R1 QA accuracy, complements AEPO's entropy-based training on AIME24/25 with a Python interpreter, and exceeds OmniOPD in controlled tool-free reasoning comparisons at 1.7B and 4B. Ablations show that step-level credit, dependency attenuation, and graph topology each contribute. On ALFWorld, a distilled 8B annotator matches the API annotator, enabling DARS to run efficiently without a frontier judge. Code is available at https://github.com/JianhuiWei7/DARS.
Figures & tables
| Method | 1.5B | 7B |
|---|---|---|
| GRPO | † | |
| RLOO | † | |
| GiGPO (w/ std) † | ||
| GiGPO (w/o std) † | ||
| GiGPO (same budget) | ||
| DARS (ours) |
| Arm | Step 5 | Step 10 | Step 15 | Grid mean | Best | vs. init. |
|---|---|---|---|---|---|---|
| ARPO ( Dong et al., 2026b ) | 61.7 | 61.3 | 62.9 | 62.0 | 62.9 | |
| AEPO ( Dong et al., 2026a ) | 63.3 | 61.7 | 60.0 | 61.7 | 63.3 | |
| ARPO DARS | 59.2 | 61.3 | 64.2 | 61.6 | 64.2 | |
| AEPO DARS | 60.4 | 67.5 | 62.9 | 63.6 | 67.5 |
| Qwen3-1.7B | Qwen3-4B | |||||||
|---|---|---|---|---|---|---|---|---|
| Method | AMC | AIME24 | AIME25 | Avg. | AMC | AIME24 | AIME25 | Avg. |
| Untrained model | 84.38 | 50.56 | 37.50 | 57.48 | 97.21 | 73.75 | 66.55 | 79.17 |
| StaRPO | 84.37 | 50.42 | 37.50 | 57.43 | 97.03 | 72.40 | 67.40 | 78.94 |
| On-policy distillation | 84.45 | 50.62 | 36.88 | 57.32 | 96.83 | 71.64 | 67.39 | 78.62 |
| OmniOPD | 85.52 | 49.86 | 36.34 | 57.24 | 96.33 | 72.43 | 66.56 | 78.44 |
| DARS w/o step credit | 84.94 | 49.25 | 36.50 | 56.90 | 97.48 | 74.20 | 65.71 | 79.13 |
| Qwen3-1.7B | Qwen3-4B | |||||||
|---|---|---|---|---|---|---|---|---|
| Variant | AMC | AIME24 | AIME25 | Avg. | AMC | AIME24 | AIME25 | Avg. |
| Scalar only ( ) | 84.94 | 49.25 | 36.50 | 56.90 | 97.48 | 74.20 | 65.71 | 79.13 |
| DARS ( ) | 86.53 | 49.92 | 38.25 | 58.23 | 96.72 | 74.64 | 67.12 | 79.50 |
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
| Measure | ungated | gated |
|---|---|---|
| Failed trajectories with (over-credit) | 0.12 | 0.02 |
| Failed trajectories with | 0.14 | 0.04 |
| Count-aware graphs on multi-object tasks | 0/48 | 48/48 |
| Mean on successful trajectories (preserve) | 0.91 | 0.89 |
| Usable verdicts / median latency | — | 122/123 / 23 s |
| Arm / configuration [checkpoint] | Draws | Metric | Table |
|---|---|---|---|
| ALFWorld — Qwen2.5-1.5B-Instruct 500 steps, 128 unseen games | |||
| GiGPO (local) [final 500] | 4 | 86.9 | 8 |
| GiGPO (local) [milestones 100 / 150 / 200 / 300 / 400] | 4 | 67.7 / 86.5 / 90.2 / 84.6 / 86.3 | 9 |
| DARS ( , ) [final 500] | 4 | 96.9 | 1 |
| DARS [milestones 100 / 150 / 200 / 300 / 400] | 4 | 77.9 / 91.6 / 91.0 / 90.1 / 93.4 | 9 |
| ALFWorld — Qwen2.5-1.5B-Instruct 200 steps, 128 unseen games | |||
| Setting | DARS | Local GiGPO reference | vs. reference | vs. published |
|---|---|---|---|---|
| ALFWorld 1.5B | (same 500-step budget) | ( ; 4/4) | ||
| ALFWorld 7B | (same 500-step budget) | ( ; 9W/2L/5T) | ||
| WebShop 1.5B ‡ | / | / (parent, 400 steps) | / ( ; 11/12) | |
| WebShop 7B ‡ | / | / (parent, 600 steps) | / ( ; 12/12) | |
| Search-R1 shared | (662 steps; DARS 360) | (3/3 groups) | – | |
| Search-R1 held-out | (one epoch each) | (3/3 groups) | – |
| Environment | Method | Success | Task score | ( ) |
| ALFWorld 1.5B, 500 steps | GiGPO (local) | – | – | |
| DARS (ours) | – | ( ; 4/4 draws) | ||
| ALFWorld 1.5B, 200 steps | GRPO (local) | – | ||
| RLOO (local) | – | |||
| GiGPO (local) | – | – | ||
| DARS (ours) | – | ( ; 4/4 draws) |
| Step | 100 | 150 | 200 | 300 | 400 | 500 (final) |
|---|---|---|---|---|---|---|
| GiGPO (local) | ||||||
| DARS | ||||||
| (draws won) | (4/4) | (4/4) | (3/4) | (4/4) | (4/4) | (4/4) |
| Method | Pick | Look | Clean | Heat | Cool | Pick2 | All |
|---|---|---|---|---|---|---|---|
| GiGPO | 72.6 | 100.0 | 82.6 | 100.0 | 93.5 | 97.5 | 88.7 |
| + DARS | 88.3 | 91.9 | 93.2 | 90.3 | 92.4 | 100.0 | 92.0 |
| Selection | Method | Step | Draws | Success | W/8 | |
|---|---|---|---|---|---|---|
| Peak | GiGPO | 470 | 8 | – | – | |
| GRPO | 460 | 16 | 0/8 | |||
| RLOO | 460 | 16 | 0/8 | |||
| Final | GiGPO | 500 | 8 | – | – | |
| GRPO | 500 | 16 | 0/8 | |||
| RLOO | 500 | 16 | 0/8 |
| Size | Checkpoint | Success | Task score | success (won; ) | task score (won; ) |
| 1.5B, selection draws (seeds 0–11) | |||||
| 1.5B | GiGPO parent (400 steps) | – | – | ||
| DARS, graded, seed 1, step 400 (locked) | (11/12; .001) | (12/12; .0005) | |||
| DARS, graded, seed 0, step 400 | (9/12; .004) | (10/12; .003) | |||
| DARS, graded, mixing weight , step 400 ( ) | (10/11; .002) | (11/11; .001) | |||
| 1.5B, confirmation draws (fresh seeds 12–23) | |||||
| Size (added steps) | GiGPO continuation parent | DARS continuation parent | DARS GiGPO continuation |
|---|---|---|---|
| success / task score | success / task score | success; task score | |
| 7B (+250) | / | / | (20/2/2; ); ( ) |
| 1.5B (+400) | / | / | ( ); ( ) |
| Success | Task score | |||||
| Step | DARS | GiGPO twin | (W/L; ) | DARS | GiGPO twin | (W/L; ) |
| 50 | (7/4; ) | (12/0; ) | ||||
| 100 | (12/0; ) | (12/0; ) | ||||
| 150 | (11/0; ) | (11/1; ) | ||||
| 200 a | (3/4; ) | (3/5; ) | ||||
| 50 b | (9/3; ) | (12/0; ) | ||||
| Track | Topology | Node instantiation (per rollout) | Optional terms | |||
|---|---|---|---|---|---|---|
| ALFWorld 1.5B/7B | Object chains | objects, treatment | none | |||
| WebShop 1.5B (fixed budget) | Parallel-AND | attributes, options | / | optional verifier | ||
| WebShop 7B (fixed budget) | Parallel-AND | attributes, options | none | |||
| Search-R1 7B | Hop chain | required hops | answer-commitment term | |||
| Math with Python | Derivation chain | required steps | none | |||
| Math tool-free | Derivation chain | required steps | none |
| Variant | Success | Task score |
|---|---|---|
| GiGPO (local) | ||
| DARS, | ||
| DARS, | ||
| DARS, completion verifier |
| Arm | NQ | TriviaQA | PopQA | HotpotQA | 2Wiki | MuSiQue | Bamboogle | All |
|---|---|---|---|---|---|---|---|---|
| GiGPO (step 662) | 33.3 | 66.0 | 49.5 | 37.5 | 46.8 | 15.0 | 28.5 | 39.5 |
| DARS (step 360) | 39.0 | 66.2 | 54.7 | 50.2 | 42.7 | 17.7 | 37.2 | 44.0 |
| Arm | Selected step | test350 | vs. GiGPO (per group) |
|---|---|---|---|
| GiGPO (our reproduction) | 600 | – | |
| Commitment terms only | 100 | ( , , ) | |
| Graph credit only | 80 | ( , , ) | |
| DARS, | 410 | ||
| DARS (Table 1 ) | 490 | ( , , ) |
| Step | Arm | MATH500 | AIME24 | AIME25 | Mean AIME |
|---|---|---|---|---|---|
| 5 | ARPO | — | 68.3 | 55.0 | 61.7 |
| AEPO | 93.3 | 72.5 | 54.2 | 63.3 | |
| ARPO DARS | 95.8 | 64.2 | 54.2 | 59.2 | |
| AEPO DARS | 93.3 | 65.8 | 55.0 | 60.4 | |
| 10 | ARPO | — | 65.8 | 56.7 | 61.3 |
| AEPO | 90.0 | 65.8 | 57.5 | 61.7 |
| Arm | MATH500 | GSM8K | MATH |
|---|---|---|---|
| AEPO | 94.8 | 95.6 | 97.3 |
| AEPO DARS | 94.9 | 96.1 | 96.9 |
| ARPO DARS | 94.6 | 95.8 | 96.2 |
| Signal source | Arm | cells | Avg. | vs. DARS ( /wins) |
| Dependency graph | DARS ( , ) | 10 | 58.23 (1.09) | – |
| DARS, (same graphs) | 14 | 57.50 (1.13) | (7/10) | |
| DARS ( ) | 3 | 57.27 (1.75) | ||
| DARS, wider group admission | 3 | 57.14 (0.44) | ||
| DARS, broken-node penalty ( ) | 3 | 57.10 (0.15) | ||
| DARS, answer-commit gates | 3 | 57.08 (0.33) |
| Gate | 1.7B | 4B |
|---|---|---|
| Usable induced chains | 96.9% (prior judge 61.4%) | 91.2% |
| Answer-node agreement with the outcome verifier | 79% | 80.2% |
| Groups with nonzero variance | 142/200 | 263 groups |
| Segments carrying credit | 28.6% | 28.3% |
| Negative credit share | 3.5% | 1.2% |
| Arm | Avg. | AIME | |
|---|---|---|---|
| DARS ( ) | 20 | 79.50 (1.06) | 70.88 |
| Untrained model | 14 | 79.17 (0.77) | 70.15 |
| DARS w/o step credit | 13 | 79.13 (0.81) | 69.95 |
| StaRPO (paper segmentation) | 4 | 78.94 (1.17) | 69.90 |
| Rejection fine-tuning | 4 | 78.79 (1.16) | 69.64 |
| On-policy distillation | 15 | 78.62 (0.96) | 69.51 |
| Success | vs. | |
|---|---|---|
| (milestone counting) | 90.2 | – |
| 94.3 | ||
| (default) | 91.4 | |
| (hard masking) | 84.2 |
| Corpus | Variant | Changed | Mean change | 10th pct. | lost |
|---|---|---|---|---|---|
| WebShop 1.5B (371 graphs) | sequential chain | 12.7% | 25.2% | ||
| 8.6% | 0.0% | ||||
| ALFWorld 1.5B (2000 graphs) | sequential chain | 10.6% | 0.5% | ||
| 11.7% | 1.2% | ||||
| ALFWorld 7B (27,520 graphs) | 2.0% | – | – | – | |
| Math 1.7B (775 graphs) | 10.6% | – | – | – |
| Environment | Goal structure | First-error prefix credit vs. GiGPO | Observed behavior |
|---|---|---|---|
| ALFWorld | ordered chain | points success | chain credit and graph credit nearly coincide |
| WebShop | parallel-AND | points task score | credits 7 browsing turns without a purchase |
| Search-R1 | instrumental prefix | success falls to | of trajectories never commit an answer |
| Executed action | Recorded events | Credit | ||
|---|---|---|---|---|
| 0 | search[blu ray gold plated high speed heavy duty HDMI cable blue braided 20-pack price < 40.00] | — | ||
| 1 | click[b08qsnm69h] | Verify: T, A1, A2, A3, A4, P | ||
| 2 | click[20-pack] | Verify: S | ||
| 3 | click[blue braided] | Verify: C | ||
| 4 | click[buy now] | — |
| Executed action | Recorded events | Credit | ||
|---|---|---|---|---|
| 0 | search[Vanity bronze finish vanity light price < 110.00] | — | ||
| 1 | click[b09j9y9h95] | Verify: T, B, L, P | ||
| 2 | click[back to search] | Error: B, L | ||
| 3 | search[ Kira Home Ainsley 21.5" 3-Light Farmhouse Vanity/Bathroom Light + Clear Cylinder Glass Shades, Oil-Rubbed Bronze Finish $77.99] | — | ||
| 4 | click[b09j9y9h95] | Verify: T, B, L, P; Repair: B, L | ||
| 5 | click[buy now] | — |
| DARS | GiGPO | |
|---|---|---|
| 0 | Search: Uruguayan former footballer who managed team with Charlie Oatway | Search: What is the birthdate of this Uruguayan former footballer, who management team consists of Charlie Oatway |
| 1 | Search: Gus Poyet Uruguayan football manager | Search: Gastón Pereyra birthdate uruguayan former footballer |
| 2 | Answer: 15 November 1967 | Search: Gastón Pereyra birthdate uruguayan former footballer |
| 3 | — | Answer: 30 May 1984 |
| DARS | GiGPO | |
|---|---|---|
| 0 | go to countertop 1 | go to countertop 1 |
| 1 | go to toilet 1 | take soapbottle 1 from countertop 1 |
| 2 | take soapbar 1 from toilet 1 | go to garbagecan 1 |
| 3 | go to garbagecan 1 | move soapbottle 1 to garbagecan 1 |
| 4 | move soapbar 1 to garbagecan 1 | go to countertop 1 |
| 5 | go to toilet 1 | go to toilet 1 |
| DARS | GiGPO | |
|---|---|---|
| 0 | go to sidetable 1 | go to sidetable 1 |
| 1 | go to shelf 1 | go to dresser 1 |
| 2 | go to dresser 1 | go to sofa 1 |
| 3 | go to sofa 1 | go to sidetable 1 |
| 4 | go to sidetable 1 | go to armchair 1 |
| 5 | go to armchair 1 | take pillow 1 from armchair 1 |
| DARS | GiGPO | |
|---|---|---|
| 0 | Search: Team USA baseball at 2000 Summer Olympics | Search: Which of the starting pitchers for Team USA in the Baseball at the 2000 Summer Olympics was not born August 29, 1977 |
| 1 | Search: Ben Sheets born date 1977 August 29 | Answer: Juan Pe~na |
| 2 | Answer: Ben Sheets | — |