Beyond Instruction Following: Learning Grounded Skill-Following with Skill Contracts
Organizations: Nanjing University · Shanghai Artificial Intelligence Laboratory · Peking University · Yiyue Technology · Tsinghua University · Northeastern University
Abstract
Instruction following typically enforces discrete, response-level requirements, whereas an expert-authored skill prescribes procedural requirements spanning multiple phases and environment interactions. Given such a skill, we train the executor to execute all required phases instead of focusing solely on the final answer. We therefore introduce Grounded Skill-Following, which requires an agent to execute a fixed, expert-authored skill across its required phases by grounding decisions in environment observations. To achieve verifiable procedural execution, we formulate each skill as a skill contract combining visible skill instructions with an explicit contract runtime. The runtime specifies required phases, admissible actions, permitted transitions, and accepted termination. This structure provides a dense, verifiable training signal throughout execution. We leverage this by introducing Verified Progress Credit, which assigns rewards upon the initial completion of contract milestones and aggregates them into the trajectory return to guide policy optimization. During rollout, the contract runtime continuously tracks state transitions to provide Contract-State Feedback, which indicates whether the latest action is accepted and guides the agent toward valid next actions. To measure procedural compliance, we introduce the Protocol Completion Rate (PCR), defined as reaching accepted termination through all required phases, and decouple it from the final Task Outcome. Jointly trained with our framework, Qwen3.5-4B achieves Protocol Completion Rates of 99.27% on Math and 99.96% on Search, while slightly outperforming original baselines in Task Outcome (82.95% and 46.61%, respectively). Controlled studies examine how skill instructions, training signals, and contract-state feedback affect both metrics, while withholding interventions evaluate behavioral dependence on observation content.
Figures & tables
| Math | Search | ||||||||||||||
| Setting | MATH tr | GSM tr | Avg. | PCR | RF-PCR | NQ tr | TQA ho | Pop ho | HQA tr | 2W ho | MuS ho | Bam ho | Avg. | PCR | RF-PCR |
| Qwen2.5-3B | |||||||||||||||
| Direct | 17.00 | 26.08 | 21.54 | n/a | n/a | 2.27 | 6.31 | 3.11 | 2.84 | 5.16 | 0.12 | 1.60 | 3.06 | n/a | n/a |
| Prompt | 0.16 | 0.00 | 0.08 | 0.16 | 0.00 | 0.08 | 0.14 | 0.14 | 0.09 | 0.25 | 0.00 | 0.00 | 0.10 | 0.20 | 0.00 |
| Recipe | 45.16 | 81.20 | 63.18 | n/a | n/a | 45.26 | 61.73 | 44.00 | 32.14 | 29.21 | 6.37 | 10.40 | 32.73 | n/a | n/a |
| Ours-S | 48.62 | 80.59 | 64.61 | 97.29 | 94.54 | 46.34 | 60.53 | 45.67 | 30.34 | 29.87 | 6.91 | 16.00 | 33.67 | 99.82 | 98.54 |
| A: Recovery coefficient | B: Feedback removal | ||||||||
| Domain | Task Outcome | PCR | RF-PCR | Task Outcome | PCR | ||||
| On | Off | Change | On | Off | |||||
| Math | 0.90 | 62.12 | 94.46 | 93.78 | 64.61 | 58.25 | 97.29 | 86.93 | |
| 1.00 | 64.61 | 96.23 | 0.00 | ||||||
| Search | 0.90 | 28.81 | 98.90 | 98.18 | 30.11 | 13.15 | 98.90 | 43.05 | |
| 1.00 | 29.51 | 99.53 | 88.33 | ||||||
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Phase | Accepted action | Next phase |
|---|---|---|
| NEED_SKILL | Expected skill call | NEED_PLAN |
| NEED_PLAN | Substantive planning thought | NEED_ACTION |
| NEED_ACTION | Admissible environment action | NEED_EVIDENCE_THINK |
| NEED_EVIDENCE_THINK | Substantive evidence reasoning | CAN_CONTINUE_OR_ANSWER after a valid observation, otherwise NEED_ACTION |
| CAN_CONTINUE_OR_ANSWER | Repeated environment action, first optional Search evidence-organization action with <summary> , or final answer | NEED_EVIDENCE_THINK , CAN_CONTINUE_OR_ANSWER , or DONE , respectively. A repeated evidence-organization action is rejected |
| DONE | None | Terminal |
| Order | Operation | Update |
|---|---|---|
| 1 | Initialize | Set initial contract state . Initialize context with system scaffold and task prompt (Appendix B.1 ). |
| 2 | Generate | Sample action . |
| 3 | Transition | Execute state transition via Equation 4 . Set when no observation is generated. Action rejection preserves and , setting when feedback is enabled. Action acceptance records new milestones scaled by , then resets . |
| 4 | Record | Record newly verified milestones and update context via Equation 5 , omitting and prepending skill instructions upon first accepted skill call. |
| 5 | Repeat | Repeat Order 2–4 until accepted termination or budget exhaustion. Exhaustion does not constitute Protocol Completion. |
| 6 | Score | Compute final trajectory return using Equation 7 . |
| Milestone | Scope | Weight | Contract location | First-completion condition |
|---|---|---|---|---|
| Skill | Both | 0.10 | NEED_SKILL | The expected <skill_call> is accepted. |
| Plan | Both | 0.10 | NEED_PLAN | A substantive <think> is accepted. |
| Action | Both | 0.20 | NEED_ACTION | The first admissible <code> , <search> , or <research> passes its task adapter. |
| Observation | Both | 0.20 | NEED_ACTION NEED_EVIDENCE_THINK | Math execution produces nonempty output, or Search returns at least one nonempty document. |
| Evidence reasoning | O-Math | 0.20 | NEED_EVIDENCE_THINK | A substantive <think> is accepted. |
| Evidence reasoning | O-Search | 0.10 | NEED_EVIDENCE_THINK | A substantive <think> is accepted. |
| Phase | O-Math | O-Search |
|---|---|---|
| NEED_SKILL | Accept the expected code-interpreter-protocol skill call. | Accept the expected qa-search-protocol skill call. |
| NEED_PLAN | Accept a substantive <think> that plans the calculation. | Accept a substantive <think> that plans retrieval. |
| NEED_ACTION | Accept executable Python that prints a result in <code>...</code> . The restricted interpreter returns <interpreter> . The observation is valid when execution succeeds and printed output is nonempty. The accepted action enters NEED_EVIDENCE_THINK . | Accept the first admissible query in <search> and later nonrepeated admissible queries in <research> . The frozen retriever returns <information> . The observation is valid when at least one retrieved document is nonempty. The accepted action enters NEED_EVIDENCE_THINK . |
| NEED_EVIDENCE_ THINK | Accept a substantive <think> that interprets the interpreter result. A valid observation opens continuation or answering. An invalid one returns to NEED_ACTION . | Accept a substantive <think> that assesses the retrieved evidence. A valid observation opens continuation or answering. An invalid one returns to NEED_ACTION . |
| CAN_CONTINUE_ OR_ANSWER | Repeat <code>...</code> and re-enter evidence reasoning, or accept a nonempty boxed answer and enter DONE . | Repeat retrieval with <research> and re-enter evidence reasoning. Optionally accept one substantive <summary> as evidence organization, with no retriever call, no observation, and no phase change. Alternatively, accept a nonempty <answer> and enter DONE . |
| DONE | Terminal. Score the accepted answer by normalized numerical match. | Terminal. Score the accepted answer by dataset-normalized exact match. |
| Condition | GPU | TP | Rounds | P/R | S/P | U/R | Batch |
|---|---|---|---|---|---|---|---|
| O-Math / Math interventions | 4 | 2 | 200 | 72 | 5 | 5 | 72 |
| O-Search / Search interventions | 4 | 2 | 200 | 512 | 5 | 10 | 256 |
| O-Joint | 4 | 2 | 400 | 72 | 5 | 5 | 72 |
| Recovery-coefficient comparison | 4 | 2 | 400 | 72 | 5 | 5 | 72 |
| Domain | Condition | Feedback | Credit | Outcome | PCR |
|---|---|---|---|---|---|
| Search | B00 Outcome-only | 0 | 0 | 0.00 | 0.00 |
| Search | B10 Feedback-only | 1 | 0 | 0.00 | 0.00 |
| Search | B01 Credit-only | 0 | 1 | 29.28 | 98.69 |
| Search | B11 Feedback + Credit | 1 | 1 | 30.11 | 98.90 |
| Math | B00 Outcome-only | 0 | 0 | 0.00 | 0.00 |
| Math | B10 Feedback-only | 1 | 0 | 0.00 | 0.00 |
| Domain | Visible skill | MATH | GSM | 2Wiki | Hotpot | Avg. | PCR | Hint Rate |
|---|---|---|---|---|---|---|---|---|
| Math | Original | 48.62 | 80.59 | – | – | 64.61 | 97.29 | 5.59 |
| Math | Removed | 42.76 | 74.68 | – | – | 58.72 | 91.34 | 11.70 |
| Math | Paraphrased | 43.92 | 75.59 | – | – | 59.75 | 94.36 | 7.27 |
| Math | Wrong Search | 1.48 | 2.35 | – | – | 1.92 | 4.71 | 99.69 |
| Search | Original | – | – | 29.87 | 30.34 | 30.11 | 98.90 | 0.99 |
| Search | Removed | – | – | 22.57 | 18.51 | 20.54 | 63.81 | 14.86 |