Beyond Instruction Following: Learning Grounded Skill-Following with Skill Contracts
Authors: Jianghan Shen, Zhenjie Liu, Yue Li, Jie Huang, Siqi Luo, Yiming Cheng, Yizhi Yao, Kaijie Zhang, +5 more
Organizations: Nanjing University · Shanghai Artificial Intelligence Laboratory · Peking University · Yiyue Technology · Tsinghua University · Northeastern University
Instruction following typically enforces discrete, response-level requirements, whereas an expert-authored skill prescribes procedural requirements spanning multiple phases and environment interactions. Given such a skill, we train the executor to execute all required phases instead of focusing solely on the final answer. We therefore introduce Grounded Skill-Following, which requires an agent to execute a fixed, expert-authored skill across its required phases by grounding decisions in environment observations. To achieve verifiable procedural execution, we formulate each skill as a skill contract combining visible skill instructions with an explicit contract runtime. The runtime specifies required phases, admissible actions, permitted transitions, and accepted termination. This structure provides a dense, verifiable training signal throughout execution. We leverage this by introducing Verified Progress Credit, which assigns rewards upon the initial completion of contract milestones and aggregates them into the trajectory return to guide policy optimization. During rollout, the contract runtime continuously tracks state transitions to provide Contract-State Feedback, which indicates whether the latest action is accepted and guides the agent toward valid next actions. To measure procedural compliance, we introduce the Protocol Completion Rate (PCR), defined as reaching accepted termination through all required phases, and decouple it from the final Task Outcome. Jointly trained with our framework, Qwen3.5-4B achieves Protocol Completion Rates of 99.27% on Math and 99.96% on Search, while slightly outperforming original baselines in Task Outcome (82.95% and 46.61%, respectively). Controlled studies examine how skill instructions, training signals, and contract-state feedback affect both metrics, while withholding interventions evaluate behavioral dependence on observation content.
Figures & tables
Figure 1: Complementary views of skills. Grounded skill-following fixes an expert-authored skill and trains its executor, alongside local instruction compliance and skill-lifecycle optimization. The views overlap and are not a historical taxonomy.
Figure 2: Skill-contract training for grounded skill-following. The contract Cj pairs visible skill instructions zj with a contract runtime. The runtime checks each action, credits newly verified progress, and provides contract-state feedback after each nonterminal transition. The feedback describes the current phase without exposing internal bookkeeping.
Math
Search
Setting
MATH tr
GSM tr
Avg.
PCR
RF-PCR
NQ tr
TQA ho
Pop ho
HQA tr
2W ho
MuS ho
Bam ho
Avg.
PCR
RF-PCR
Qwen2.5-3B
Direct
17.00
26.08
21.54
n/a
n/a
2.27
6.31
3.11
2.84
5.16
0.12
1.60
3.06
n/a
n/a
Prompt
0.16
0.00
0.08
0.16
0.00
0.08
0.14
0.14
0.09
0.25
0.00
0.00
0.10
0.20
0.00
Recipe
45.16
81.20
63.18
n/a
n/a
45.26
61.73
44.00
32.14
29.21
6.37
10.40
32.73
n/a
n/a
Ours-S
48.62
80.59
64.61
97.29
94.54
46.34
60.53
45.67
30.34
29.87
6.91
16.00
33.67
99.82
98.54
Table 1: Task Outcome, PCR, and RF-PCR on Math and Search (%). RF-PCR excludes samples that receive rejected-action feedback. Direct evaluates frozen checkpoints without the contract runtime, while Prompt preloads the skill and uses the runtime. Recipe pairs task-specific adaptations, Ours-S trains each domain separately, Ours-J trains them jointly, and n/a denotes an absent protocol metric. Bold marks the highest value in each metric column within each model block, including ties. Averages are unweighted, while tr and ho mark training and held-out datasets.
Figure 3: Qwen2.5-3B learning dynamics under the same contract runtime with feedback disabled. Each point averages a fixed monitor of 250 examples per dataset: (a) MATH and GSM8K, (b) Natural Questions and MuSiQue. Verified progress credit yields nonzero PCR and Task Outcome in both domains, while Outcome-only remains at zero.
A: Recovery coefficient
B: Feedback removal
Domain
ρ
Task Outcome
PCR
RF-PCR
Task Outcome
PCR
On
Off
Change
On
Off
Math
0.90
62.12
94.46
93.78
64.61
58.25
−6.36
97.29
86.93
1.00
64.61
96.23
0.00
Search
0.90
28.81
98.90
98.18
30.11
13.15
−16.96
98.90
43.05
1.00
29.51
99.53
88.33
Table 2: Training- and inference-time recovery interventions. Panel A compares independently trained Qwen2.5-3B joint conditions with ρ=0.90 and ρ=1.00 . Panel B removes feedback at inference from frozen Qwen2.5-3B single-task checkpoints. Both panels report unweighted averages over MATH and GSM8K for Math and 2WikiMultiHopQA and HotpotQA for Search. Values are percentages, and changes are percentage points.
Figure 4: Qwen2.5-3B interventions on (a) contract-state feedback and verified progress credit during training, (b) visible skill instructions at inference, and (c) observation content at inference. Values are percentages or percentage-point changes.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Phase
Accepted action
Next phase
NEED_SKILL
Expected skill call
NEED_PLAN
NEED_PLAN
Substantive planning thought
NEED_ACTION
NEED_ACTION
Admissible environment action
NEED_EVIDENCE_THINK
NEED_EVIDENCE_THINK
Substantive evidence reasoning
CAN_CONTINUE_OR_ANSWER after a valid observation, otherwise NEED_ACTION
CAN_CONTINUE_OR_ANSWER
Repeated environment action, first optional Search evidence-organization action with <summary> , or final answer
NEED_EVIDENCE_THINK , CAN_CONTINUE_OR_ANSWER , or DONE , respectively. A repeated evidence-organization action is rejected
DONE
None
Terminal
Appendix
Table 3: Phase set Qj shared by Math and Search. A rejected action preserves the current phase and returns one admissible next action. Table 6 defines substantive thought.
Order
Operation
Update
1
Initialize
Set initial contract state S0j=(NEED_SKILL,∅,0) . Initialize context x0j with system scaffold and task prompt (Appendix B.1 ).
2
Generate
Sample action at∼πθ(⋅∣xtj) .
3
Transition
Execute state transition via Equation 4 . Set ot+1=⊥ when no observation is generated. Action rejection preserves qtj and Mtj , setting ht+1=1 when feedback is enabled. Action acceptance records new milestones scaled by ρht , then resets ht+1=0 .
4
Record
Record newly verified milestones ΔMi,tj and update context via Equation 5 , omitting ⊥ and prepending skill instructions zj upon first accepted skill call.
5
Repeat
Repeat Order 2–4 until accepted termination or budget exhaustion. Exhaustion does not constitute Protocol Completion.
6
Score
Compute final trajectory return Ri using Equation 7 .
Appendix
Table 4: Runtime execution order for one skill-contract trajectory.
Milestone
Scope
Weight
Contract location
First-completion condition
Skill
Both
0.10
NEED_SKILL
The expected <skill_call> is accepted.
Plan
Both
0.10
NEED_PLAN
A substantive <think> is accepted.
Action
Both
0.20
NEED_ACTION
The first admissible <code> , <search> , or <research> passes its task adapter.
Observation
Both
0.20
NEED_ACTION → NEED_EVIDENCE_THINK
Math execution produces nonempty output, or Search returns at least one nonempty document.
Evidence reasoning
O-Math
0.20
NEED_EVIDENCE_THINK
A substantive <think> is accepted.
Evidence reasoning
O-Search
0.10
NEED_EVIDENCE_THINK
A substantive <think> is accepted.
Appendix
Table 5: Milestone set Mj and first-completion weights used for verified progress credit. Each contract location maps the credit event to the shared phase or transition in Table 3 . The task reward is a separate binary term. Table 6 defines substantive thought. For Search, evidence organization through <summary> is optional for accepted termination. Its verified progress credit encourages this optional action, so permitted trajectories need not receive equal progress credit.
Phase
O-Math
O-Search
NEED_SKILL
Accept the expected code-interpreter-protocol skill call.
Accept the expected qa-search-protocol skill call.
NEED_PLAN
Accept a substantive <think> that plans the calculation.
Accept a substantive <think> that plans retrieval.
NEED_ACTION
Accept executable Python that prints a result in <code>...</code> . The restricted interpreter returns <interpreter> . The observation is valid when execution succeeds and printed output is nonempty. The accepted action enters NEED_EVIDENCE_THINK .
Accept the first admissible query in <search> and later nonrepeated admissible queries in <research> . The frozen retriever returns <information> . The observation is valid when at least one retrieved document is nonempty. The accepted action enters NEED_EVIDENCE_THINK .
NEED_EVIDENCE_ THINK
Accept a substantive <think> that interprets the interpreter result. A valid observation opens continuation or answering. An invalid one returns to NEED_ACTION .
Accept a substantive <think> that assesses the retrieved evidence. A valid observation opens continuation or answering. An invalid one returns to NEED_ACTION .
CAN_CONTINUE_ OR_ANSWER
Repeat <code>...</code> and re-enter evidence reasoning, or accept a nonempty boxed answer and enter DONE .
Repeat retrieval with <research> and re-enter evidence reasoning. Optionally accept one substantive <summary> as evidence organization, with no retriever call, no observation, and no phase change. Alternatively, accept a nonempty <answer> and enter DONE .
DONE
Terminal. Score the accepted answer by normalized numerical match.
Terminal. Score the accepted answer by dataset-normalized exact match.
Appendix
Table 6: Task-specific instantiation of Stepj by phase. Table 3 defines the shared transitions, and Table 5 locates verified progress credit. For <think> , “substantive” requires at least 12 characters after whitespace normalization, including at least 6 alphanumeric characters. These are form-based checks and do not assess reasoning quality.
Condition
GPU
TP
Rounds
P/R
S/P
U/R
Batch
O-Math / Math interventions
4
2
200
72
5
5
72
O-Search / Search interventions
4
2
200
512
5
10
256
O-Joint
4
2
400
72
5
5
72
Recovery-coefficient comparison
4
2
400
72
5
5
72
Appendix
Table 7: Online training configurations. P/R is prompts per rollout, S/P is samples per prompt, and U/R is optimizer updates per rollout. Model-scale runs reuse the corresponding task profile, and the Feedback + Credit conditions in RQ2 reuse the main single-task runs.
Domain
Condition
Feedback
Credit
Outcome
PCR
Search
B00 Outcome-only
0
0
0.00
0.00
Search
B10 Feedback-only
1
0
0.00
0.00
Search
B01 Credit-only
0
1
29.28
98.69
Search
B11 Feedback + Credit
1
1
30.11
98.90
Math
B00 Outcome-only
0
0
0.00
0.00
Math
B10 Feedback-only
1
0
0.00
0.00
Appendix
Table 8: Training-time feedback-credit factorial on Qwen2.5-3B. Each condition is trained independently from the same base checkpoint within its domain. The task reward and contract runtime are fixed, while Feedback toggles contract-state feedback and Credit toggles verified progress credit. Outcome-only uses the contract runtime, in which rejected actions do not receive a penalty or terminate the trajectory, but only a final answer accepted after all required phases can receive the task reward. Task Outcome and PCR are unweighted averages over MATH and GSM8K for Math and 2WikiMultiHopQA and HotpotQA for Search.
Domain
Visible skill
MATH
GSM
2Wiki
Hotpot
Avg.
PCR
Hint Rate
Math
Original
48.62
80.59
–
–
64.61
97.29
5.59
Math
Removed
42.76
74.68
–
–
58.72
91.34
11.70
Math
Paraphrased
43.92
75.59
–
–
59.75
94.36
7.27
Math
Wrong Search
1.48
2.35
–
–
1.92
4.71
99.69
Search
Original
–
–
29.87
30.34
30.11
98.90
0.99
Search
Removed
–
–
22.57
18.51
20.54
63.81
14.86
Appendix
Table 9: Inference-time visible skill instruction interventions on frozen Qwen2.5-3B single-task checkpoints. Only the visible skill instructions zj change. Paraphrased rewrites wording, plan descriptions, and transition language while preserving action tags, action order, environment-owned observations, repeated-action permission, and the termination format. Appendix E.2.1 gives the exact Math and Search bodies. Outcome is pass@1 for Math and exact match for Search.