Large Language Models (LLMs) have shown strong performance on tool-use agentic tasks when given a fixed tool schema. Yet a tool schema is not the action space of an agent; it is merely one interface representation of it. The same executable action can be exposed through many different, functionally equivalent tool definitions, and an agent that has truly learned a task should behave consistently across them. We show that current agents often do not, a phenomenon we term schema bias. To study this systematically, we introduce an executable transformation framework that rewrites a native tool schema using nine operators, including merging and splitting tools, altering how a single tool is expressed, and distributing one action across several dependent calls. The tasks, executable actions, and reachable states remain fixed, so any change in success is attributable to the interface alone. Evaluating eleven LLMs, including two closed models, on up to 32 schema variants, we ask how large schema bias is, how it manifests, whether the difficulty of a schema variant can be predicted without a full evaluation, and whether training removes it. We find that schema bias is substantial even for the newest models: success rates range from complete failure to 97% depending solely on the schema. To reliably estimate schema difficulty, it requires running a small sample of the target queries. Training repairs a schema variant only when that variant appears in the training data.
Figures & tables
Figure 1: Schema variants, schema bias, and training mitigation. (A) All schema variants Si are just different representations of the same action space Ω : a call in any of them decodes to the same native action a∗ and reaches the same final state. Each space is shown with an abbreviated schema for one action. (B) Schema bias is large and model-specific. Success rates of four LLMs under four representative variants; the dashed line is each model’s native-schema score. The same variant can match native for one model and collapse to zero for another (full grid in Figure 2 ). (C) Qwen3-4B after SFT or RL on native-only or variant-mixture data. Hatched segments are gains (green) or losses (red). Native-only data leaves most of the bias under either method; RL on variant-mixture data mitigates bias while keeping the native score (§ 8 ).
Family
Operators
Mechanisms
Single-call
merge split
Tool-set partitioning . Re-designs tool boundaries over the native operations: several functions are fused into one tool behind a discriminator, or one tool is split by a condition on its arguments. Representative methods (Figure 2 ): fully merged and class dispatch for merge; fully split and interval split for split.
Per-tool representation . Rewrites the surface form of a single tool expression: parameter nesting structure, tool and parameter identifiers, presence of descriptions, and parameter order. Representative methods: nested args for nest and namespaced names for rename; strip descriptions and reorder arguments have a single method each.
Multi-call
transaction reference resolution schema discovery
Cross-call protocol . Spreads one native action over several dependent calls whose state must be carried across turns: transaction replaces one native action with multiple intermediate calls. reference resolution requires an earlier additional reference resolving call. schema discovery requires schema query before tool invoke.
Table 1: The schema operator space, grouped by what each mechanism changes. Each operator is a class of mechanism; a schema variant fixes its method (Eq. 1 ).
Figure 2: The model-by-schema performance grid: exact success and failure mode composition. Each cell is one model (column) on one verified-equivalent schema (row), 2,748 identical tasks. The number is the exact success rate (%), represented as the white portion of composition bar. The coloured segments are the absolute rates of the automatically detected failure modes analysed in § 6 . Row labels are the representative variant names following Table 1 . The two rightmost columns are closed models evaluated through their APIs. The remaining variants are reported in Appendix Figure 4 .
Figure 3: Merging is generally more harmful than splitting. Left: colored lines show each model’s change from its native-schema success rate. The black line is the median across the nine models. The x-axis represents ratio of function count between variant and native schema. We sample ten variants to span the granularity ladder from full merge to full split. Each ratio is the query-weighted mean across 12 schema domains. Right: a compact and abbreviated schema example shows how merging and enum splitting transform the native schema.
Failure mode
Rule
What went wrong
Livelock
The episode hits the turn limit.
The agent loops on retries or clarifications without committing an action.
Transaction handle
A transaction flag fires: a write or commit names a transaction that was never opened, an opened transaction is never committed, a commit precedes the required writes, or a transaction is committed twice.
The begin / write / commit protocol is broken.
Rejected and stuck
At least one call was rejected as off-schema and fewer native actions executed than required.
The agent calls a name or argument the schema does not expose and cannot recover.
Wrong class
A call is routed through a dispatcher or namespace of the wrong class.
The operation is chosen from the wrong group.
Under-execution
Fewer native actions executed than the task requires, with no rejection.
Required steps are skipped.
Over-execution
More native actions executed than the task requires.
Extra or repeated actions are taken.
Table 2: The seven failure modes. Every failed episode is assigned the first mode whose rule fires. The first four surface as explicit runtime errors or routing errors; the last three are silent failure: the episode ends with the wrong set of native actions or the wrong argument values.
Training-free
SFT
RL (GRPO)
Variant
Base
Instruction
Decoding
Native
Mixed
Native
Mixed
Native
0.77
–
0.00
+0.02
− 0.08
+0.05
+0.04
Tool-set partitioning
Fully merged
0.36
+0.28
+0.29
+0.11
+0.34
+0.11
+0.43
Class dispatch
0.66
+0.01
− 0.02
− 0.26
0.00
+0.05
+0.05
Fully split
0.76
0.00
0.00
+0.05
− 0.06
+0.06
+0.05
Table 3: Change in success from each mitigation on Qwen3-4B over the twelve representative variants of Figure 2 . Base is the untrained model; every other column show delta success rate from it. changes of at least 0.03 (about 2.5 standard errors) are shaded green or red, darker for larger changes. Instruction appends variant-specific calling convention describing instructions to input prompt (Table 11 ); Decoding forces the first call to be schema-valid via constrained decoding. SFT and RL columns are results averaged from three seeds; Mixed data rotate over seven schema variants. Red frames mark the variants each training set contains.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Domain
Representative operations
Ops
RQ4
Car cabin controller
windows, sunroof, seats, cabin climate
18
train
Home HVAC / thermostat
modes, zones, holds, fan settings
13
train
Multi-room media player
playback, sources, queues, rooms
15
train
Smart kitchen appliances
oven, stove, timers, presets
13
train
Smart laundry
wash/dry cycles, soil and water levels
13
train
Home security
locks, alarms, sensors, access codes
13
train
Appendix
Table 4: The twelve synthetic domains, native operation counts, and the intervention-training split used in RQ4. Each domain defines a grouped native tool schema (system context, callable operations, and typed arguments); the full specification is released with the artifact.
Native catalog
Domains
12
Native tools
168 (13–18 per domain)
Arguments
397 (1–4 per tool, mean 2.4; all required)
enumerated (string enum)
261 (2–19 values, mean 3.8)
free string / integer / boolean
43 / 92 / 1
Queries
Appendix
Table 5: Statistics of the controlled synthetic environment: the native tool catalog (top) and the 2,748 evaluation queries (bottom). Every query is a single user utterance whose target is one to six native calls within one domain; there are no multi-turn dialogues. Word counts are for the colloquial queries used in all experiments, with the templated source query in parentheses.
Figure 4: Scores of the remaining methods of each operator , in the format of Figure 2 (nine models, 2,748 queries). Row labels give the operator and method (Table 1 ): partial merges along the granularity ladder (function count relative to native), the nested and union argument structures of the fully merged schema, the three grouping rules at a fixed group size, the membership and naming controls of class dispatch together with the whole-catalog control, the split ladder, and opaque names.
model
native
flat
nested
union
Qwen3-4B
.774
.367
.324
.036
Qwen3-30B-A3B
.774
.070
.044
.031
Qwen2.5-7B
.531
.357
.341
.174
Qwen2.5-14B
.774
.554
.361
.389
Qwen2.5-32B
.763
.754
.430
.547
Llama-3.1-8B
.739
.688
.454
.497
Appendix
Table 6: Merge penalty by argument-structure encoding. Exact success on the fully merged schema (one dispatcher per domain, the same member sets as the fully merged row in Figure 2 ) under three typed single-call argument structures, next to each model’s native score from the same run. 2,748 queries per cell; a native re-run after the grid stays within 0.3 points. Bold marks each model’s lowest argument structure.
model
ratio
flat
nested
union
mixed
H [95% CI]
Qwen3-4B
0.3×
.612
.383
.583
.411
−.115 [ −.124,−.105 ]
0.5×
.600
.417
.555
.352
−.172 [ −.183,−.161 ]
Qwen3-30B-A3B
0.3×
.621
.185
.624
.296
−.181 [ −.192,−.170 ]
0.5×
.619
.212
.580
.240
−.231 [ −.242,−.219 ]
Qwen2.5-7B
0.3×
.333
.133
.309
.202
−.056 [ −.067,−.046 ]
0.5×
.332
.100
.316
.144
−.105 [ −.115,−.095 ]
Appendix
Table 7: Matched convention-mixture test on all nine models. At each granularity target, every arm exposes the same dispatcher groups, names, and function count; the three homogeneous arms use one argument structure throughout, and the mixed score averages three balanced rotations of the structures across groups. H is mixed minus the mean of the three homogeneous arms, with a paired task bootstrap 95% interval (2,748 tasks).
model
class-sem.
class-neut.
random
anti-class
M [95% CI]
N [95% CI]
Qwen3-4B
.662
.655
.233
.211
+.421 [ .404,.438 ]
+.007 [ −.004,.019 ]
Qwen3-8B
.636
.532
.198
.198
+.335 [ .317,.352 ]
+.103 [ .085,.123 ]
Qwen3-30B-A3B
.668
.683
.283
.260
+.400 [ .384,.415 ]
−.015 [ −.025,−.004 ]
Qwen2.5-7B
.176
.098
.036
.031
+.062 [ .052,.073 ]
+.078 [ .064,.091 ]
Qwen2.5-14B
.523
.460
.172
.156
+.288 [ .271,.306 ]
+.063 [ .046,.081 ]
Qwen2.5-32B
.730
.741
.442
.439
+.299 [ .285,.313 ]
−.012 [ −.020,−.002 ]
Appendix
Table 8: Whole-catalog class-grouping controls. All arms expose the 168 native operations through twelve flat dispatchers with the same group-size profile. M (membership) is class-neutral minus the mean of five random-neutral partitions; N (naming) is class-semantic minus class-neutral. Paired task bootstrap 95% intervals over 2,748 tasks. Qwen3-8B, not one of the nine main models, is included in this control only.
airline
retail
model
native
merged
nested
txn.
ref.
disc.
native
merged
nested
txn.
ref.
disc.
Qwen2.5-7B
.24
.28
.14
.18
.30
.30
.08
.04
.06
.06
.04
.03
Qwen2.5-14B
.20
.44
.16
.22
.42
.42
.33
.05
.30
.17
.16
.04
Qwen2.5-32B
.32
.46
.30
.22
.44
.42
.42
.05
.42
.28
.20
.05
Qwen3-4B
.26
.46
.28
.30
.38
.40
.28
.09
.20
.13
.12
.06
Qwen3-30B-A3B
.30
.46
.38
.28
.40
.42
.39
.05
.40
.37
.28
.04
Appendix
Table 9: Schema bias on τ2 -bench: task success of nine models under six verified-equivalent schemas in the airline (50 tasks) and retail (114 tasks) domains. In the cell marked † , one task’s model request timed out on all four attempts and the task is scored as a failure; every other cell finished all tasks without error. Fully merged collapses the domain’s tools into one dispatcher and nested args groups each tool’s arguments into objects (the same operators as the corresponding rows of Figure 2 ); the three protocols are the cross-call operators of Table 1 . Bold marks each model’s best schema per domain.
Figure 5: Schema-form gap per model and variant. Change in success (percentage points) when off-schema calls are executed as the native action they decode to rather than rejected. Positive values (blue) are schema-form failures; negative values (red) mean rejection with retry does better than executing the decoded call. The colour scale saturates at +25 ; the printed number is the true gap. Nested args is a near-zero control; interval split uses one threshold.
model
transaction
schema discovery
opaque names
fully merged
Qwen3-4B
.97 / .28
.81 / .07
1.00 / .75
.33 / .36
Qwen3-30B-A3B
1.00 / .72
.89 / .11
1.00 / .76
.00 / .07
Qwen2.5-7B
.00 / .25
.67 / .03
1.00 / .27
.36 / .35
Qwen2.5-14B
.03 / .10
.33 / .05
1.00 / .74
.22 / .54
Qwen2.5-32B
.00 / .77
.92 / .29
1.00 / .77
.25 / .75
Llama-3.1-8B
.03 / .00
.78 / .04
1.00 / .71
1.00 / .71
Appendix
Table 10: Isolated-call compliance vs. full-task exact success (compliance / exact) on the variants with the largest task-conditional gaps. Compliance is the fraction of 36 schema-derived single-operation probes executed with no rejection and no task query; exact is the full 2,748-query score of Figure 2 .
Variant
Instruction
Fully merged
All operations go through a single dispatcher tool. Set operation to the operation you need and pass each of its arguments under the key <operation>::<argument> .
Class dispatch
Each class has one dispatcher tool named after the class. Call the tool of the right class, set operation to the method you need, and pass each argument under the key <method>::<argument> .
Fully split
The values of enumerable arguments are part of the tool names. Call the tool whose name contains the values you need and pass only the remaining arguments.
Interval split
Numeric arguments are split into ranges across several tools. Call the tool whose name gives the range that contains your value, and still pass the value itself.
Nested args
Some arguments are grouped into nested objects. Put each argument inside the object that the tool’s parameter schema defines for it.
Namespaced names
Tool names are prefixed with their class name. Call each tool by its full prefixed name, using the class you intend to act on.
Appendix
Table 11: Training-free interventions used in § 8 . Instruction appends the listed sentence for each variant to the default system prompt, which already requires using only the provided tools; the native schema needs no instruction. Decoding sets vLLM tool_choice=required on the first turn only, so the first completion is grammar-constrained to a schema-valid tool name and arguments; later turns use the default auto policy.
method
configuration
LoRA SFT
LoRA ( Hu et al., 2022 ) rank r=32 , α=64 , dropout 0.05 ; targets q,k,v,o,gate,up,down_proj ; loss on assistant turns only. Two epochs, learning rate 2×10−4 with cosine decay and 3% warmup, AdamW ( Loshchilov and Hutter, 2019 ) , bf16, batch size 1 with gradient accumulation 8 . Maximum sequence length 16 k tokens (shorter cutoffs truncated transaction traces), 32 k for the seven-variant mixture (the namespaced-names and fully split tool lists reach about 22 k tokens), and 8 k for the phrasing-and-volume study. Seeds 42 to 45 .
GRPO
On-policy GRPO in slime ( THUDM, 2025 ) with Megatron-LM ( Shoeybi et al., 2019 ) (tensor parallel 2 ) and SGLang ( Zheng et al., 2024 ) rollouts. Learning rate 1×10−6 (constant), eight samples per prompt, 24 prompts per rollout, global batch 96 , KL coefficient in {0,0.01,0.1} (low-variance estimator), PPO-style clipping ( Schulman et al., 2017 ) 0.2 / 0.28 , no entropy bonus, rollout temperature 1.0 , at most 2,048 generated tokens per turn and 12 turns per episode; episodes and rollout contexts are capped at 11 k tokens, or 32 k for the seven-variant mixture. Every run uses 100 rollouts ( 19,200 episodes, 200 optimizer steps); an SFT-then-RL run takes about 24 GPU-hours on GH200 GPUs.
Appendix
Table 12: Training hyperparameters (§ 8 ). The base model is Qwen3-4B-Instruct-2507 unless noted; the Qwen2.5-7B run uses the same LoRA recipe.
Training data
n
Native
Fully split
Fully merged
Ref. res.
Trans.
Schema disc.
Untrained
0.77
0.76
0.37
0.77
0.28
0.07
SFT
native only
2
0.79
0.81
0.47
0.64
0.18
0.02
transaction only
2
0.81
0.79
0.75
0.83
0.57
0.35
schema discovery only
2
0.61
0.55
0.59
0.57
0.39
0.74
mixed
3
0.70
0.71
0.69
0.74
0.50
0.70
Appendix
Table 13: Training ablations on Qwen3-4B: exact success on the full 2,748-query set for six of the twelve representative variants, by training method, training data, and KL coefficient ( n = seeds averaged). Mixed data rotate over fully merged, transaction, reference resolution, schema discovery, opaque names, and a one-argument enum split. The last three RL rows change the environment or the reward: rejections without an explanation, no rejections (invalid calls are executed as their decoded native call), and a reward of exact success only. The SFT → RL rows start GRPO from the mixed SFT adapter.
held out
evaluated on
result
Domains
transaction and schema discovery, four held-out domains
SFT keeps 43% (transaction) and 67% (schema discovery) of its gain; RL keeps 65% and 71%
Schema changes (SFT)
new combination / stronger version / new operator / class-level structure
Table 14: Transfer of training gains on Qwen3-4B (Appendix G.5 ). Changes are in exact success relative to the untrained model; in the last row, pairs of values are two seeds of SFT followed by RL.
Large language model (LLM) agents increasingly interleave natural language reasoning with external tools such as web search and code execution. These tool-use policies are often optimized via reinforcement learning (RL), which can amplify spurious correlations in the training data. In this work, we study when and why RL-trained agents learn shortcut tool-selection policies: invoking tools based on superficial prompt cues rather than genuine task requirements. We construct controlled synthetic environments combining factual question answering and mathematical reasoning tasks, and inject cues that are strongly correlated with specific tools during training but causally irrelevant to tool necessity. Across counterfactual evaluations where cues are present but the associated tools are not required, agents exhibit substantial shortcut behavior, with spurious tool invocation rates increasing by up to 39 percent. However, shortcut formation is not universal: across the conditions we test, it arises only when the agent has already learned to use the target tool reliably, suggesting that task competence, rather than dataset imbalance alone, is a key factor in shortcut learning. A swapped-cue analysis further shows that semantic alignment between cues and tools substantially amplifies this effect. To mitigate these failures, we introduce a dense, decision-level reward in which an LLM judge evaluates the necessity of each tool call. This tool-necessity reward effectively suppresses cue-driven tool use while preserving task performance, providing a practical approach to improving the robustness of LLM agent tool-use policies.
Yiwei Yang, Haoxiang Zhang, Bingbing Wen +6
University of Washington · University of California San Diego · Stanford University
Learning to complete tasks in unfamiliar environments with unknown rules remains a key challenge for LLM agents. Current LLM agents often record their discoveries in prose, which may not provide a compact, explicit account of how the environment works. Inspired by how scientists organize observations into testable, predictive theories, we introduce Schema, an agent harness that organizes learning and action through interactive program induction. The LLM agent decides what to investigate and how to act, expressing its evolving understanding of the environment as executable programs. The harness consists of a persistent program workspace and a small set of interfaces for checking these programs against the interaction history, planning within them, and executing plans under step-by-step verification. Schema raises ARC-AGI-3 RHAE from 58.7% to 99.2% with the same base model, solves 100% of the public DiG-bench games, and reaches the median performance of the top-50 human players on MazeBench. Extensive analysis shows the effectiveness of Schema in unknown mechanism discovery, and ablations confirm the contribution of each component.
Guanning Zeng, Jiani Wang, Wenjie Ma +8
Carnegie Mellon University · Impossible Research · UC Berkeley
Production agent frameworks (OpenAI Function Calling, Anthropic Tool Use, MCP) transmit tool schemas as JSON, a format designed for machine parsing, not for interpretation by language models. For small models (4B-14B), this protocol mismatch accounts for the majority of tool-use failure at production catalog sizes. We present TSCG, a deterministic tool-schema compiler that resolves this mismatch at the API boundary, converting JSON schemas into token-efficient structured text without model access, fine-tuning, or runtime search. TSCG combines eight composable operators with a formal compression bound (>=51% on well-formed schemas). On TSCG-Agentic-Bench (about 19,000 calls, 12 models, 5 scenarios), TSCG restores Phi-4 14B from 0% to 84.4% accuracy at 20 tools (90.3% at 50 tools) and achieves 108-181% accuracy-retained ratio across three models on BFCL. Format-versus-compression decomposition (R^2=0.88 -> 0.03) establishes representation change as the dominant mechanism. Per-operator isolation across three frontier models reveals three distinct operator-response profiles: operator-hungry (Opus 4.7), operator-sensitive (GPT-5.2), and operator-robust (Sonnet 4), providing per-model deployment guidance. Scaling experiments show accuracy advantages persisting on heavy production MCP schemas (+5.0 pp at about 10,500 input tokens) despite saturation on light synthetic catalogs, with 52-57% token savings throughout. The synthetic benchmark generalizes to real MCP schemas within 0.1 accuracy points. TSCG ships as a 1,200-line zero-dependency TypeScript package.