Action-Space Shaping for LLM Agents: Measuring and Mitigating Tool-Schema Bias
Organizations: University of Cambridge · Yinwang · Huawei · LARK, HKUST (GZ) · HKUST
Abstract
Large Language Models (LLMs) have shown strong performance on tool-use agentic tasks when given a fixed tool schema. Yet a tool schema is not the action space of an agent; it is merely one interface representation of it. The same executable action can be exposed through many different, functionally equivalent tool definitions, and an agent that has truly learned a task should behave consistently across them. We show that current agents often do not, a phenomenon we term schema bias. To study this systematically, we introduce an executable transformation framework that rewrites a native tool schema using nine operators, including merging and splitting tools, altering how a single tool is expressed, and distributing one action across several dependent calls. The tasks, executable actions, and reachable states remain fixed, so any change in success is attributable to the interface alone. Evaluating eleven LLMs, including two closed models, on up to 32 schema variants, we ask how large schema bias is, how it manifests, whether the difficulty of a schema variant can be predicted without a full evaluation, and whether training removes it. We find that schema bias is substantial even for the newest models: success rates range from complete failure to 97% depending solely on the schema. To reliably estimate schema difficulty, it requires running a small sample of the target queries. Training repairs a schema variant only when that variant appears in the training data.
Figures & tables
| Family | Operators | Mechanisms |
|---|---|---|
| Single-call | merge split | Tool-set partitioning . Re-designs tool boundaries over the native operations: several functions are fused into one tool behind a discriminator, or one tool is split by a condition on its arguments. Representative methods (Figure 2 ): fully merged and class dispatch for merge; fully split and interval split for split. |
| nest arguments rename strip descriptions reorder arguments | Per-tool representation . Rewrites the surface form of a single tool expression: parameter nesting structure, tool and parameter identifiers, presence of descriptions, and parameter order. Representative methods: nested args for nest and namespaced names for rename; strip descriptions and reorder arguments have a single method each. | |
| Multi-call | transaction reference resolution schema discovery | Cross-call protocol . Spreads one native action over several dependent calls whose state must be carried across turns: transaction replaces one native action with multiple intermediate calls. reference resolution requires an earlier additional reference resolving call. schema discovery requires schema query before tool invoke. |
| Failure mode | Rule | What went wrong |
| Livelock | The episode hits the turn limit. | The agent loops on retries or clarifications without committing an action. |
| Transaction handle | A transaction flag fires: a write or commit names a transaction that was never opened, an opened transaction is never committed, a commit precedes the required writes, or a transaction is committed twice. | The begin / write / commit protocol is broken. |
| Rejected and stuck | At least one call was rejected as off-schema and fewer native actions executed than required. | The agent calls a name or argument the schema does not expose and cannot recover. |
| Wrong class | A call is routed through a dispatcher or namespace of the wrong class. | The operation is chosen from the wrong group. |
| Under-execution | Fewer native actions executed than the task requires, with no rejection. | Required steps are skipped. |
| Over-execution | More native actions executed than the task requires. | Extra or repeated actions are taken. |
| Training-free | SFT | RL (GRPO) | |||||
| Variant | Base | Instruction | Decoding | Native | Mixed | Native | Mixed |
| Native | 0.77 | – | 0.00 | +0.02 | 0.08 | +0.05 | +0.04 |
| Tool-set partitioning | |||||||
| Fully merged | 0.36 | +0.28 | +0.29 | +0.11 | +0.34 | +0.11 | +0.43 |
| Class dispatch | 0.66 | +0.01 | 0.02 | 0.26 | 0.00 | +0.05 | +0.05 |
| Fully split | 0.76 | 0.00 | 0.00 | +0.05 | 0.06 | +0.06 | +0.05 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Domain | Representative operations | Ops | RQ4 |
| Car cabin controller | windows, sunroof, seats, cabin climate | 18 | train |
| Home HVAC / thermostat | modes, zones, holds, fan settings | 13 | train |
| Multi-room media player | playback, sources, queues, rooms | 15 | train |
| Smart kitchen appliances | oven, stove, timers, presets | 13 | train |
| Smart laundry | wash/dry cycles, soil and water levels | 13 | train |
| Home security | locks, alarms, sensors, access codes | 13 | train |
| Native catalog | |
|---|---|
| Domains | 12 |
| Native tools | 168 (13–18 per domain) |
| Arguments | 397 (1–4 per tool, mean 2.4; all required) |
| enumerated (string enum) | 261 (2–19 values, mean 3.8) |
| free string / integer / boolean | 43 / 92 / 1 |
| Queries |
| model | native | flat | nested | union |
|---|---|---|---|---|
| Qwen3-4B | .774 | .367 | .324 | .036 |
| Qwen3-30B-A3B | .774 | .070 | .044 | .031 |
| Qwen2.5-7B | .531 | .357 | .341 | .174 |
| Qwen2.5-14B | .774 | .554 | .361 | .389 |
| Qwen2.5-32B | .763 | .754 | .430 | .547 |
| Llama-3.1-8B | .739 | .688 | .454 | .497 |
| model | ratio | flat | nested | union | mixed | [95% CI] |
|---|---|---|---|---|---|---|
| Qwen3-4B | .612 | .383 | .583 | .411 | [ ] | |
| .600 | .417 | .555 | .352 | [ ] | ||
| Qwen3-30B-A3B | .621 | .185 | .624 | .296 | [ ] | |
| .619 | .212 | .580 | .240 | [ ] | ||
| Qwen2.5-7B | .333 | .133 | .309 | .202 | [ ] | |
| .332 | .100 | .316 | .144 | [ ] |
| model | class-sem. | class-neut. | random | anti-class | [95% CI] | [95% CI] |
|---|---|---|---|---|---|---|
| Qwen3-4B | .662 | .655 | .233 | .211 | [ ] | [ ] |
| Qwen3-8B | .636 | .532 | .198 | .198 | [ ] | [ ] |
| Qwen3-30B-A3B | .668 | .683 | .283 | .260 | [ ] | [ ] |
| Qwen2.5-7B | .176 | .098 | .036 | .031 | [ ] | [ ] |
| Qwen2.5-14B | .523 | .460 | .172 | .156 | [ ] | [ ] |
| Qwen2.5-32B | .730 | .741 | .442 | .439 | [ ] | [ ] |
| airline | retail | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| model | native | merged | nested | txn. | ref. | disc. | native | merged | nested | txn. | ref. | disc. |
| Qwen2.5-7B | .24 | .28 | .14 | .18 | .30 | .30 | .08 | .04 | .06 | .06 | .04 | .03 |
| Qwen2.5-14B | .20 | .44 | .16 | .22 | .42 | .42 | .33 | .05 | .30 | .17 | .16 | .04 |
| Qwen2.5-32B | .32 | .46 | .30 | .22 | .44 | .42 | .42 | .05 | .42 | .28 | .20 | .05 |
| Qwen3-4B | .26 | .46 | .28 | .30 | .38 | .40 | .28 | .09 | .20 | .13 | .12 | .06 |
| Qwen3-30B-A3B | .30 | .46 | .38 | .28 | .40 | .42 | .39 | .05 | .40 | .37 | .28 | .04 |
| model | transaction | schema discovery | opaque names | fully merged |
|---|---|---|---|---|
| Qwen3-4B | .97 / .28 | .81 / .07 | 1.00 / .75 | .33 / .36 |
| Qwen3-30B-A3B | 1.00 / .72 | .89 / .11 | 1.00 / .76 | .00 / .07 |
| Qwen2.5-7B | .00 / .25 | .67 / .03 | 1.00 / .27 | .36 / .35 |
| Qwen2.5-14B | .03 / .10 | .33 / .05 | 1.00 / .74 | .22 / .54 |
| Qwen2.5-32B | .00 / .77 | .92 / .29 | 1.00 / .77 | .25 / .75 |
| Llama-3.1-8B | .03 / .00 | .78 / .04 | 1.00 / .71 | 1.00 / .71 |
| Variant | Instruction |
|---|---|
| Fully merged | All operations go through a single dispatcher tool. Set operation to the operation you need and pass each of its arguments under the key <operation>::<argument> . |
| Class dispatch | Each class has one dispatcher tool named after the class. Call the tool of the right class, set operation to the method you need, and pass each argument under the key <method>::<argument> . |
| Fully split | The values of enumerable arguments are part of the tool names. Call the tool whose name contains the values you need and pass only the remaining arguments. |
| Interval split | Numeric arguments are split into ranges across several tools. Call the tool whose name gives the range that contains your value, and still pass the value itself. |
| Nested args | Some arguments are grouped into nested objects. Put each argument inside the object that the tool’s parameter schema defines for it. |
| Namespaced names | Tool names are prefixed with their class name. Call each tool by its full prefixed name, using the class you intend to act on. |
| method | configuration |
|---|---|
| LoRA SFT | LoRA ( Hu et al., 2022 ) rank , , dropout ; targets q,k,v,o,gate,up,down_proj ; loss on assistant turns only. Two epochs, learning rate with cosine decay and warmup, AdamW ( Loshchilov and Hutter, 2019 ) , bf16, batch size with gradient accumulation . Maximum sequence length k tokens (shorter cutoffs truncated transaction traces), k for the seven-variant mixture (the namespaced-names and fully split tool lists reach about k tokens), and k for the phrasing-and-volume study. Seeds to . |
| GRPO | On-policy GRPO in slime ( THUDM, 2025 ) with Megatron-LM ( Shoeybi et al., 2019 ) (tensor parallel ) and SGLang ( Zheng et al., 2024 ) rollouts. Learning rate (constant), eight samples per prompt, prompts per rollout, global batch , KL coefficient in (low-variance estimator), PPO-style clipping ( Schulman et al., 2017 ) / , no entropy bonus, rollout temperature , at most generated tokens per turn and turns per episode; episodes and rollout contexts are capped at k tokens, or k for the seven-variant mixture. Every run uses rollouts ( episodes, optimizer steps); an SFT-then-RL run takes about GPU-hours on GH200 GPUs. |
| Training data | Native | Fully split | Fully merged | Ref. res. | Trans. | Schema disc. | |
| Untrained | 0.77 | 0.76 | 0.37 | 0.77 | 0.28 | 0.07 | |
| SFT | |||||||
| native only | 2 | 0.79 | 0.81 | 0.47 | 0.64 | 0.18 | 0.02 |
| transaction only | 2 | 0.81 | 0.79 | 0.75 | 0.83 | 0.57 | 0.35 |
| schema discovery only | 2 | 0.61 | 0.55 | 0.59 | 0.57 | 0.39 | 0.74 |
| mixed | 3 | 0.70 | 0.71 | 0.69 | 0.74 | 0.50 | 0.70 |
| held out | evaluated on | result |
|---|---|---|
| Domains | transaction and schema discovery, four held-out domains | SFT keeps 43% (transaction) and 67% (schema discovery) of its gain; RL keeps 65% and 71% |
| Schema changes (SFT) | new combination / stronger version / new operator / class-level structure | / / / points (trained changes ) |
| Benchmarks | -bench task-weighted / -bench retail, transaction rewrite / BFCL state-independent | , / , / , points |