What Does a Skill Actually Do? Estimands and Evaluation Validity for Tool and Skill Use in LLM Agents: A Critical Review
Organizations: The Hong Kong Polytechnic University, Hong Kong, China
Abstract
Reported improvements from tools and reusable skills in large language model agents refer to different comparisons. This critical narrative review examines what these evaluations estimate and which conclusions their designs support. The review checks the roles of one hundred cited papers and extracts focal evaluation designs in detail from thirty-five studies. Targeted readings of thirty-five additional published or accepted studies broaden coverage of tool creation, memory, interactive benchmarks, reliability, and risk. Designs are characterized by treatment contrast, target population, outcome, budget constraint, summary measure, and identification assumptions. Analytic decompositions and counterexamples show that pairing runs on the same task does not itself identify an invocation effect when evaluation conditions on a trigger within the treated run. Paired gain and regression counts describe discordance under the coupling protocol rather than the share of tasks whose expected outcomes worsen. Total effects of deploying a module answer a different question from efficiency under a common budget. Comparisons across studies distinguish curated skill provision from retriever replacement, task populations from triggered subsets, and preparation costs from marginal usage costs. Publication status and reading depth are recorded. The review provides a methodological synthesis and a reporting checklist to help align claims about tools and skills with the comparisons their evaluation designs support.
Figures & tables
| Aspect | Sources | Evaluation question |
|---|---|---|
| Tool learning, creation, and retrieval | [ toolformer , react , toolllm , gorilla , toolret , reinvoke , toolkengpt , creator , craft , latm ] | Is the comparison about tool availability, learned calling, generated tools, or replacement of a retrieval component? |
| Skill and memory construction | [ awm , asi , voyager , reflexion , expel ] | Is the object a fixed library, within-task adaptation, or transfer of experience across tasks? |
| Benchmarks, outcomes, and protocols | [ appworld , stabletoolbench , swebench , webarena , mind2web , workarena , apibank , toolsandbox , taubench , agentboard , agentbench , toolemu , bfcl , osworld ] | Which tasks, histories, environments, success rules, and risks does the score represent? |
| Frontier component and trajectory studies | [ skillfollowing , worththeirtokens , regressiontax , skillsbench , skillapt , radeg , shadowing , confgated , replaygap , vasudev2026 , menu , capabilitypages , tabagent , betterturns ] | What do trigger conditioning, paired flips, gates, and replay establish in the reported protocols? |
| Evaluation methodology | [ agentsmatter , huang2026leaderboard ] | Which resource constraints and target populations make system comparisons interpretable? |
| Related reviews | [ toolsurvey , yehudai2026 , mohammadi2025 , nageshwaran2026 , kehkashan2026 , wang2026tse ] | How does this component-level analysis relate to tool-learning, benchmarking, and trajectory-analysis syntheses? |
| Axis | Values found in the literature |
|---|---|
| Treatment contrast | Availability (module on vs. off); component swap (retriever, shortlist head, menu constructor); activation policy (load vs. abstain, gate on vs. off); library composition (size, content, admission rule) |
| Target population | All tasks; pre-treatment subgroup (domain, difficulty tier); post-treatment subgroup (tasks where retrieval was triggered); states within trajectories |
| Outcome | Success or reward; tokens, calls, latency; exposure to risky candidates; intermediate failures |
| Budget constraint | None; ex-ante cap with a stated allocation rule; approximate matching of realized spending |
| Summary measure | Difference in means; paired discordance counts; share of tasks whose expected outcome falls; state-level value contrast |
| Identification | Randomized or paired runs over a task distribution; stated coupling between arms; principal ignorability for post-treatment strata; sequential ignorability and positivity for state-level contrasts from logs |
| Quantity | Contrast; population | Supports | Does not support | Examples |
|---|---|---|---|---|
| Component metric (Recall@k, nDCG) | None; labeled query pool | Ranking quality on that pool | Any effect on agent outcomes; transfer to other query sources | [ toolret , reinvoke ] |
| Retrieved vs. skipped contrast | None; post-treatment groups | Description of where the agent retrieves | Any causal effect; groups differ in difficulty | criticized in [ skillfollowing ] |
| Total availability effect | Module on vs. off; all tasks | Deploying the module as specified | Resource efficiency; which path produced the effect | [ skillsbench , regressiontax , vasudev2026 ] |
| Component-swap effect | Component vs. ; all tasks, or fixed evidence paths | Replacing by in that agent | Effect of skills vs. no skills; value of an additional action | [ menu , capabilitypages , tabagent , confgated ] |
| Budget-constrained comparison | Configurations under a common cap; all tasks | Choice between configurations at that budget | Total effect without the cap; exactness if the cap is approximate | [ worththeirtokens ] |
| Composition effect | Library versions; all tasks | Effect of growing or curating the library | Effects of individual skills; later library states | [ shadowing , awm , asi ] |
| Study / venue | Comparison | Interpretive boundary |
|---|---|---|
| ToolRet [ toolret ] ; Findings ACL 2025 | Retrieved vs. oracle toolsets; trained vs. untrained retrievers with downstream agents | Supports tested replacements, not a general mapping from recall to success |
| AWM [ awm ] ; ICML 2025 | Workflow memory vs. baseline agents; offline and online variants | Adaptation package and task sequence matter; observed token overhead is not a common resource cap |
| ASI [ asi ] ; COLM 2025 | Static agent, text-skill AWM, and programmatic skills; verification/representation ablations | Main gain bundles induction, verification, and action-space changes; one high-level step may contain several primitive actions |
| AI Agents That Matter [ agentsmatter ] ; TMLR 2025 | Agent architectures vs. retry baselines; cost–accuracy trade-offs | Resource comparison under stated tasks and prices, not a skill-specific invocation effect |
| StableToolBench [ stabletoolbench ] ; Findings ACL 2024 | API and evaluator changes; a fixed solvable-task subset | Changes the measurement environment and population; repeated grading is not repeated execution |
| Study | Quantity (review interpretation) | Condition that limits interpretation |
|---|---|---|
| Skill Following [ skillfollowing ] [A] | Overall paired effect; trigger-conditioned paired contrast (RAE) | Restricts to tasks where the enabled run returned a skill; shared-seed coupling with unknown selection terms |
| Regression Tax [ regressiontax ] [P] | Total availability effect; paired discordance | One run per task and condition; variance not estimated |
| Budget study [ worththeirtokens ] [A] | Comparison under an approximate budget constraint | Budget matched through a step cap, not exactly; control arm also adds pruning |
| Skill shadowing [ shadowing ] [P] | Composition effect with counterfactual decomposition | Population restricted to task–model pairs whose skills each raised pass rate by at least 4 percentage points; bounds need a monotonicity assumption |
| SkillsBench [ skillsbench ] [A] | Total availability effect of curated per-task bundles | Task-specific bundles supplied in the evaluation; not retrieval from a shared library |
| SkillApt [ skillapt ] [P] | Task-conditional load vs. abstain utility | State represented by question features; complete-case analysis (111 of 113 states) |
| Item | What to report | Applies to |
|---|---|---|
| Question and contrast | Deployment, component choice, or state-level decision; what each arm sees | All designs |
| Target population | Task distribution and whether it matches the query distribution of any retrieval benchmark used | All designs |
| Unit, coupling, runs | Unit of comparison; how arms are coupled (independent, shared seed, branched state); runs per task and arm; run-to-run variance | All designs with stochastic agents |
| Identification | Assumptions under which the reported summary is causal, or a statement that it is protocol-specific | Designs that condition on post-treatment events or use logs |
| Trigger stratification | Trigger rate and paired contrasts on triggered and non-triggered tasks, with the trigger definition | Designs in which retrieval or invocation is optional |
| Discordance and degradation | Gain and regression counts under the stated protocol; if harm to tasks is claimed, confidently and possibly degraded shares from simultaneous per-task intervals, or a stated hierarchical model | Paired designs; degradation shares only for harm claims |