cs.SESep 27, 2026

What Does a Skill Actually Do? Estimands and Evaluation Validity for Tool and Skill Use in LLM Agents: A Critical Review

Authors: Shuyang Zhang

Organizations: The Hong Kong Polytechnic University, Hong Kong, China

Abstract

Reported improvements from tools and reusable skills in large language model agents refer to different comparisons. This critical narrative review examines what these evaluations estimate and which conclusions their designs support. The review checks the roles of one hundred cited papers and extracts focal evaluation designs in detail from thirty-five studies. Targeted readings of thirty-five additional published or accepted studies broaden coverage of tool creation, memory, interactive benchmarks, reliability, and risk. Designs are characterized by treatment contrast, target population, outcome, budget constraint, summary measure, and identification assumptions. Analytic decompositions and counterexamples show that pairing runs on the same task does not itself identify an invocation effect when evaluation conditions on a trigger within the treated run. Paired gain and regression counts describe discordance under the coupling protocol rather than the share of tasks whose expected outcomes worsen. Total effects of deploying a module answer a different question from efficiency under a common budget. Comparisons across studies distinguish curated skill provision from retriever replacement, task populations from triggered subsets, and preparation costs from marginal usage costs. Publication status and reading depth are recorded. The review provides a methodological synthesis and a reporting checklist to help align claims about tools and skills with the comparisons their evaluation designs support.

Figures & tables

Explore similar work

CardsList
  1. A Framework for Evaluating Agentic Skills at Scale

    Jun 16, 2026Maksim Shaposhnikov, Nicolas Fortuin, Simon Stipcich +3Large Language Model AgentsSkills

  2. Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents

    Aug 12, 2026Gen Dong, Yanjie Gao, Liqun Li +3Large Language Model AgentsTask Success Rate

  3. Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents

    Sep 1, 2026Seonghyeon Cho, Chanjun ParkSkills