As agents evolve from single-tool systems into modular, composite architectures, skills are becoming an important mechanism for capability development and distribution. However, the academic community lacks a structured framework for systematically analyzing and evaluating skills. Drawing on information gain and behavioral constraint, we propose the four-dimensional MCRI Framework and operationalize it as MCRI-Eval, a large language model-based evaluation method. We evaluate MCRI-Eval using 63,812 public skills from the OpenClaw skill Hub, with 58,275 skill-conditioned model executions across BigCodeBench, BFCL-Fundamental, and Mind2Web. MCRI-Eval scores are positively associated with community popularity signals and achieve the highest downstream ranking agreement among the evaluated methods. MCRI-Eval also improves top-1 skill selection across all three benchmarks: compared with the strongest baseline on each benchmark, the skills selected by MCRI-Eval advance by 17.7, 22.8, and 19.6 percentile points in downstream performance rank on BigCodeBench, BFCL-Fundamental, and Mind2Web, respectively. These results indicate that MCRI-Eval provides a useful pre-execution signal for prioritizing promising skills before costly execution-based evaluation.
Figures & tables
Figure 1: Constraint strength spectrum from natural language to skills and formal programs.
Figure 2: The MCRI framework.
Figure 3: Example of the MCRI Framework.
Signal
Top 3% ( N=1,823 )
Sig.
Top 1% ( N=606 )
Sig.
Top 0.5% ( N=305 )
Sig.
Stars
0.166 (GPT-5.4-nano)
0.213 (GPT-5.4-nano)
0.302 (MiMo-v2.5)
Downloads
0.125 (GPT-5-mini)
0.211 (GPT-5.4-nano)
0.293 (MiMo-v2.5)
All-time installs
0.197 (GPT-5.4-nano)
0.329 (GPT-5.4-nano)
0.393 (MiMo-v2.5)
Table 1: Strongest Spearman correlations at different filtering thresholds. The corresponding scoring model is shown in parentheses.
Signal
MiMov2.5Pro
MiMov2.5
GPT5.4-nano
GPT5-mini
Stars
+0.203∗∗∗
+0.152∗∗
+0.213∗∗∗
+0.134∗
Downloads
+0.165∗∗
+0.183∗∗
+0.211∗∗∗
+0.082ns
Comments
+0.156∗∗
+0.152∗∗
+0.270∗∗∗
+0.208∗∗∗
All-time installs
+0.280∗∗∗
+0.267∗∗∗
+0.329∗∗∗
+0.226∗∗∗
Table 2: Spearman correlations at the top-1% threshold across four scoring models. Statistical significance is denoted by ∗p<0.05 , ∗∗p<0.01 , and ∗∗∗p<0.001 ; ns denotes a nonsignificant result.
Variant
MiMo v2.5Pro
MiMo v2.5
GPT5.4- nano
GPT5- mini
Full
+0.203∗∗∗
+0.152∗∗
+0.213∗∗∗
+0.134∗
w/o M
+0.199∗∗∗
+0.169∗∗
+0.182∗∗
+0.124∗
w/o R
+0.036ns
+0.057ns
+0.172∗∗
+0.039ns
w/o C
+0.169∗∗
+0.140∗
+0.172∗∗
+0.096ns
w/o I
+0.192∗∗∗
+0.098ns
+0.162∗∗
+0.127∗
Table 3: Dimension-level ablation results for stars and all-time installs across four scoring models.
Model
AUC
Accuracy
F1
MiMo-v2.5-Pro
0.758
70.5%
0.644
MiMo-v2.5
0.719
62.9%
0.581
GPT-5.4-nano
0.819
76.2%
0.747
GPT-5-mini
0.716
68.6%
0.593
Average
0.753
69.6%
0.641
Table 4: Extreme-class identification within the top-1% samples. High- and low-signal groups correspond to the top and bottom 15%, respectively.
Method
C-index
Spearman ρ
p -value
Baseline
0.5101
0.0248
0.7765
Full-text
0.5522
0.1496
0.0856
MCRI-Eval
0.5777
0.2119
0.0143
Table 5: Downstream evaluation results on BFCL-Fundamental. Top- k results report downstream performance (Perf.) and the corresponding percentile score within the candidate set.
Method
C-index
Spearman ρ
p -value
Baseline
0.5686
0.1824
0.0988
Full-text
0.5648
0.1891
0.0869
MCRI-Eval
0.6237
0.3355
0.0019
Table 6: Downstream evaluation results on BigCodeBench. Top- k results report downstream performance (Perf.) and the corresponding percentile score within the candidate set.
Method
C-index
Spearman ρ
p -value
Baseline
0.5781
0.1931
0.09687
Full-text
0.6065
0.2754
0.01678
MCRI-Eval
0.6418
0.3611
0.001458
Table 7: Downstream evaluation results on Mind2Web. Top- k results report Step Success Rate (Perf.) and the corresponding percentile score within the candidate set.
Benchmark
M SD
C SD
R SD
I SD
Mean SD
Mind2Web
0.113
0.119
0.118
0.103
0.113
BigCodeBench
0.139
0.147
0.163
0.141
0.148
BFCL-Fundamental
0.144
0.107
0.119
0.158
0.132
Table 8: Standard deviations of the four MCRI dimension scores.
Benchmark
Baseline
Full-text
MCRI
ΔBase
ΔFull
Mind2Web
0.119
0.197
0.085
28.5%
56.7%
BigCodeBench
0.476
0.396
0.061
87.1%
84.5%
BFCL-Fundamental
0.095
0.280
0.078
18.4%
72.3%
Cross-bench. avg.
0.230
0.291
0.075
67.5%
74.3%
Table 9: Score stability of the three skill evaluation methods. The final two columns report the relative reduction in standard deviation achieved by MCRI-Eval.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Method
C-ind.
ρ
p
BFCL
Baseline
0.5101
0.0248
0.7765
FullTextBaseline
0.5522
0.1496
0.0856
heuristic_a
0.5682
0.1846
0.0334
heuristic_b
0.5699
0.1849
0.0331
MCRI-Eval
0.5777
0.2119
0.0143
BCB
Baseline
0.5686
0.1824
0.0988
Appendix
Table 10: Ranking agreement of the heuristic comparison. BFCL, BCB, and M2W denote BFCL-Fundamental, BigCodeBench, and Mind2Web, respectively. MCRI-Eval attains the highest C-index and Spearman correlation on all three benchmarks.
Benchmark
Method
Top-1
Top-3
Top-5
Top-10
Top-20
BFCL
Baseline
79.60 / 51.82
79.60 / 51.82
79.60 / 51.82
79.20 / 43.64
78.80 / 35.91
FullTextBaseline
78.00 / 24.62
79.56 / 50.88
79.47 / 49.02
79.73 / 54.62
79.37 / 46.91
heuristic_a
78.67 / 33.33
79.33 / 46.21
79.73 / 54.62
79.20 / 43.64
79.33 / 46.21
heuristic_b
78.67 / 33.33
80.00 / 60.23
80.00 / 60.23
79.13 / 42.35
79.60 / 51.82
MCRI-Eval
80.67 / 74.62
79.56 / 50.88
80.00 / 60.23
80.13 / 63.11
79.28 / 45.25
BigCodeBench
Baseline
53.33 / 35.98
55.78 / 75.81
55.33 / 71.34
55.33 / 71.34
55.00 / 65.85
Appendix
Table 11: Top- k results for the heuristic comparison. Each cell reports downstream performance (%) / normalized rank score. Bold marks the highest downstream performance for each benchmark and value of k .
Benchmark
R1 / R2 / R3
Mean / Rank
BFCL-Fundamental
0.800 / 0.800 / 0.800
80.00% / 54.00
BigCodeBench
0.560 / 0.540 / 0.500
53.33% / 35.98
Mind2Web
0.519 / 0.533 / 0.513
52.17% / 6.05
Appendix
Table 12: No-skill downstream performance over three repeated runs. The final column is the percentile of the mean in the corresponding distribution of skill-injected outcomes.
Method
C-index
ρ
p -value
Top-1
Top-3
Top-5
Top-10
Baseline (LLM)
0.5101
0.0248
0.7765
79.60 / 51.82
79.60 / 51.82
79.60 / 51.82
79.20 / 43.64
FullTextBaseline (LLM)
0.5522
0.1496
0.0856
78.00 / 24.62
79.56 / 50.88
79.47 / 49.02
79.73 / 54.62
MCRI-Eval (LLM)
0.5777
0.2119
0.0143
80.67 / 74.62
79.56 / 50.88
80.00 / 60.23
80.13 / 63.11
Stars
0.5107
0.0315
0.7186
76.67 / 15.15
76.89 / 16.41
78.93 / 38.48
78.87 / 37.20
Downloads
0.5161
0.0467
0.5938
76.67 / 15.15
77.33 / 18.94
78.93 / 38.48
78.93 / 38.48
All-time installs
0.4430
-0.1557
0.0736
76.00 / 10.61
77.33 / 18.94
78.93 / 38.48
78.40 / 29.85
Appendix
Table 13: Non-LLM selector comparison on BFCL-Fundamental ( N=133 ). Top- k cells report pass rate (%) / normalized rank score. MCRI-Eval has the strongest positive global ranking agreement; the file-count Top-1 value is an isolated high-outcome selection despite its negative overall correlation.
Method
C-index
ρ
p -value
Top-1
Top-3
Top-5
Top-10
Baseline (LLM)
0.5686
0.1824
0.0988
53.33 / 35.98
55.78 / 75.81
55.33 / 71.34
55.33 / 71.34
FullTextBaseline (LLM)
0.5648
0.1891
0.0869
54.67 / 60.37
54.22 / 50.61
55.20 / 69.15
54.60 / 58.90
MCRI-Eval (LLM)
0.6237
0.3355
0.0019
56.00 / 78.05
56.00 / 78.05
56.40 / 82.07
55.67 / 74.70
Stars
0.5093
0.0321
0.7733
54.67 / 60.37
54.00 / 45.73
53.47 / 37.93
55.03 / 66.40
Downloads
0.4762
-0.0646
0.5620
53.33 / 35.98
54.00 / 45.73
53.60 / 39.88
53.93 / 44.76
All-time installs
0.5238
0.0679
0.5418
53.33 / 35.98
52.89 / 31.10
53.60 / 39.88
54.80 / 62.56
Appendix
Table 14: Non-LLM selector comparison on BigCodeBench ( N=83 ). Top- k cells report pass rate (%) / normalized rank score. Text length is predictive in this benchmark, but MCRI-Eval retains the strongest global ranking agreement.
Method
C-index
ρ
p -value
Top-1
Top-3
Top-5
Top-10
Baseline (LLM)
0.5781
0.1931
0.0969
54.88 / 49.55
54.88 / 49.55
54.90 / 50.14
55.14 / 57.55
FullTextBaseline (LLM)
0.6065
0.2754
0.0168
56.23 / 79.73
56.14 / 77.93
55.30 / 62.43
55.51 / 66.89
MCRI-Eval (LLM)
0.6418
0.3611
0.001458
57.39 / 99.32
56.71 / 92.79
56.14 / 78.11
55.54 / 67.43
Stars
0.4508
-0.1450
0.2144
56.52 / 87.84
53.72 / 23.65
54.55 / 39.59
54.41 / 35.61
Downloads
0.4916
-0.0123
0.9169
56.52 / 87.84
54.98 / 52.48
54.20 / 30.41
54.26 / 31.89
All-time installs
0.5206
0.0578
0.6226
56.52 / 87.84
53.72 / 23.65
54.20 / 30.41
54.61 / 41.35
Appendix
Table 15: Non-LLM selector comparison on Mind2Web ( N=75 ). Top- k cells report Step SR (%) / normalized rank score. MCRI-Eval provides the strongest global ranking agreement and the best Top- k outcomes.
Dim.
Name
Functional layer
Definition
M
Metadata
Identification
Task definition, functional scope, applicable scenarios, triggering conditions, and declarative input/output specifications that tell the host agent what the skill does and when it should be used.
C
Constraints
Constraint
Rules that explicitly restrict the model’s admissible behavior, including prohibitions, preconditions, invariants, handling policies, and boundary conditions.
R
Resources
Information
Task-specific information beyond parametric knowledge, including templates, scripts, example data, reference documents, concrete API details, and previously observed error traces.
I
Instructions
Interface
Executable guidance for invoking internal scripts, external tools, or sub-skills and interacting with the environment, including call forms, orchestration order, and subtask decomposition.
Appendix
Table 16: The four MCRI dimensions and their functional roles.
Table 17: Stage-A sub-item weights. The weights within each dimension sum to one.
Method
Reference information
Additional input
Baseline
A statement that the experiment measures the LLM’s unaided ability to assess skill quality.
skill name, summary, and benchmark task description.
FullTextBaseline
The same reference statement as Baseline.
Baseline input plus the complete skill text.
MCRI-Eval
The MCRI theoretical description, including information gain, behavioral constraint, and the four dimension definitions.
The mean Stage-A values M=x.xx , C=x.xx , R=x.xx , and I=x.xx , accompanied by the instruction that dimension importance can vary by benchmark and that the values are advisory rather than a fixed aggregation formula.
Appendix
Table 18: Information supplied to each Stage-B method. All other output and aggregation rules are shared.
The growth of agent skills has transformed how agentic systems are built, evaluated, and deployed. As skill libraries continue to scale, rigorous evaluation becomes critical to ensuring their utility, quality, and safety in real-world applications. Consequently, the field is undergoing an emerging paradigm shift from isolated skill creation to automated, evaluation-driven skill evolution. In this survey, we systematically examine the landscape of skill evolution and evaluation beyond foundational skill creation. We categorize evolution into four distinct paradigms, spanning execution feedback, trajectory distillation, compression, and reinforcement learning, showing how each element contributes to improving skill utility and reliability. We also provide an analysis of six skill-centric benchmark categories, identifying structural gaps in benchmark coverage, trade-offs, and metric richness to advance skill research. Finally, we identify open directions for building skill ecosystems that are generalizable, efficient, and verifiably safe. The project URL is https://github.com/Cassie07/AgentSkill_Survey
Kexin Ding, Yang Zhou, Can Jin +3
Rutgers University · University of North Carolina at Charlotte
Agent skills -- structured, reusable knowledge artifacts that augment LLM agent capabilities -- have been rapidly adopted in industry, yet their cross-domain impact and use across commercial and open-source models remain under-studied, and no reusable methodology exists for evaluating an individual skill. In this work, we present an evaluation framework that lets a skill author construct realistic tasks to rigorously assess the aspects of a skill that matter most to them, and that estimates skill utility by solving those tasks. Further, we apply our evaluation approach at scale to 500 real-world skills, generating 1,000 tasks derived from the skills' content, along with instruction-following and goal-completion scoring rubrics. Using these metrics, we evaluate how 19 agent-model configurations, both proprietary and open-source, perform on the tasks. Our results show that models vary widely in how closely they adhere to the instructions encoded in skills, leading to substantial differences in their performance gains. Furthermore, we show that access to a skill significantly changes model behavior compared to the no-skill setup, providing an essential mechanism for encoding opinionated workflows into LLM agents. We release our evaluation dataset to support future work on agent skills.
Maksim Shaposhnikov, Nicolas Fortuin, Simon Stipcich +3
Agent skills provide reusable procedural knowledge that helps agents solve specialized tasks. As their use expands, evaluating skill quality becomes increasingly important. Existing evaluations often measure skill quality by testing whether a skill improves performance on specific downstream tasks. However, a reusable skill may apply to multiple task scenarios. Downstream evaluation mainly reflects the compatibility between a skill and the evaluated task, provides only a partial view of skill quality, and does not identify which aspect of the skill should be improved. We find that general properties of the \texttt{SKILL.md} document play an important role in skill quality. To evaluate these properties, we propose \textbf{SkillEval}, an interpretable framework for document-level skill evaluation. SkillEval evaluates each property using a fixed and inspectable scoring direction, producing interpretable scores. It further measures and reduces the influence of unrelated document features, such as length and formatting, so that each score captures its intended semantic property more specifically. Specifically, SkillEval learns an interpretable direction for each quality property from controlled positive--negative skill pairs in the hidden representation space of the model, and scores a new skill by projecting its representation onto these fixed directions. We use SkillEval to evaluate skills in controlled quality tests and show that SkillEval reliably distinguishes skills of different quality. In addition, SkillEval scores closely reflect downstream task performance, providing an early indication of whether a skill is likely to help an agent complete a task. We further explore SkillEval for diagnosing weaknesses in skill documents and guiding targeted revisions. The revised skills improve the targeted properties and achieve higher pass rates on downstream tasks.
Jiahui Han, Qinuo Li, Ziheng Peng +6
1Xi’an Jiaotong University · 2Shanghai AI Laboratory · 5Harbin Institute of Technology +2