MCRI: A Four-Dimensional Framework for Analyzing and Evaluating Agent Skills
Abstract
As agents evolve from single-tool systems into modular, composite architectures, skills are becoming an important mechanism for capability development and distribution. However, the academic community lacks a structured framework for systematically analyzing and evaluating skills. Drawing on information gain and behavioral constraint, we propose the four-dimensional MCRI Framework and operationalize it as MCRI-Eval, a large language model-based evaluation method. We evaluate MCRI-Eval using 63,812 public skills from the OpenClaw skill Hub, with 58,275 skill-conditioned model executions across BigCodeBench, BFCL-Fundamental, and Mind2Web. MCRI-Eval scores are positively associated with community popularity signals and achieve the highest downstream ranking agreement among the evaluated methods. MCRI-Eval also improves top-1 skill selection across all three benchmarks: compared with the strongest baseline on each benchmark, the skills selected by MCRI-Eval advance by 17.7, 22.8, and 19.6 percentile points in downstream performance rank on BigCodeBench, BFCL-Fundamental, and Mind2Web, respectively. These results indicate that MCRI-Eval provides a useful pre-execution signal for prioritizing promising skills before costly execution-based evaluation.
Figures & tables
| Signal | Top 3% ( ) | Sig. | Top 1% ( ) | Sig. | Top 0.5% ( ) | Sig. |
|---|---|---|---|---|---|---|
| Stars | 0.166 (GPT-5.4-nano) | 0.213 (GPT-5.4-nano) | 0.302 (MiMo-v2.5) | |||
| Downloads | 0.125 (GPT-5-mini) | 0.211 (GPT-5.4-nano) | 0.293 (MiMo-v2.5) | |||
| All-time installs | 0.197 (GPT-5.4-nano) | 0.329 (GPT-5.4-nano) | 0.393 (MiMo-v2.5) |
| Signal | MiMov2.5Pro | MiMov2.5 | GPT5.4-nano | GPT5-mini |
|---|---|---|---|---|
| Stars | ||||
| Downloads | ||||
| Comments | ||||
| All-time installs |
| Variant | MiMo v2.5Pro | MiMo v2.5 | GPT5.4- nano | GPT5- mini |
|---|---|---|---|---|
| Full | ||||
| w/o | ||||
| w/o | ||||
| w/o | ||||
| w/o |
| Model | AUC | Accuracy | F1 |
|---|---|---|---|
| MiMo-v2.5-Pro | 0.758 | 70.5% | 0.644 |
| MiMo-v2.5 | 0.719 | 62.9% | 0.581 |
| GPT-5.4-nano | 0.819 | 76.2% | 0.747 |
| GPT-5-mini | 0.716 | 68.6% | 0.593 |
| Average | 0.753 | 69.6% | 0.641 |
| Method | C-index | Spearman | -value |
|---|---|---|---|
| Baseline | 0.5101 | 0.0248 | 0.7765 |
| Full-text | 0.5522 | 0.1496 | 0.0856 |
| MCRI-Eval | 0.5777 | 0.2119 | 0.0143 |
| Method | C-index | Spearman | -value |
|---|---|---|---|
| Baseline | 0.5686 | 0.1824 | 0.0988 |
| Full-text | 0.5648 | 0.1891 | 0.0869 |
| MCRI-Eval | 0.6237 | 0.3355 | 0.0019 |
| Method | C-index | Spearman | -value |
|---|---|---|---|
| Baseline | 0.5781 | 0.1931 | 0.09687 |
| Full-text | 0.6065 | 0.2754 | 0.01678 |
| MCRI-Eval | 0.6418 | 0.3611 | 0.001458 |
| Benchmark | SD | SD | SD | SD | Mean SD |
|---|---|---|---|---|---|
| Mind2Web | 0.113 | 0.119 | 0.118 | 0.103 | 0.113 |
| BigCodeBench | 0.139 | 0.147 | 0.163 | 0.141 | 0.148 |
| BFCL-Fundamental | 0.144 | 0.107 | 0.119 | 0.158 | 0.132 |
| Benchmark | Baseline | Full-text | MCRI | ||
|---|---|---|---|---|---|
| Mind2Web | 0.119 | 0.197 | 0.085 | 28.5% | 56.7% |
| BigCodeBench | 0.476 | 0.396 | 0.061 | 87.1% | 84.5% |
| BFCL-Fundamental | 0.095 | 0.280 | 0.078 | 18.4% | 72.3% |
| Cross-bench. avg. | 0.230 | 0.291 | 0.075 | 67.5% | 74.3% |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Benchmark | Method | C-ind. | ||
|---|---|---|---|---|
| BFCL | Baseline | 0.5101 | 0.0248 | 0.7765 |
| FullTextBaseline | 0.5522 | 0.1496 | 0.0856 | |
| heuristic_a | 0.5682 | 0.1846 | 0.0334 | |
| heuristic_b | 0.5699 | 0.1849 | 0.0331 | |
| MCRI-Eval | 0.5777 | 0.2119 | 0.0143 | |
| BCB | Baseline | 0.5686 | 0.1824 | 0.0988 |
| Benchmark | Method | Top-1 | Top-3 | Top-5 | Top-10 | Top-20 |
|---|---|---|---|---|---|---|
| BFCL | Baseline | 79.60 / 51.82 | 79.60 / 51.82 | 79.60 / 51.82 | 79.20 / 43.64 | 78.80 / 35.91 |
| FullTextBaseline | 78.00 / 24.62 | 79.56 / 50.88 | 79.47 / 49.02 | 79.73 / 54.62 | 79.37 / 46.91 | |
| heuristic_a | 78.67 / 33.33 | 79.33 / 46.21 | 79.73 / 54.62 | 79.20 / 43.64 | 79.33 / 46.21 | |
| heuristic_b | 78.67 / 33.33 | 80.00 / 60.23 | 80.00 / 60.23 | 79.13 / 42.35 | 79.60 / 51.82 | |
| MCRI-Eval | 80.67 / 74.62 | 79.56 / 50.88 | 80.00 / 60.23 | 80.13 / 63.11 | 79.28 / 45.25 | |
| BigCodeBench | Baseline | 53.33 / 35.98 | 55.78 / 75.81 | 55.33 / 71.34 | 55.33 / 71.34 | 55.00 / 65.85 |
| Benchmark | R1 / R2 / R3 | Mean / Rank |
|---|---|---|
| BFCL-Fundamental | 0.800 / 0.800 / 0.800 | 80.00% / 54.00 |
| BigCodeBench | 0.560 / 0.540 / 0.500 | 53.33% / 35.98 |
| Mind2Web | 0.519 / 0.533 / 0.513 | 52.17% / 6.05 |
| Method | C-index | -value | Top-1 | Top-3 | Top-5 | Top-10 | |
|---|---|---|---|---|---|---|---|
| Baseline (LLM) | 0.5101 | 0.0248 | 0.7765 | 79.60 / 51.82 | 79.60 / 51.82 | 79.60 / 51.82 | 79.20 / 43.64 |
| FullTextBaseline (LLM) | 0.5522 | 0.1496 | 0.0856 | 78.00 / 24.62 | 79.56 / 50.88 | 79.47 / 49.02 | 79.73 / 54.62 |
| MCRI-Eval (LLM) | 0.5777 | 0.2119 | 0.0143 | 80.67 / 74.62 | 79.56 / 50.88 | 80.00 / 60.23 | 80.13 / 63.11 |
| Stars | 0.5107 | 0.0315 | 0.7186 | 76.67 / 15.15 | 76.89 / 16.41 | 78.93 / 38.48 | 78.87 / 37.20 |
| Downloads | 0.5161 | 0.0467 | 0.5938 | 76.67 / 15.15 | 77.33 / 18.94 | 78.93 / 38.48 | 78.93 / 38.48 |
| All-time installs | 0.4430 | -0.1557 | 0.0736 | 76.00 / 10.61 | 77.33 / 18.94 | 78.93 / 38.48 | 78.40 / 29.85 |
| Method | C-index | -value | Top-1 | Top-3 | Top-5 | Top-10 | |
|---|---|---|---|---|---|---|---|
| Baseline (LLM) | 0.5686 | 0.1824 | 0.0988 | 53.33 / 35.98 | 55.78 / 75.81 | 55.33 / 71.34 | 55.33 / 71.34 |
| FullTextBaseline (LLM) | 0.5648 | 0.1891 | 0.0869 | 54.67 / 60.37 | 54.22 / 50.61 | 55.20 / 69.15 | 54.60 / 58.90 |
| MCRI-Eval (LLM) | 0.6237 | 0.3355 | 0.0019 | 56.00 / 78.05 | 56.00 / 78.05 | 56.40 / 82.07 | 55.67 / 74.70 |
| Stars | 0.5093 | 0.0321 | 0.7733 | 54.67 / 60.37 | 54.00 / 45.73 | 53.47 / 37.93 | 55.03 / 66.40 |
| Downloads | 0.4762 | -0.0646 | 0.5620 | 53.33 / 35.98 | 54.00 / 45.73 | 53.60 / 39.88 | 53.93 / 44.76 |
| All-time installs | 0.5238 | 0.0679 | 0.5418 | 53.33 / 35.98 | 52.89 / 31.10 | 53.60 / 39.88 | 54.80 / 62.56 |
| Method | C-index | -value | Top-1 | Top-3 | Top-5 | Top-10 | |
|---|---|---|---|---|---|---|---|
| Baseline (LLM) | 0.5781 | 0.1931 | 0.0969 | 54.88 / 49.55 | 54.88 / 49.55 | 54.90 / 50.14 | 55.14 / 57.55 |
| FullTextBaseline (LLM) | 0.6065 | 0.2754 | 0.0168 | 56.23 / 79.73 | 56.14 / 77.93 | 55.30 / 62.43 | 55.51 / 66.89 |
| MCRI-Eval (LLM) | 0.6418 | 0.3611 | 0.001458 | 57.39 / 99.32 | 56.71 / 92.79 | 56.14 / 78.11 | 55.54 / 67.43 |
| Stars | 0.4508 | -0.1450 | 0.2144 | 56.52 / 87.84 | 53.72 / 23.65 | 54.55 / 39.59 | 54.41 / 35.61 |
| Downloads | 0.4916 | -0.0123 | 0.9169 | 56.52 / 87.84 | 54.98 / 52.48 | 54.20 / 30.41 | 54.26 / 31.89 |
| All-time installs | 0.5206 | 0.0578 | 0.6226 | 56.52 / 87.84 | 53.72 / 23.65 | 54.20 / 30.41 | 54.61 / 41.35 |
| Dim. | Name | Functional layer | Definition |
|---|---|---|---|
| M | Metadata | Identification | Task definition, functional scope, applicable scenarios, triggering conditions, and declarative input/output specifications that tell the host agent what the skill does and when it should be used. |
| C | Constraints | Constraint | Rules that explicitly restrict the model’s admissible behavior, including prohibitions, preconditions, invariants, handling policies, and boundary conditions. |
| R | Resources | Information | Task-specific information beyond parametric knowledge, including templates, scripts, example data, reference documents, concrete API details, and previously observed error traces. |
| I | Instructions | Interface | Executable guidance for invoking internal scripts, external tools, or sub-skills and interacting with the environment, including call forms, orchestration order, and subtask decomposition. |
| Dim. | Output name | Sub-item weights |
|---|---|---|
| M | task_spec_clarity | task_definition (0.25), use_conditions (0.25), scope_boundaries (0.20), declarative_io (0.20), consistency (0.10) |
| C | constraint_sufficiency | explicit_rules (0.25), decision_coverage (0.25), operability (0.20), failure_boundaries (0.20), non_conflict (0.10) |
| R | information_gain | extra_parametric_value (0.30), specificity (0.25), critical_coverage (0.20), reusability (0.15), validity (0.10) |
| I | interface_completeness | callable_entities (0.20), invocation_contract (0.25), orchestration (0.25), state_artifacts (0.15), environment_handoff (0.15) |
| Method | Reference information | Additional input |
|---|---|---|
| Baseline | A statement that the experiment measures the LLM’s unaided ability to assess skill quality. | skill name, summary, and benchmark task description. |
| FullTextBaseline | The same reference statement as Baseline. | Baseline input plus the complete skill text. |
| MCRI-Eval | The MCRI theoretical description, including information gain, behavioral constraint, and the four dimension definitions. | The mean Stage-A values , , , and , accompanied by the instruction that dimension importance can vary by benchmark and that the values are advisory rather than a fixed aggregation formula. |