HEXIS: Compiling Skills into Extended Finite State Machines
Abstract
Agent skills provide reusable knowledge and instructions, yet agents must repeatedly infer how to apply them and which operation should follow. This couples task reasoning with control decisions, allowing prescribed steps to be omitted or applied incorrectly. We introduce HEXIS, which compiles agent skills into extended finite state machines that separate knowledge from control flow. Skill knowledge is incorporated into local instructions that guide reasoning and generation within states. The machine records execution progress and intermediate results, while explicit transition conditions determine subsequent operations. Our incremental compiler first maps skill clauses and tool interfaces to state operations, local instructions, data bindings, and transitions. It then aligns development traces with existing states to identify missing operations and dependencies. These are incorporated by adding or reusing states and refining their connections. Updates are accepted only after static checks and replay of the current and all previously accepted traces. Across four benchmarks and four executors, HEXIS improves success over Skill + ReAct by 16.1 percentage points on average. Qwen3.8-27B reduces execution tokens by 38.4-88.9% across benchmarks.
Figures & tables
| Benchmark | Method | Qwen3.6-flash | GLM-4.7-FlashX | Qwen3.5-9B | Qwen3.8-27B |
|---|---|---|---|---|---|
| Spreadsheet Bench | Skill + ReAct | 45.6% | 22.8% | 33.3% | 56.1% |
| AWM | 61.4% | 31.6% | 35.1% | 59.6% | |
| ReasoningBank | 43.9% | 15.7% | 33.3% | 63.2% | |
| SkillOpt | 59.6% | 40.4% | 47.4% | 66.7% | |
| AFlow | 36.8% | 22.8% | 38.6% | 73.7% | |
| HEXIS | 75.4% | 47.4% | 38.6% | 71.9% |
| Skill document | Execution | Success (%) | Tokens (k) |
|---|---|---|---|
| Original | Skill + ReAct | 45.6% | 213 |
| Original | HEXIS | 75.4% | 223 |
| SkillOpt | Skill + ReAct | 59.6% | 257 |
| SkillOpt | HEXIS | 84.2% | 69 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Benchmark | Skill + ReAct | HEXIS Fable 5.1 | HEXIS Sonnet 5 low |
|---|---|---|---|
| SpreadsheetBench | 45.6% | 75.4% | 71.9% |
| LiveMath | 45.0% | 76.7% | 72.7% |
| DABench | 78.4% | 82.4% | 80.4% |
| LongSeal | 7.8% | 21.6% | 19.6% |
| Benchmark | States | Edges | Recovery | Fallback | Retry events | Steps | Revisits | |
|---|---|---|---|---|---|---|---|---|
| (tasks) | (tasks) | (total) | (mean) | (mean) | ||||
| SpreadsheetBench | 57 | 17 | 35 | 0 (0.0%) | 0 (0.0%) | 0 | 24.56 | 10.60 |
| LiveMath | 120 | 12 | 19 | 3 (2.5%) | 1 (0.8%) | 3 | 9.04 | 0.11 |
| DABench | 51 | 15 | 29 | 0 (0.0%) | 0 (0.0%) | 0 | 13.57 | 1.57 |
| LongSeal | 51 | 16 | 30 | 1 (2.0%) | 1 (2.0%) | 1 | 30.20 | 17.31 |
| State | Assigned operation | Ordered outgoing rules |
|---|---|---|
| s1 | Model: analyze the question, hypotheses, option support, and relative strength; write analysis . | Always |
| s2 | Model: select an option using request , analysis , and ; write answer_letter and justification . | ; else |
| s2m | Judge: check whether the selected option covers the stated conclusion; write . | , ; else |
| s3 | Model: generate write_cmd to write the selected letter in the required boxed format to output_path , using edit_log when available. | ; ; else |
| s4 | Tool: execute bash with command=write_cmd ; bind the returned stdout to edit_log . | ; else , |
| s5 | Tool: execute read with filePath=output_path ; bind the returned stdout to file_content . | ; ; else , |
| Benchmark | Success | Full compliance | ||||
|---|---|---|---|---|---|---|
| Skill + ReAct | Prompt-only | HEXIS | Skill + ReAct | Prompt-only | HEXIS | |
| SpreadsheetBench | 45.6% | 63.2% | 75.4% | 96.5% | 75.4% | 100.0% |
| LiveMath | 45.0% | 33.9% | 76.7% | 45.5% | 82.6% | 91.7% |
| DABench | 78.4% | 80.4% | 82.4% | 88.2% | 96.1% | 98.0% |
| LongSeal | 7.8% | 11.8% | 21.6% | 33.3% | 19.6% | 96.1% |
| Resource | License | URL |
| SpreadsheetBench ( Ma et al., 2024 ) | CC BY-SA 4.0 | https://huggingface.co/datasets/KAKA22/SpreadsheetBench |
| LiveMathematicianBench ( He et al., 2026 ) | Available online | https://huggingface.co/datasets/LiveMathematicianBench/ |
| InfiAgent-DABench ( Hu et al., 2024 ) | Apache 2.0 | https://github.com/InfiAgent/InfiAgent |
| SealQA ( Pham et al., 2026 ) | Apache 2.0 | https://huggingface.co/datasets/vtllms/sealqa |
| SigLeak skills ( Geng et al., 2026 ) | Available online | https://anonymous.4open.science/r/SigLeak-D1DB |
| Pandas Pro ( Jeffallan, 2026 ) | MIT | https://github.com/jeffallan/claude-skills |