SkillFM: Generating Skills for LLM Agents via Latent Flow Matching
Organizations: Nanyang Technological University · University of Edinburgh · Meta
Abstract
Textual skills provide reusable guidance for large language model agents, but existing approaches often rely on manually curated skill banks or reinforcement learning with indirect and delayed feedback. We introduce SkillFM (Skill Flow Matching), a generative framework that synthesizes task-conditioned textual skills directly without test-time skill retrieval. Our framework combines a codec for encoding and reconstructing textual skills in a continuous latent space with a conditional flow model trained using improved MeanFlow. At inference time, the learned velocity field enables single-step latent sampling, and an LLM-based decoder converts the sampled representation into textual guidance for a frozen downstream agent. We evaluate the framework on embodied tasks, question answering, and web shopping. On ALFWorld and Search-QA, our method achieves the best overall performance among the compared vector-based skill approaches. Our analyses further demonstrate that latent skill generation is an effective alternative to retrieval-based skill augmentation. Our code and training skill libraries are available at https://github.com/lulushang999/SkillFM.
Figures & tables
| Method | ALFWorld task | SR | Step | |||||
|---|---|---|---|---|---|---|---|---|
| Pick | Look | Clean | Heat | Cool | Pick2 | |||
| Seen split | ||||||||
| Vanilla ∗ | 82.9 | 46.2 | 18.5 | 37.5 | 32.0 | 29.2 | 43.6 | 35.0 |
| Full SFT ∗ | 82.9 | 38.5 | 70.4 | 43.8 | 24.0 | 37.5 | 53.6 | 36.0 |
| In-Context Skill ∗ | 85.7 | 69.2 | 70.4 | 31.3 | 12.0 | 33.3 | 52.9 | 30.8 |
| SHINE ∗ | 88.6 | 69.2 | 59.3 | 6.25 | 36.0 | 70.8 | 59.3 | 29.8 |
| Single-hop | Multi-hop | |||||||
|---|---|---|---|---|---|---|---|---|
| Method | NQ † | Triv ⋆ | Pop ⋆ | Hotp † | 2WK ⋆ | MuS ⋆ | Bam ⋆ | Avg |
| Vanilla | 25.2 | 50.6 | 35.2 | 26.2 | 26.8 | 4.2 | 28.8 | 28.06 |
| CoT | 19.2 | 50.8 | 20.0 | 22.4 | 26.2 | 5.0 | 37.6 | 24.48 |
| Few-Shot | 34.6 | 57.6 | 39.4 | 30.0 | 25.8 | 6.0 | 17.6 | 31.65 |
| R1-Instruct | 27.0 | 55.0 | 31.6 | 26.6 | 33.6 | 5.8 | 34.4 | 30.11 |
| RAG | 39.0 | 64.0 | 45.0 | 32.4 | 21.2 | 6.8 | 27.2 | 34.43 |
| ALFWorld | Search-QA | WebShop | |||
|---|---|---|---|---|---|
| Skill source | Seen | Unseen | Single-hop | Multi-hop | SR |
| Embedding retrieval | 38.57 | 30.60 | 42.20 | 27.45 | 15.20 |
| Ours | 74.29 | 77.61 | 45.73 | 31.82 | 18.40 |
| (percentage points) | +35.71 | +47.01 | +3.53 | +4.37 | +3.20 |
| Training examples | 100% | 75% | 50% |
|---|---|---|---|
| 500 | 78.10% | 64.96% | 63.50% |
| 1,000 | 79.56% | 72.26% | 68.98% |
| 3,553 | 83.21% | 78.83% | 64.23% |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Domain | Training data | Codec training | Flow training |
|---|---|---|---|
| ALFWorld | 3,553 fixed training chains | Encoder/decoder LoRA and projector training | Step 10,000; training-domain holdout selection |
| Single-hop QA | NQ; 4,000 pairs | Batch 48; 26 epochs | Batch 144 to step 10,000, then 2,304 to step 30,000 |
| Multi-hop QA | 2,723 pairs | Batch 48; 24 epochs; selected epoch 9 | Batch 192; 30,000 steps; selected step 20,000 |
| WebShop | 2,000 pairs | Batch 24; 24 epochs | Batch 144; selected step 30,000 |
| Domain | Output tokens | Actions | Skill use and evidence retrieval |
|---|---|---|---|
| ALFWorld | 8,192 | 50 | One skill; one rollout |
| Single-hop QA | 2,048 | 4 | Top-3 evidence retrieval |
| Multi-hop QA | 8,192 | 4 | Top-3 evidence retrieval; grounded hybrid |
| WebShop | 4,096 | 30 | One skill; one rollout |
| Method | SR | Score | Method | SR | Score |
|---|---|---|---|---|---|
| ReAct | 19.50 | 46.20 | MemP | 6.40 | 25.30 |
| Reflexion | 28.80 | 58.10 | SimpleMem | 8.59 | 33.20 |
| Mem0 | 2.00 | 23.90 | MemRL | 9.20 | 29.50 |
| ExpeL | 11.20 | 30.90 | EvolveR | 17.60 | 42.50 |
| Ours | 20.00 | 55.35 | |||
| Actor | NQ | Triv | Pop | Hotp | 2WK | MuS | Bam | Avg |
|---|---|---|---|---|---|---|---|---|
| Qwen3-4B | 37.4 | 57.0 | 40.4 | 33.2 | 37.2 | 11.0 | 37.6 | 36.26 |
| Qwen3-8B | 38.8 | 62.0 | 41.4 | 38.4 | 43.2 | 11.0 | 44.8 | 39.94 |
| Qwen3-32B | 39.2 | 66.4 | 41.8 | 42.6 | 46.2 | 16.0 | 48.8 | 43.00 |
| Actor | SR (%) | Score |
|---|---|---|
| Qwen3-4B | 14.60 | 53.36 |
| Qwen3-8B | 20.00 | 55.35 |
| Qwen3-32B | 21.60 | 51.88 |
| Skills | Source success | Seen | Unseen | Overall |
|---|---|---|---|---|
| 500 | 100% | 75.00% | 81.34% | 78.10% |
| 500 | 75% | 58.57% | 71.64% | 64.96% |
| 500 | 50% | 62.86% | 64.18% | 63.50% |
| 1,000 | 100% | 76.43% | 82.84% | 79.56% |
| 1,000 | 75% | 68.57% | 76.12% | 72.26% |
| 1,000 | 50% | 65.00% | 73.13% | 68.98% |
| Observable pattern | Seen | Unseen |
|---|---|---|
| Failed episodes | 25/140 (17.86%) | 21/134 (15.67%) |
| Never executed a take action | 16 | 13 |
| Took objects, but only of another category | 4 | 6 |
| Acquired the correct category, but did not finish | 5 | 2 |
| At least one inadmissible action | 12 | 9 |
| Prefix length | Epoch 1 | Epoch 8 | Epoch 16 | Epoch 24 |
|---|---|---|---|---|
| 4 | 1.631252 | 0.007894 | 0.004181 | 0.001952 |
| 8 | 1.516875 | 0.008462 | 0.004854 | 0.002112 |
| 16 | 1.231258 | 0.007750 | 0.003986 | 0.001594 |
| 32 | 1.328236 | 0.007550 | 0.003797 | 0.001281 |
| Updates | Training loss | Holdout Chamfer | Mean angle | Angle P95 |
|---|---|---|---|---|
| 10,000 (selected) | 5.670 | 0.014282 | 0.09467 | 0.28164 |
| 20,000 | 4.315 | 0.014569 | 0.09415 | 0.28416 |
| 30,000 | 3.827 | 0.014731 | 0.09498 | 0.28319 |