Textual skills provide reusable guidance for large language model agents, but existing approaches often rely on manually curated skill banks or reinforcement learning with indirect and delayed feedback. We introduce SkillFM (Skill Flow Matching), a generative framework that synthesizes task-conditioned textual skills directly without test-time skill retrieval. Our framework combines a codec for encoding and reconstructing textual skills in a continuous latent space with a conditional flow model trained using improved MeanFlow. At inference time, the learned velocity field enables single-step latent sampling, and an LLM-based decoder converts the sampled representation into textual guidance for a frozen downstream agent. We evaluate the framework on embodied tasks, question answering, and web shopping. On ALFWorld and Search-QA, our method achieves the best overall performance among the compared vector-based skill approaches. Our analyses further demonstrate that latent skill generation is an effective alternative to retrieval-based skill augmentation. Our code and training skill libraries are available at https://github.com/lulushang999/SkillFM.
Figures & tables
Figure 1: Motivation for SkillFM. (a) External skill reuse requires retrieval and bank maintenance, with ambiguous matches and delayed feedback. (b) SkillFM learns from a skill bank offline and generates readable guidance through one-step latent sampling and decoding, supporting reuse across frozen actors without test-time bank lookup.
Figure 2: Training and inference pipeline. We train the framework in two stages: the codec learns to encode and decode skills, and conditional flow matching learns to generate latent skills for the decoder. At inference, we reuse the trained flow model and codec to produce textual skill guidance for the agent with one flow evaluation.
Method
ALFWorld task
SR ↑
Step ↓
Pick
Look
Clean
Heat
Cool
Pick2
Seen split
Vanilla ∗
82.9
46.2
18.5
37.5
32.0
29.2
43.6
35.0
Full SFT ∗
82.9
38.5
70.4
43.8
24.0
37.5
53.6
36.0
In-Context Skill ∗
85.7
69.2
70.4
31.3
12.0
33.3
52.9
30.8
SHINE ∗
88.6
69.2
59.3
6.25
36.0
70.8
59.3
29.8
Table 1: Performance on ALFWorld in success rate (%). Results are reported on the seen and unseen splits with a per-task breakdown. Step denotes the average number of interaction steps per episode. ∗ denotes results from Yu et al. (2026) , version 3. † denotes the three-run means reported by Ouyang et al. (2026) under their evaluation protocol. The best results are highlighted in blue .
Single-hop
Multi-hop
Method
NQ †
Triv ⋆
Pop ⋆
Hotp †
2WK ⋆
MuS ⋆
Bam ⋆
Avg ↑
Vanilla
25.2
50.6
35.2
26.2
26.8
4.2
28.8
28.06
CoT
19.2
50.8
20.0
22.4
26.2
5.0
37.6
24.48
Few-Shot
34.6
57.6
39.4
30.0
25.8
6.0
17.6
31.65
R1-Instruct
27.0
55.0
31.6
26.6
33.6
5.8
34.4
30.11
RAG
39.0
64.0
45.0
32.4
21.2
6.8
27.2
34.43
Table 2: Performance on Search-QA in exact match (%). Avg is micro-averaged over all examples. Each dataset contains 500 evaluation examples, except Bamboogle, which contains 125. † and ⋆ indicate in-domain and out-of-domain datasets, respectively, under the training protocol of Yu et al. (2026) . The best results are highlighted in blue .
Figure 3: Transfer across ALFWorld actors. Success rates (%); differences in pp.
Table 6Table 7
Figure 4: Flow training analysis. (a) PCA of condition-sensitive reverse directions; arrows show task-type means. (b) The logged JVP squared-norm statistic and its binned training trends. Both panels analyze ALFWorld flow training. Additional training and JVP diagnostics appear in Appendices F.4 and F.5 .
ALFWorld
Search-QA
WebShop
Skill source
Seen
Unseen
Single-hop
Multi-hop
SR
Embedding retrieval
38.57
30.60
42.20
27.45
15.20
Ours
74.29
77.61
45.73
31.82
18.40
Δ (percentage points)
+35.71
+47.01
+3.53
+4.37
+3.20
Table 8: Generated versus retrieved skill guidance under 25% injected irrelevant skills. ALFWorld and WebShop report success (%); Search-QA reports EM (%). Δ is generation minus retrieval in percentage points.
Training examples
100%
75%
50%
500
78.10%
64.96%
63.50%
1,000
79.56%
72.26%
68.98%
3,553
83.21%
78.83%
64.23%
Table 9: Training-bank size and source-trajectory success. Cells report success rates (%) on 274 tasks. Column headings indicate the fraction of source trajectories that succeeded. Bold marks the highest success at each training size.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Domain
Training data
Codec training
Flow training
ALFWorld
3,553 fixed training chains
Encoder/decoder LoRA and projector training
Step 10,000; training-domain holdout selection
Single-hop QA
NQ; 4,000 pairs
Batch 48; 26 epochs
Batch 144 to step 10,000, then 2,304 to step 30,000
Multi-hop QA
2,723 pairs
Batch 48; 24 epochs; selected epoch 9
Batch 192; 30,000 steps; selected step 20,000
WebShop
2,000 pairs
Batch 24; 24 epochs
Batch 144; selected step 30,000
Appendix
Table 10: Domain-specific training recipes for SkillFM.
Domain
Output tokens
Actions
Skill use and evidence retrieval
ALFWorld
8,192
50
One skill; one rollout
Single-hop QA
2,048
4
Top-3 evidence retrieval
Multi-hop QA
8,192
4
Top-3 evidence retrieval; grounded hybrid
WebShop
4,096
30
One skill; one rollout
Appendix
Table 11: Execution budgets for the frozen Qwen3-8B actor.
Method
SR ↑
Score ↑
Method
SR ↑
Score ↑
ReAct
19.50
46.20
MemP
6.40
25.30
Reflexion
28.80
58.10
SimpleMem
8.59
33.20
Mem0
2.00
23.90
MemRL
9.20
29.50
ExpeL
11.20
30.90
EvolveR
17.60
42.50
Ours
20.00
55.35
Appendix
Table 12: Performance on WebShop in success rate (%) and task score. SR denotes complete success, and Score denotes the mean task reward.
Actor
NQ
Triv
Pop
Hotp
2WK
MuS
Bam
Avg
Qwen3-4B
37.4
57.0
40.4
33.2
37.2
11.0
37.6
36.26
Qwen3-8B
38.8
62.0
41.4
38.4
43.2
11.0
44.8
39.94
Qwen3-32B
39.2
66.4
41.8
42.6
46.2
16.0
48.8
43.00
Appendix
Table 13: Search-QA across executor scales. Scores are EM (%); Avg is the unweighted mean over seven datasets.
Actor
SR (%)
Score
Qwen3-4B
14.60
53.36
Qwen3-8B
20.00
55.35
Qwen3-32B
21.60
51.88
Appendix
Table 14: WebShop across executor scales. SR is success (%) over 500 tasks; Score is mean reward on a 0–100 scale.
Skills N
Source success
Seen
Unseen
Overall
500
100%
75.00%
81.34%
78.10%
500
75%
58.57%
71.64%
64.96%
500
50%
62.86%
64.18%
63.50%
1,000
100%
76.43%
82.84%
79.56%
1,000
75%
68.57%
76.12%
72.26%
1,000
50%
65.00%
73.13%
68.98%
Appendix
Table 15: ALFWorld training-bank size and quality. Success rates (%) on 140 seen and 134 unseen tasks; Overall pools all 274 tasks.
Observable pattern
Seen
Unseen
Failed episodes
25/140 (17.86%)
21/134 (15.67%)
Never executed a take action
16
13
Took objects, but only of another category
4
6
Acquired the correct category, but did not finish
5
2
At least one inadmissible action
12
9
Appendix
Table 16: Observable failure patterns on ALFWorld. Counts refer to failed episodes. The three acquisition categories partition the failures; the invalid-action row can overlap with them.
Prefix length K
Epoch 1
Epoch 8
Epoch 16
Epoch 24
4
1.631252
0.007894
0.004181
0.001952
8
1.516875
0.008462
0.004854
0.002112
16
1.231258
0.007750
0.003986
0.001594
32
1.328236
0.007550
0.003797
0.001281
Appendix
Table 17: Training reconstruction cross-entropy by prefix length. Values are epoch means; bold marks the lowest loss per epoch.
Figure 6: Latent-to-prefix reconstruction on ALFWorld. The supplied learning curves show training cross-entropy and clean/full-noise validation cross-entropy on a logarithmic scale. Shading marks the noise ramp through epoch 8, and stars mark selected checkpoints. The plot’s L denotes prefix length, written as K in this paper. These are reconstruction diagnostics, not downstream success curves.
Figure 7: ALFWorld flow-training diagnostics. Left: training-loss medians and 10th–90th percentiles within 200-update bins from one run. Right: condition-matched, macro-averaged EMA endpoint errors on the flow-stage holdout; the star marks the selected checkpoint.
Updates
Training loss
Holdout Chamfer
Mean angle
Angle P95
10,000 (selected)
5.670
0.014282
0.09467
0.28164
20,000
4.315
0.014569
0.09415
0.28416
30,000
3.827
0.014731
0.09498
0.28319
Appendix
Table 18: Recorded EMA checkpoints and endpoint diagnostics. Loss is the mean over the preceding 1,000 optimizer updates. Chamfer is in rad 2 ; angular statistics are in radians. Bold marks the lowest loss or Chamfer among the three checkpoints.
Figure 8: Directional-derivative diagnostics during iMF training. Left: per-update Jk (faint trace), 200-update means and medians, and within-bin 10th–90th percentiles. Right: the 1,000 updates preceding each evaluated checkpoint; boxes show the interquartile range, whiskers the 10th–90th percentiles, horizontal lines medians, diamonds means, and dots values beyond the whiskers, on a logarithmic axis. Each observation is a weighted minibatch aggregate, not an individual-example derivative or an independent training run.