Organizations: The Chinese University of Hong Kong, Shenzhen · Fudan University · University of Illinois at Urbana-Champaign · University of Oxford · National University of Singapore
LLM-based agents can improve themselves across tasks by reusing and revising the skills they orchestrate into executable procedures. Flow-based training fits this loop: it samples procedures in proportion to reward, and the flow through each skill credits it for the next library revision. Three obstacles stand in the way of making this self-improvement reliable: flow training suffers strategy collapse over tree-structured histories; nonnegative flow-based credit rewards frequent use as if it were benefit; and library edits rest on the task reward the policy optimizes. We introduce R2 Flow, a recursive self-improvement framework that alternates policy learning, independent verification, and versioned skill-library updates on a shared-state orchestration graph. The graph merges histories that differ only in the order of independent steps, allowing flow training to pool evidence across equivalent executions. A flow-share readout of the trained flow, invariant to the backward policy, and a separate signed utility rank which skills to change, verifier evidence decides whether an edit is warranted, and a residual-variance plateau sets when to update. Committed edits reshape the graph the next policy learns on, realizing recursive skill evolution. Across question answering, mathematical reasoning, interactive decision making, and code generation, R2 Flow improves task accuracy and library-edit precision over heuristic orchestration, reinforcement learning, and skill-evolution baselines, and transfers across executors. Code is available at https://github.com/beita6969/r2flow.
Figures & tables
Figure 1: R 2 Flow trains a flow on shared states (inner loop), reads share and signed utility separately, and gates library edits on verified evidence (outer loop).
Figure 2: Four orchestration paradigms. (a) Workflow graph optimization hides a harmful skill in the return. (b) Flow-driven orchestration learns AB and BA twice on a history tree, so evidence never pools. (c) Posterior-guided evolution scores reliability, not signed utility, so success reads as benefit. (d) R 2 Flow pools evidence on shared states, separates share from utility, and commits verified edits.
Figure 3: The R 2 Flow framework. (1) The Supervisor πθ acts in a structured environment where histories differing only in the order of independent steps share one state. (2) PF , PB and Fψ are trained with R 2 TB; flow share Ψ0(u) and signed utility Ae are read separately. (3) Verifier evidence, the two readouts and the variance plateau decide whether, which and when to edit.
Variant
IID
OOD
HotpotQA
TriviaQA
AIME
HealthB.
MBPP+
ALFWorld
MuSiQue
NQ-Open
MATH
GPQA
SWE
WebShop
Ans EM
Ans EM
Acc.
Rubric
Pass@1
SR
Ans EM
Ans EM
Acc.
Acc.
Resolved
SR
Qwen3.5-9B (frozen)
57.97
44.84
51.33
33.84
78.75
46.56
39.06
23.28
88.44
61.09
16.41
33.28
R 2 Flow architecture, untrained
77.03
77.81
54.00
39.53
83.75
63.44
68.28
61.09
89.84
65.94
22.97
55.31
− Shared states (history tree)
84.06
91.09
62.00
53.75
88.12
81.09
77.19
83.12
93.28
79.06
38.12
79.38
− Learned PB (uniform PB )
86.56
91.88
67.33
57.19
88.59
85.31
81.25
84.06
94.84
83.91
40.31
84.84
Table 2: Component ablations, five-run mean, protocol of § 5 . Each row removes one component of R 2 Flow, with everything else fixed. The frozen backbone and R 2 Flow (Full) rows are reproduced from Table 1 . The untrained row runs the R 2 Flow architecture with an untrained Supervisor. PB is the backward policy, and the flow share is read from the compatible reference flow; the signed utility says whether a skill helps; verifier labels, credible bounds and the paired non-inferiority check gate each library edit; a plateau in the residual variance ends a phase, and the flow and verifier posterior carry into the next. Row blocks follow §4.1–4.3: graph, flow training and readouts, skill evolution.
Figure 4: Executor transfer, significance and training dynamics, five-run mean. (a) IID scores of six frozen executors answering directly (dashed) and under R 2 Flow with the same trained Supervisor (solid); the radius is linear from 0 to 100. (b) Gain of R 2 Flow over each baseline of Table 1 , averaged over the six IID or six OOD datasets on each primary metric, with 95% confidence intervals; *** marks p<10−3 (Welch’s t -test, five runs). (c) Training accuracy of R 2 Flow and of − Shared states over 250 steps; shaded: spread across runs; numbers: mean over the last 50 steps. (d) OOD scores averaged per domain: QA (MuSiQue, NQ-Open), Math/Sci (MATH, GPQA), Code (SWE-bench), Interactive (WebShop); each bar stacks the frozen score and the R 2 Flow gain.
Figure 5: Training objectives and recursive self-improvement, five-run mean. (a) IID scores and per-step cost of each Supervisor objective (TB, Malkin et al., 2022 ; DB, Bengio et al., 2023 ; SubTB without bd , Madan et al., 2023 ; TTB, Zhang et al., 2026e ). (b) Row minus column objective in mean OOD score. (c) Mean OOD score per phase; − Verifier uses reward-derived labels, − Library freezes the library, and Fixed Phase replaces the plateau trigger. (d) Training-reward gain against verified gain per committed edit; shaded: held-out verified score falls. (e) Phase at which each of twelve injected harmful skills is pruned. (f) Distinct strategies per domain, reorderings counted once.
In recent years, a variety of powerful LLM-based agentic systems have been applied to automate complex tasks through task orchestration. However, existing orchestration methods still face key challenges, including strategy collapse under reward maximization, high gradient variance with opaque credit assignment, and unguided skill evolution whose decisions are typically made by directly prompting an LLM to judge rather than derived from principled training signals. To address these challenges, we propose SkillFlow, a flow-based framework that takes a trainable Supervisor as the agent and a structured environment with dynamic skill library and frozen executor, automating task orchestration through multi-turn interaction. SkillFlow employs Tempered Trajectory Balance (TTB), a regression-based flow-matching loss that samples trajectories proportional to reward, preserving diverse orchestration strategies rather than collapsing to a single mode. The same flow objective yields a jointly learned backward policy that provides transparent per-step credit assignment at zero additional inference cost. Building on these flow diagnostics, a recursive skill evolution mechanism determines when to evolve, what skills to create or prune, and where decision gaps lie -- closing the loop from training signal to autonomous capability growth. Experimental results on 14 datasets show that SkillFlow significantly outperforms baselines across question answering, mathematical reasoning, code generation, and real-world interactive decision making tasks. Our code is available at https://anonymous.4open.science/r/SkillFlow-E850.
Mingda Zhang, Tiesunlong Shen, Haoran Luo +4
The Chinese University of Hong Kong, Shenzhen · National University of Singapore · Nanyang Technological University +1
Agentic large language models often rely on skills, reusable natural language procedures that guide planning, action, and tool use. In practice, skills are typically improved through prompt engineering or by aligning the task LLM itself, which is costly, model-specific, and often infeasible for closed-source models. Skill optimization is not a one-step problem but a recurrent process with two coupled levels of credit assignment: a useful skill must improve rollout quality under current conditioning, while a useful revision must turn observed outcomes into a better skill for the next round. We propose Skill-R1, a reinforcement learning framework for instance-level recurrent skill optimization from verifiable rewards. Rather than updating the task LLM, Skill-R1 trains a lightweight skill generator that conditions on the task context, prior rollouts, and their verified outcomes to produce skills that steer a frozen task LLM. This preserves black-box compatibility with both open- and closed-source models while making adaptation substantially cheaper than model-level updates. Skill-R1 proceeds over multiple generations: at each step, the current skill induces rollouts whose verified outcomes are fed back to produce the next revision. To optimize this recurrent process, we introduce a bi-level group-relative policy optimization objective combining intra-generation and inter-generation advantages. The intra-generation term compares rollouts under shared skill conditioning, while the inter-generation term rewards revisions that improve behavior across successive generations. Together, these provide a principled objective for directional skill evolution rather than one-shot self-refinement. Empirically, Skill-R1 achieves consistent gains over no-skill baselines and standard GRPO across benchmarks with verifiable rewards, with particularly strong improvements on complex, multi-step tasks.
Skill documents, structured natural-language instructions that guide Large Language Model (LLM) agents, are critical to modern agent frameworks, yet LLMs struggle to write skills that actually work. On SkillsBench, human-authored skills improve pass rates by 16.2 percentage points, while LLM-authored skills provide no measurable gain. We introduce SkillAxe, a fully unsupervised framework that enables LLMs to iteratively diagnose and refine their own skills. SkillAxe decomposes skill quality into four interpretable dimensions (quality impact, trigger precision, instruction compliance with fault attribution, and solution-path coverage), producing structured improvement briefs that require no ground-truth labels, test suites, or environment rewards. On SkillsBench, SkillAxe improves pass rates by 28% relative over unimproved LLM skills and closes 47--67% of the gap to human-authored skills. We validate the approach as a continuous improvement engine in the wild on SpreadsheetBench, where a SkillAxe-built skill library learns from past agent trajectories and raises pass rate from 16.0% to 52.0% using only 22 skills.