Experience-Guided Initiation Search for Learned Skills in Skill Composition
Authors: Qixuan Li, Yanhong Zhao, Jincheng Yu
Organizations: Shenzhen International Graduate School, Tsinghua University, Shenzhen, China · School of Electronic Science and Engineering, Nanjing University, Nanjing, China · Department of Electronic Engineering, and the Institute for Embodied Intelligence and Robotics, Tsinghua University, Beijing, China
Deploying frozen learned skills, such as Vision-Language-Action (VLA) policies, in new environments requires identifying initiation configurations that support reliable execution. In skill composition, an initiation configuration affects not only the current skill but also the physical state passed to subsequent skills, so successful execution of an individual skill does not necessarily imply successful completion of the composed task. Estimating target-specific capability through extensive rollouts is costly in real-world deployment, while directly reusing historical experience can be unreliable under environment changes. We propose EVIS, an Experience-Guided and Behavior-Validated Initiation Search framework for discovering reliable initiation configurations under limited target interaction. EVIS uses historical execution experience to prioritize promising candidates and target-environment behavior to validate whether they remain effective. We evaluate EVIS on single-skill and two-stage manipulation tasks with frozen VLA policies. EVIS reduces mean target-environment queries and improves reliable candidate discovery under small interaction budgets. These results show that combining historical guidance with target-side behavioral validation can reduce the interaction cost of deploying frozen learned skills in new environments.
Figures & tables
Fig. 1: Experience-guided and behavior-validated initiation search for frozen learned skills. Historical execution experience from prior contexts is used to propose promising initiation configurations for an unseen target context. EVIS then uses a small number of target-environment rollouts to validate these proposals and identify a reliable initiation configuration for the desired execution objective.
Fig. 2: Overview of EVIS. Historical execution experience from prior contexts is used to construct multiple scoring hypotheses over initiation configurations, which provide complementary proposals for a new unseen target context. Target-environment rollouts then behaviorally validate the proposed configurations under objective G and determine whether a candidate should be accepted or the search should continue. The underlying learned skill policy remains frozen throughout the process.
Fig. 3: Representative historical contexts and unseen target contexts in RoboCasa365, illustrating variations in visual appearance and kitchen layout.
Fig. 4: Representative single-skill initiation search on an unseen target context. Historical guidance prioritizes promising configurations, allowing EVIS to observe a successful candidate earlier than Scratch-Sobol.
Metric
EVIS
EVIS-Fixed B20
History Reuse B5
Scratch-Sobol B20
Mean Queries ↓
8.125
20.000
5.000
20.000
Confirmed SR ↑
76.7%
76.7%
70.8%
72.5%
TABLE I: Single-skill initiation search.
Budget B
→ DrawerUtensilSort
→ ChooseMeasuringCup
EVIS
Scratch-Sobol
EVIS
Scratch-Sobol
Reliable Rate
Confirmed SR
Reliable Rate
Confirmed SR
Reliable Rate
Confirmed SR
Reliable Rate
Confirmed SR
6
25.0%
12.1%
0.0%
0.0%
33.3%
14.6%
8.3%
3.8%
10
41.7%
20.4%
8.3%
5.4%
58.3%
29.2%
16.7%
7.1%
15
41.7%
20.4%
33.3%
20.0%
66.7%
33.3%
41.7%
17.9%
20
66.7%
32.9%
58.3%
30.8%
66.7%
33.3%
58.3%
24.6%
TABLE II: Multi-stage initiation search.
Fig. 5: Effect of experience-guided proposals. (a) Cross-context candidate ranking for single-skill drawer opening. (b) Multi-stage initiation search with and without historical initialization under increasing target-interaction budgets.
Task
Method
False Accept ↓
Rel. Precision ↑
ChooseMeasuringCup
History-only
7
41.7%
+ Qualification
0
100%
DrawerUtensilSort
History-only
11
8.3%
+ Qualification
0
100%
TABLE III: Effect of behavioral validation.
Method
Recovery Rate ↑
Final Reliable Rate ↑
Mean Queries ↓
Qualification Only
0.0%
8.3%
2.33
Sobol Fallback
42.9%
33.3%
13.83
EVIS Fallback
57.1%
41.7%
13.00
TABLE IV: Recovery after rejected historical reuse.
Skill-augmented agents improve sample efficiency by distilling successful trajectories into reusable strategies. Yet most existing approaches remain text-centric, linearizing spatial layouts and action-state correspondences into language that loses critical geometric structure. Recent efforts have begun incorporating visual evidence, but construct and update skills separately from policy optimization, leaving their mutual improvement underexplored. We propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents. Retrieved skills guide both inference and reward shaping, while successful trajectories are distilled back into the library, forming a closed feedback loop in which skill accumulation and policy improvement reinforce each other. An optional cold-start mechanism further accelerates early-stage learning. Evaluated on Sokoban, FrozenLake, and PrimitiveSkill, ViSkill achieves an overall success rate of 0.89, rising to 0.91 with cold-start initialization, outperforming all evaluated proprietary and open-source baselines while converging faster than standard PPO. Our code is available at https://github.com/ZJU-REAL/ViSkill.
Vision-language-action (VLA) models map visual observations and language instructions directly to robot actions, but long-horizon tasks require more than action prediction. An agent must coordinate perception, planning, execution, progress verification, and recovery as the physical state evolves. An action prediction or a model-generated skill decision does not, by itself, guarantee that the proposed operation is valid in the current state or that its outcome will be verified. We propose EmbodiedSkills, a unified framework that treats each skill decision as an execution proposal: the runtime checks its prerequisites before execution and verifies the outcome afterward. A shared executable-skill interface connects high-level skill selection, bounded low-level VLA execution, and post-action verification within a single agent loop. Because this interface remains fixed, low-level VLA policies can be replaced or adapted without changing the agent loop. The interface also records planning, execution, verification, and recovery events as structured trajectories, which provide supervision for individual components and can support optional online adaptation when interactive feedback is available. We instantiate EmbodiedSkills with Qwen3-VL and OpenPI/pi0.5 on RoboTwin 2.0 and LIBERO. Task-adapted low-level VLA policies achieve an average success rate of 86.20% across 50 RoboTwin 2.0 tasks and 97.40% across the four LIBERO suites. These results establish the execution performance of the task-adapted low-level VLA policies used in EmbodiedSkills. On four memory-dependent RMBench tasks, the same task-adapted execution approach achieves 12.5% average success. The framework provides a trainable and inspectable agent layer for turning these policies into closed-loop embodied systems.
Wei Wang, Wenqiao Zhang, Yutong Lin +14
College of Computer Science and Technology, Zhejiang University · Nanjing University of Aeronautics and Astronautics · Cornell University +3
Large language model (LLM) agents accumulate rich episodic trajectories while solving real-world tasks, but it remains unclear whether such experience can be distilled into reusable procedural skills. We introduce SkillEvolBench, a diagnostic benchmark for evaluating this step from experience reuse to skill formation. It contains 180 tasks across six real-world agent environments, organized into role-conditioned task families with shared latent procedures. Agents learn from acquisition tasks, update an external skill library using compacted trajectories and verifier feedback, and then face frozen deployment tasks testing context shift, adversarial shortcuts, and composition. By comparing self-generated and curated-start skill evolution against no-skill and raw-trajectory controls, SkillEvolBench separates procedural abstraction from base capability, curated prior knowledge, and direct reuse of episodic traces. Across ten model configurations and three agent harnesses, we find that current agents often adapt locally but rarely form robust reusable skills. Skill-based conditions can improve acquisition or replay, and individual models sometimes gain on specific deployment axes, but these gains are unstable under frozen deployment. Raw-trajectory reuse frequently outperforms distilled skills, suggesting that current abstraction procedures discard contextual and procedural cues that remain useful for future tasks. Capacity and cost analyses further show that writing more skills or larger Tier-3 resource libraries is not sufficient: additional updates can improve coverage while introducing episode-specific drift and procedural clutter. These findings position SkillEvolBench as a testbed for measuring when one-off experience becomes durable procedural knowledge rather than task-local memory.
Yingtie Lei, Zhongwei Wan, Jiankun Zhang +13
1The Ohio State University · 2The University of Chicago · University College London +4