Experience-Guided Initiation Search for Learned Skills in Skill Composition
Authors: Qixuan Li, Yanhong Zhao, Jincheng Yu
Organizations: Shenzhen International Graduate School, Tsinghua University, Shenzhen, China · School of Electronic Science and Engineering, Nanjing University, Nanjing, China · Department of Electronic Engineering, and the Institute for Embodied Intelligence and Robotics, Tsinghua University, Beijing, China
Deploying frozen learned skills, such as Vision-Language-Action (VLA) policies, in new environments requires identifying initiation configurations that support reliable execution. In skill composition, an initiation configuration affects not only the current skill but also the physical state passed to subsequent skills, so successful execution of an individual skill does not necessarily imply successful completion of the composed task. Estimating target-specific capability through extensive rollouts is costly in real-world deployment, while directly reusing historical experience can be unreliable under environment changes. We propose EVIS, an Experience-Guided and Behavior-Validated Initiation Search framework for discovering reliable initiation configurations under limited target interaction. EVIS uses historical execution experience to prioritize promising candidates and target-environment behavior to validate whether they remain effective. We evaluate EVIS on single-skill and two-stage manipulation tasks with frozen VLA policies. EVIS reduces mean target-environment queries and improves reliable candidate discovery under small interaction budgets. These results show that combining historical guidance with target-side behavioral validation can reduce the interaction cost of deploying frozen learned skills in new environments.
Figures & tables
Fig. 1: Experience-guided and behavior-validated initiation search for frozen learned skills. Historical execution experience from prior contexts is used to propose promising initiation configurations for an unseen target context. EVIS then uses a small number of target-environment rollouts to validate these proposals and identify a reliable initiation configuration for the desired execution objective.
Fig. 2: Overview of EVIS. Historical execution experience from prior contexts is used to construct multiple scoring hypotheses over initiation configurations, which provide complementary proposals for a new unseen target context. Target-environment rollouts then behaviorally validate the proposed configurations under objective G and determine whether a candidate should be accepted or the search should continue. The underlying learned skill policy remains frozen throughout the process.
Fig. 3: Representative historical contexts and unseen target contexts in RoboCasa365, illustrating variations in visual appearance and kitchen layout.
Fig. 4: Representative single-skill initiation search on an unseen target context. Historical guidance prioritizes promising configurations, allowing EVIS to observe a successful candidate earlier than Scratch-Sobol.
Metric
EVIS
EVIS-Fixed B20
History Reuse B5
Scratch-Sobol B20
Mean Queries ↓
8.125
20.000
5.000
20.000
Confirmed SR ↑
76.7%
76.7%
70.8%
72.5%
TABLE I: Single-skill initiation search.
Budget B
→ DrawerUtensilSort
→ ChooseMeasuringCup
EVIS
Scratch-Sobol
EVIS
Scratch-Sobol
Reliable Rate
Confirmed SR
Reliable Rate
Confirmed SR
Reliable Rate
Confirmed SR
Reliable Rate
Confirmed SR
6
25.0%
12.1%
0.0%
0.0%
33.3%
14.6%
8.3%
3.8%
10
41.7%
20.4%
8.3%
5.4%
58.3%
29.2%
16.7%
7.1%
15
41.7%
20.4%
33.3%
20.0%
66.7%
33.3%
41.7%
17.9%
20
66.7%
32.9%
58.3%
30.8%
66.7%
33.3%
58.3%
24.6%
TABLE II: Multi-stage initiation search.
Fig. 5: Effect of experience-guided proposals. (a) Cross-context candidate ranking for single-skill drawer opening. (b) Multi-stage initiation search with and without historical initialization under increasing target-interaction budgets.
Task
Method
False Accept ↓
Rel. Precision ↑
ChooseMeasuringCup
History-only
7
41.7%
+ Qualification
0
100%
DrawerUtensilSort
History-only
11
8.3%
+ Qualification
0
100%
TABLE III: Effect of behavioral validation.
Method
Recovery Rate ↑
Final Reliable Rate ↑
Mean Queries ↓
Qualification Only
0.0%
8.3%
2.33
Sobol Fallback
42.9%
33.3%
13.83
EVIS Fallback
57.1%
41.7%
13.00
TABLE IV: Recovery after rejected historical reuse.