Self-evolving agents distill task-solving experience into skills for future reuse, but these skills can encode incorrect procedures or non-transferable knowledge. It is therefore critical to verify each skill's reusability: whether its guidance remains useful beyond the experience from which it was distilled. Such verification requires observing how a skill affects execution in new tasks, yet existing tasks may not expose the situations where the target skill can actually be exercised. To construct such situations, we propose SkillSandbox, a framework that dynamically synthesizes a task and its environment for each skill that are skill-relevant yet novel. A Proposer specifies the conditions to preserve and the source-specific details to vary, a Builder constructs an executable scenario, and a Verifier compares executions with and without the skill. The Verifier assesses executability, utility, and efficiency to assign a Keep or Reject verdict, determining whether the skill enters the library. Across ALFWorld and WebShop with three models, SkillSandbox consistently yields the strongest downstream performance and improved execution efficiency. Further analyses examine whether these gains reflect accurate assessment of skill reusability and identify which components of SkillSandbox contribute to them.
Figures & tables
Executor
ALFWorld
WebShop
N
140
200
300
500
Gemini 3.1 Flash
5.96
5.96
5.90
6.10
Qwen3.5-27B
6.00
6.60
6.58
6.66
Table 1: Mean skill-relevant tasks out of 10 retrieved tasks. N : candidate task count.
Figure 1: Percentage of skills used in (a) at least one task and (b) at least 5 tasks as downstream tasks are streamed and solved sequentially (Gemini 3.1 Flash-Lite, WebShop). Mean of three runs of the same task stream; bands show one standard deviation.
Figure 2: Overview of skill verification via dynamic scenario synthesis.
Table 2: Main results on held-out ALFWorld and WebShop tasks. Vanilla Skill, ReasoningBank, and MemP retain skills without verification; ExpeL uses source-task replay, ACE and SkillOS use downstream feedback, and SkillSandbox uses synthesized scenarios for verification. Bold marks the best value per model; colored percentages are relative changes from No Skill.