Robot foundation models provide strong visuomotor control, yet their performance can degrade when object positions or task instructions change. Further improvements often require post-training on substantial robot data, which can be costly to collect through methods such as teleoperation. Agentic harnesses can adapt around the model, but current self-evolving harnesses use robot trials inefficiently when deciding which code and skill changes to pursue. We introduce EmbodiedRSI, a self-evolving agentic harness that autonomously decides where to explore next and turns the resulting physical interaction into improved code and skills. EmbodiedRSI realizes this through a Fast-Slow Dual-System Architecture, in which competing code and skill hypotheses are maintained in a Hypothesis Graph. Value-of-Information Experiment Selection chooses physical experiments that can distinguish these hypotheses. Their outcomes guide Code-Skill Co-Evolution. The Slow System builds Hierarchical Memory, and Reward-Grounded Memory Learning selects effective memory according to their value for later Fast-System improvement. On RoboCasa365, EmbodiedRSI reaches 77.0% overall success and 71.3% on Composite-Unseen, compared with 40.1% for the best baseline. EmbodiedRSI also reaches 86.8% overall success on LIBERO-Pro. Beyond benchmark performance, EmbodiedRSI transfers zero-shot to real-world robot, achieving 71.3% overall success across multiple challenging tasks.
Figures & tables
Figure 1: The EmbodiedRSI Self-Evolution Loop. Physical outcomes guide code and skill updates while the robot foundation model remains frozen.
Figure 2: Overview of EmbodiedRSI. The Fast System uses a Hypothesis Graph to choose each physical experiment. Each result guides a joint code-skill update. The Slow System retains memories that improve later adaptation.
Figure 3: Hypothesis Graph Updates from Physical Trials. Before Experiments (left). Directed refines edges show how code and skill hypotheses have evolved, while dashed links mark code-skill relations awaiting evidence. After Experiments (right). Solid bidirectional works-with edges identify relations supported by measured positive joint gain.
Figure 4: RoboCasa365 Task Suites. (a) Atomic-Seen contains individual manipulation tasks observed during pretraining. (b) Composite-Seen combines multiple skills into tasks represented in pretraining. (c) Composite-Unseen combines skills into tasks absent from pretraining.
Method
Atomic-Seen
Composite-Seen
Composite-Unseen
Overall
π0
34.6
6.1
1.1
14.8
π0.5
39.6
7.1
1.2
16.9
RLDX-1
60.0
21.3
5.0
30.0
WorldDreamer
66.3
26.7
9.0
35.3
Xiaomi-Robotics-1
80.2
57.1
32.1
57.4
Harness VLA
91.3
55.8
40.1
63.6
Table 1: RoboCasa365 Success Rates. Values are percentages. “–” denotes a suite without a reported result for the corresponding method.
Method
Spat-T
Spat-S
Obj-T
Obj-S
Goal-T
Goal-S
L10-T
L10-S
Overall
OpenVLA
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
π0
0.0
0.0
0.0
2.0
0.0
0.0
0.0
0.0
0.3
π0.5
1.0
20.0
1.0
17.0
2.0
38.0
1.0
8.0
11.0
MolmoAct
0.0
0.0
0.0
6.0
0.0
0.0
6.0
0.0
1.5
NORA
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
X-VLA
0.0
0.0
8.0
2.0
9.0
1.0
10.0
0.0
3.8
Table 2: LIBERO-Pro Success Rates. The table aggregates success rates (%) across SPATIAL, OBJECT, GOAL, and LIBERO-10 under instruction-redirection (T) and position-swap (S) perturbations. “–” denotes a cell not applicable to the corresponding method.
Task
SmolVLA
EmbodiedRSI
Oreo to Red Bowl
70.0%
86.7%
Poker Chip Selection ( ×5=100 )
10.0%
90.0%
Stack Two Cakes on a Can of Luncheon Meat
63.3%
76.7%
Glasses Bridge Grasp
60.0%
56.7%
Towel Folding
26.7%
46.7%
Table 3: Real-Robot Success Rates. Success over 30 trials per task on the real robot.
Figure 6: Code-Skill Co-Evolution Ablation. Bars show RoboCasa365 success rates.
Method
Overall Success @ 50 Ep. ↑
Learning AUC ↑
Random Exploration
27.1
0.309
EmbodiedRSI (Ours)
72.9
0.635
Table 4: Learning Efficiency. EmbodiedRSI is compared with Random Exploration.
Figure 7: Representative LIBERO-Pro Learning Curves. Eight-episode rolling success on SPAT-T over 100 physical episodes. (a) Value-of-Information Experiment Selection reaches sustained 100% at episode 48 (dashed line). (b) Random Selection continues to fluctuate.
Variant
Composite-Unseen ↑
Learning AUC ↑
Retrieval-Based
61.7%
0.556
PPO + Task Reward Only
65.2%
0.589
EmbodiedRSI (Ours)
71.3%
0.635
Table 5: Reward-Grounded Memory Learning Ablation. Composite-Unseen and learning AUC on RoboCasa365.
Current embodied systems largely rely on pretrained capabilities that remain fixed after deployment, limiting their ability to learn from physical interaction. We introduce MachEmbodied-Brain (ME-Brain), a self-evolving embodied system organized around a closed loop of action execution, experience acquisition, experience evolution, and improved execution. Evolvable Memory consolidates multimodal trajectories into hierarchical, reusable experience; Cognitive Core transforms physical experience into transferable skills; and the Action Model combines event-driven keyframes, EventCell local-world prediction, and action-conditioned memory modulation to focus computation on decision-critical moments, regions, and historical evidence. Together, these modules shift embodied intelligence from train-and-freeze to deploy-and-evolve without model retraining. Cognitive Core outperforms the strongest comparison models by 8.2 and 9.6 points on embodied and agent benchmarks. The Action Model achieves 47.88% mean success on RoboMME, a 3.26-point improvement over the strongest baseline. On RoboDojo, it reaches a 21.51 mean Score and 16.03% success rate, exceeding π0.5 by 10.10 and 9.12 points. On the six-task ME-RealBench, ME-Brain achieves a 69.5 mean Score and 66.7% success rate, outperforming DM0.5 by 12.8 and 11.7 points, respectively.
Recent work extends recursive self-improvement (RSI) to agent harnesses for long-horizon coding and terminal tasks, enabling agents to improve execution mechanisms from experience. However, generalizable harness RSI remains challenging. First, evolving harnesses on evaluation benchmarks or their subsets makes it difficult to distinguish reusable improvements from benchmark-specific adaptation. Second, single-trajectory updates can conflate systematic harness deficiencies with instance-specific reasoning and solution details, producing modifications that transfer poorly to unseen tasks. Third, localizing recurring behavioral deficiencies within monolithic harnesses is difficult, while whole-harness optimization can entangle unrelated mechanisms and complicate attribution and validation. We propose ModularRSI, a benchmark-disjoint, contrastive, and modular framework for generalizable harness evolution. ModularRSI contrasts successful and failed trajectories for the same task and aggregates evidence across tasks to identify recurring behavioral deficiencies. It decomposes the evolvable harness into five functional modules: Agent Loop, Tool Use, Observation Management, Context Management, and Task Completion Detection. Each module evolves independently within a restricted modification scope, followed by an integration stage that combines the evolved modules into a unified harness and resolves potential conflicts. To support benchmark-disjoint evolution, we curate 2,000 executable evolution tasks from external sources that are disjoint from downstream evaluation benchmarks. Experiments on TB2.0 and SWE-Bench Verified show consistent improvements on unseen in-domain and cross-domain tasks, with the evolved harness also transferring across different foundation models.
Siwei Wu, Jincheng Ren, Yizhi Li +11
University of Manchester · 3IQuest Research · 4M-A-P +3
Embodied agents are increasingly built as systems around foundation models, where performance depends not only on model weights but also on the skills, context, action interfaces, and execution harness surrounding the model. While supervised fine-tuning and reinforcement learning can adapt agents to new environments, they require additional data, rewards, and training runs; meanwhile, many train-free code-centric approaches rely on programmable robot APIs that may be unavailable in fixed-interface settings. We propose SHAPER, a self-evolving framework for train-free embodied adaptation that keeps model parameters frozen and improves the non-parametric agent system by evolving reusable skills and a context-code harness through target-environment rollouts. In SHAPER, the same frozen model can serve as both planner and optimizer, refining its external skills and context-code harness without parameter updates. We evaluate SHAPER on VLABench and ESI-Bench, covering embodied agents with different low-level action interfaces, and compare against pure execution, supervised fine-tuning, and test-time-scaling baselines such as verifier-free selection and voting. Our results suggest that skill-and-harness optimization is a practical route to self-evolving embodied agents when model training is expensive, unavailable, or undesirable.