cs.ROAug 31, 2026

Zeva: In-Context Causal Learning for Generalizable Embodied Manipulation

Authors: Fu ChenXin DingBingjia HuangXiangyu LiMingju WangJiawei HeKun LiWei Sun+3 more

Organizations: Institute for AI Industry Research (AIR), Tsinghua University · AIR, Tsinghua

Abstract

Generalizable embodied manipulation remains difficult to achieve through pretraining alone, due to unseen physical conditions in the real world. We argue that robots need to learn from their own physical interactions on the fly during real-world deployment and use this knowledge to inform subsequent actions. We present Zeva, the first framework that enables in-context learning from a robot's own physical interaction experience while keeping the policy model frozen. Zeva employs a Causal Interaction Extractor to encode an executed action and its induced state change into a causal interaction signal, which is stored in a dual-timescale causal memory. For subsequent actions, relevant causal interaction signals are retrieved from memory and injected into the frozen policy model as context. Experiments in simulation and real-world manipulation demonstrate that Zeva achieves the best performance among the compared frontier VLAs and WAMs and, more importantly, enables self-evolution during deployment without gradient updates. Its success rate continues to improve as the robot accumulates interaction experience. Furthermore, the acquired interaction experience can generalize across tasks.

Explore similar work

CardsList
  1. In-Context Robot Learning with VLM Agents

    Sep 16, 2026Dongzhou Cheng, Taoran Yi, Ye Fang +12Robot SystemsIntelligence