cs.ROSep 17, 2026

LIFD: Anchored Diffusion for 3D-Aware Scene Memory in Robotic Manipulation

Authors: Wenbo LiYiteng ChenWenhao LiQingyao Wu

Organizations: School of Software Engineering, South China University of Technology

Abstract

Robotic manipulation under partial observability requires spatial information that extends beyond the current view. Geometry-aware RGB features describe visible structure, but previously observed regions may disappear as the robot or scene moves. Maintaining a useful scene representation therefore requires retaining observation history while inferring missing content without losing its connection to visible evidence. We introduce LIFD (Look, Imagine, Focus, and Do), a framework for persistent, 3D-aware scene memory. LIFD learns a scene-token representation from multi-view agreement and completes it from a single RGB view and recurrent memory. A rectified-flow model generates the tokens while Anchor-Guided Cross-Attention conditions completion on current geometric features. Compact slot features connect this representation to a manipulation policy. Multi-view and geometric supervision are used during representation learning; deployment requires one RGB camera, proprioception, and a task instruction. LIFD (Staged) reaches 91.6% average success on LIBERO and 79.8% on MetaWorld, improving LIBERO average success by 3.1 percentage points over Joint training. On four UR5e task families with ten demonstrations per family, it achieves 56.0% mean success, compared with 40.5% for OpenVLA-7B.

Explore similar work

CardsList
  1. GeomVLA: Unifying Scene, Motion, and Action in 3D

    Sep 12, 2026Ziyin Xiong, Nikolaos Gkanatsios, Moritz Reuss +13D MotionPerception-Action Loop