cs.ROSep 27, 2026

Estimate, Don't Imitate: Reusing Differentiable State-Based Policies for Visuomotor Control

Authors: Denis Shcherba, Adrian Abel, Eckart Cobo-Briesewitz, Paul Mattes, Wojciech Samek, Marc Toussaint

Organizations: Fraunhofer Heinrich-Hertz-Institut, Berlin, Germany · Technische Universität Berlin, Germany

Abstract

Simulation-trained manipulation policies can exploit privileged state information to learn effective contact-rich behaviours, but deployment requires acting from partial observations such as noisy camera images. A common solution is teacher-student distillation, in which a visuomotor policy is trained to reproduce the actions of the privileged expert. This requires the student to jointly infer the task-relevant state and relearn the expert's action mapping that is already available. An alternative is to reuse the state-based expert and learn only a perceptual interface that reconstructs its missing state inputs. However, minimising the state estimate error alone does not necessarily minimise the downstream control error induced by these estimates. To bridge this gap, we train a visual state estimator using both direct state supervision and an action-consistency loss backpropagated through the frozen, differentiable expert. A scheduled objective first establishes a physically meaningful state estimate and progressively emphasises errors that affect the expert's actions. Across five goal-conditioned manipulation tasks, retaining the expert consistently outperforms direct pixel-to-action imitation from the same expert demonstration corpus. We further demonstrate sim-to-real transfer on a physical Panda robot, achieving 76% success without retraining the underlying expert.

Figures & tables

Explore similar work

Dec 17, 2025cs.RO

ISS Policy : Scalable Diffusion Policy with Implicit Scene Supervision

Vision-based imitation learning has enabled impressive robotic manipulation skills, but action imitation alone provides limited supervision of the geometric consequences of robot behavior. To address this limitation, we introduce Implicit Scene Supervision (ISS) Policy, a 3D visuomotor diffusion policy with a DiT backbone that predicts continuous action sequences from point-cloud observations. ISS augments action diffusion with a supervised robot motion predictor that maps generated actions and robot-state context to end-effector motion, and then uses the predicted motion together with gripper intent to forecast future point-cloud representations. By explicitly modeling the intermediate transition from action to robot motion, ISS encourages the policy to capture how its actions affect the surrounding 3D scene. We further introduce asymmetric gradient routing to separate direct motion regression from scene-level policy supervision, together with a change-balanced objective that accounts for variations in scene-change magnitude. These auxiliary objectives provide dynamics-aware geometric supervision using only expert demonstrations, without requiring additional annotations or auxiliary modules at inference time. ISS Policy achieves state-of-the-art performance on single-arm manipulation tasks in MetaWorld and dexterous manipulation tasks in Adroit, while real-world dual-arm experiments further demonstrate its effectiveness on physical robotic manipulation. The resulting framework preserves the scalable DiT backbone and standard diffusion-policy control interface. Code and videos will be released.
Jun 12, 2026cs.RO

Spatially Conditioned Diffusion Policy: Learning Precise and Robust Manipulation with a Single RGB Camera

Recent visual imitation learning systems have widely adopted multi-camera setups with wrist-mounted cameras as the de facto standard. However, manipulation from a single global view remains challenging, as the policy should capture fine-grained interaction details and identify task-relevant regions without local wrist views. To address this challenge, we present Spatially Conditioned Diffusion Policy (SCDP), a diffusion-based visuomotor policy that achieves precise and robust manipulation in a single-camera setting. Our key idea is that end-effector trajectories can serve as visual attention anchors that reflect task-relevant regions. Building on this idea, SCDP consists of two key components: (i) a visual encoder that produces multi-scale feature maps to capture both broader context and fine-grained visual features, and (ii) a spatial conditioning module that samples point-wise features along intermediate end-effector trajectories in the diffusion loop. Extensive simulation experiments show that SCDP consistently outperforms strong single-view baselines and achieves performance comparable to multi-camera baselines. Real-world experiments further demonstrate precise manipulation and robustness to visual distractors, highlighting the potential of single-camera imitation learning.
Jun 8, 2026cs.RO

GHOST: Hierarchical Sub-Goal Policies for Generalizing Robot Manipulation

We present GHOST, a framework for learning visuomotor manipulation policies that generalize beyond the training distribution. GHOST factorizes control into (i) a high-level policy that predicts the next sub-goal as a distribution over 3D end-effector poses from multi-view RGB-D observations, and (ii) a low-level goal-conditioned controller that executes embodiment-specific actions. To condition image-based policies on 3D goals, we introduce a simple spatial interface that projects predicted goals into the image plane and represents them as end-effector heatmaps. Across a suite of manipulation tasks, this hierarchical factorization consistently improves performance and robustness compared to a flat Diffusion Policy. Further, we show that this hierarchical interface also makes it easy to incorporate human demonstrations without relying on (noisy) action retargeting. As sub-goals are largely embodiment-agnostic, we train the high-level policy on human video to specify how learned skills should be applied and composed, while keeping the low-level policy trained purely on robot data. This hierarchy enables adaptation to novel objects and task variations using a small number of human demonstrations.