Humanoids now walk, balance and reach with remarkable generality: one whole-body tracking policy follows references from a human, or from an end-to-end policy. That generality travels in the trajectory, and a trajectory alone carries limited information about the interaction it should produce: at contact, the executing controller determines how the robot behaves. Single-task policies usually reach hard interactions by optimising trajectory and controller together in simulation; general stacks usually assume a preset or hand-chosen controller. We present KPI, a promptable kernel for physical interaction between the trajectory source and an unmodified whole-body tracker. Instead of a controller fixed before the task, the trajectory source sends a contract: per direction, track, comply, or hold a force range. From tracking error and a wrench estimate, the kernel adapts the arms' stiffness, damping, reference and feedforward toward it at contact rate. We demonstrate KPI through an agentic framework: from one instruction, a vision-language agent writes both the reference trajectory and the contract, with no task-specific code. We demonstrate instruction-driven winch operation, door opening, and box transport, alongside scripted surface-interaction experiments. In the winch demonstration, the humanoid is able to turn a crank to hoist a second robot fully off the ground.
Figures & tables
Fig. 2: Agentic execution through KPI. (a) Analyzer writes the motion and the contract σ from Detector’s geometric evidence, and Verifier checks feasibility, stage outcomes and the remaining goal, returning rejected proposals for revision. Executor runs an accepted stage: a motion core supplies the nominal reference and KPI adapts the arm controller under σ from the interaction state. (b) One revision in the box task, numbered as in (a): Detector locates the two contact faces and the surrounding obstacles, Verifier rejects the original straight approach because it interferes with the tabletop, and Analyzer raises the hands before advancing while keeping both contacts.
Fig. 5: Winch operation. (a) Operation sequence. (b) KPI and KPI-fixed’s hand trajectories. (c) MCC ∗ failure case.
Fig. 6: Door opening and passage: grasp the handle, unlatch, step back, swing open, and walk through.
Fig. 7: Box transport and load robustness. (a) The five stages of transport and placement. (b) Maximum liftable load category. (c) SONIC fails to maintain multi-point contact with the box walls, preventing lifting.
Fig. 8: Board-writing results using different methods.
Keypoint tracking alone is insufficient for object interaction tasks such as sitting on a chair, wiping a board, or pushing furniture, where the robot can reach the correct pose without making meaningful physical contact with the object. We present CONTACTMIMIC, a learning framework that tracks explicit partlevel binary contact commands alongside keypoint trajectories. CONTACTMIMIC is made possible through the use of contact-following rewards and a trajectory augmentation scheme aimed at breaking the correlations between keypoint trajectories and contact labels. The resulting policy successfully decouples contact behavior from keypoint geometry, and achieves precise physical contact as well as contact-controllability (produce or suppress contact during deployment as desired). Simulation experiments across 10 diverse human-object interaction motions confirm that CONTACTMIMIC exhibits contact controllability that enables it to complete manipulation tasks without task-specific rewards, while also outperforming keypoint-only trackers on contact-relevant tasks. Ablations confirm the necessity of the proposed trajectory augmentation scheme and sim2real deployment validates contact controllability in the real world across 5 different motions. Video results are available on https://lixinyao11.github.io/contactmimic-page/.
Physical human-robot interaction requires yielding transiently to contact yet recovering the commanded reference under sustained load. Finite-stiffness impedance control retains a static deflection there, while predictive alternatives typically optimize a nonlinear robot or impedance model online. Operational-space cancellation instead exposes a translational error double integrator with a fixed transition matrix and a configuration-scheduled input map, making interaction a predictive quantity rather than a property re-derived per configuration. We build on it a compact offset-free interaction-error MPC for torque-controlled manipulators: a force-domain random-walk state estimates persistent interaction and model error, and a 30-variable convex QP maps the correction through the current task inertia while constraining the applied joint torque. Conditional results establish impedance equivalence of the unconstrained passive feedback, offset-free regulation at feasible frozen configurations, and quadratic stabilizability of the scheduled backbone. In a 1kHz MuJoCo simulation of a 7-DOF Franka FR3, the estimator cuts steady-state error under a repeated 15N step from 2.77mm to 0.042mm when added to the otherwise identical 100Hz MPC. A stiffness-and-damping-calibrated impedance baseline attains 2.59mm but briefly saturates and needs 3.3x the peak positive joint power. Adding ideal measured-force cancellation to that baseline gives 1.39mm, so constant-load rejection is not unique to MPC; the sensorless controller still reaches 0.042mm in the moving task, a 65x reduction without force sensing and without the baseline's saturation or power cost. Demonstrated in simulation under a shared actuator budget, the contribution is an efficient operational-space realization complementing rather than replacing broader interaction-control architectures.
Current humanoid reinforcement-learning policies excel at free-space motions but struggle with contact-rich tasks, as pure kinematic tracking cannot resolve the physical ambiguities of interacting with objects and uneven terrain. To address this, we introduce SceneBot, a unified motion-tracking framework capable of handling freespace locomotion, terrain traversal, and whole-body manipulation. SceneBot conditions a single policy on both reference motions and per-link contact labels, explicitly defining expected environmental interactions. To overcome the lack of annotated interaction data, we propose a hindsight scene reconstruction approach that infers scene-interaction graphs from retargeted human motion. Trained on 7.5 hours of this reconstructed, contact-rich data, SceneBot successfully generalizes to unseen motions and environments. Our results demonstrate that SceneBot is the first general framework to seamlessly unify free-space and contact-rich behaviors executing complex, long-horizon tasks like carrying a box upstairs and establishing contact conditioning as a powerful interface for humanoid control. All code and data will be open-sourced. More demos and information are available at: https://ericcsr.github.io/scenebot/
Sirui Chen, Shibo Zhao, Zhen Wu +3
Stanford, Amazon FAR United States · Amazon FAR United States · CMU, Amazon FAR United States