iGPC: Generative Motion Priors for Object-Aware Humanoid Interaction
Authors: Anujith Muraleedharan, Abdul Ahad Butt, Nolan Fey, Yash Prabhu, Anamika J H, Sandor Felber, Maurice Rahme, Ivan Laptev
Organizations: Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, United Arab Emirates. · Massachusetts Institute of Technology (MIT), Cambridge, MA, USA. · Minerva Humanoids, Inc.
Humanoid robots operating in unstructured environments must combine robust whole-body control with the ability to perceive and physically interact with surrounding objects. While large-scale human motion data provides powerful priors for natural and versatile humanoid control, effectively transferring such priors to perception-driven object interaction remains challenging. To address this bottleneck, we propose a framework that extends the recently proposed Generative Pretrained Controller (GPC) from general human motion to full-body humanoid-environment interaction. First, we adapt GPC into interaction experts conditioned on scene affordance cues and privileged state information. These experts leverage the pretrained human motion prior while learning task-specific contact behaviors, including reaching toward objects, grasping environmental supports for stabilization, and pushing movable objects. Second, we introduce a perception-driven student that retains the pretrained GPC policy and distills interaction skills from the experts using onboard sensory observations. To bridge the gap between privileged expert observations and sensory inputs, we propose two complementary training objectives that enable effective adaptation of the pretrained motion prior during distillation. Notably, our experiments across multiple whole-body interaction tasks demonstrate that large-scale generative human motion priors provide an effective foundation for learning deployable policies for humanoid interactions in contact-rich real-world environments.
Figures & tables
Fig. 1: iGPC adapts human-motion priors to perceptive whole-body interaction on Unitree G1. Real-world sequences illustrate rail-assisted stabilization following an external push (top) and box pushing (bottom).
Method
Object manipulation
Reactive hand- supported stabilization
Onboard- perceptive contact
Motion- reference- free execution
Frozen- backbone low-rank adaptation
GPC [ 5 ]
–
–
–
✓
✓
BeyondMimic [ 2 ]
–
–
–
✓
–
ResMimic [ 8 ]
✓
–
–
–
–
ULTRA [ 10 ]
✓
–
✓
✓
–
LessMimic [ 11 ]
✓
–
✓
✓
–
ViBe [ 12 ]
✓
–
✓
∘v
✓
TABLE I: Qualitative comparison of selected methods.
Fig. 2: Overview of the proposed generative control framework. (A) Expert learning. A G1-specific motion controller and autoregressive motor prior form the motor foundation. Motion–scene supervision initializes object-conditioned task adapters, which are subsequently refined using privileged simulation observations. (B) Student learning. Depth, LiDAR, and proprioceptive history reconstruct the expert’s observation interface. Observation, token, and action supervision train the estimator, followed by bounded input calibration and task-adapter refinement. Calibrated observations o~t condition token selection, while the physical-state estimate s^t bypasses this correction and directly conditions the decoder Dψ . Flames indicate modules trained in the corresponding step; snowflakes indicate frozen learned modules. Dashed orange arrows indicate training supervision.
Fig. 3: Qualitative comparison of commanded rail-contact acquisition. Snapshots at t=2.56s compare the policies evaluated in Table II : MaskedMimic hand targeted to rail (top left), GPC (top right), GPC w/o support observation (bottom left), iGPC (ours, bottom right).
Policy
Interaction pre-training
Affordance pre-training
Crown contact (%) ↑
Endpoint error p50/p90 (cm) ↓
Time to crown p50/p90 (s) ↓
MaskedMimic (Hand Target)
—
—
0.20
12.07/15.55
1.45/2.75
GPC (w/o Interaction)
×
×
31.30
16.03/36.43
4.90/9.00
GPC (w/o Affordance)
✓
×
48.63
15.03/32.88
3.93/8.44
iGPC (ours)
✓
✓
99.32
1.35/11.04
0.76/1.46
TABLE II: Expert policy evaluation for commanded rail-contact task.
Policy
Obs. loss
Token loss
Crown contact (%) ↑
Body collision (%) ↓
Endpoint error p50/p90 (cm) ↓
Teacher-target MAE (rad) ↓
Expert iGPC
—
—
99.32
15.19
1.35/11.04
—
Student MLP
—
—
46.53
50.73
11.19/86.78
0.178396
iGPC (no obs. or token)
×
×
25.39
17.19
51.78/146.20
0.127518
iGPC (no obs.)
×
✓
93.41
13.13
1.37/8.38
0.088503
iGPC (no token)
✓
×
90.14
10.55
2.68/16.22
0.088679
iGPC (no prior calibration)
✓
✓
91.02
10.79
1.62/13.53
0.085493
TABLE III: Perception-based student policy evaluation for commanded rail-contact task.
Fig. 4: Reactive stabilization in simulation. Top: leg-only GPC. Bottom: iGPC student with sensor visualizations. The red arrow indicates the applied disturbance.
(a) Reactive stabilization : physical falls ↓
Push Δv (m/s)
iGPC expert
iGPC student
GPC leg-only
1.5–2.0
3 (0.59%)
5 (0.98%)
12 (2.34%)
2.0–2.5
24 (4.69%)
25 (4.88%)
94 (18.36%)
2.5–3.0
116 (22.66%)
88 (17.19%)
207 (40.43%)
(b) Box pushing
Metric
iGPC expert
iGPC student
VisualMimic
TABLE IV: Whole-body interaction. (a) Physical falls, count (%), on 512 trials per push band. (b) Forward displacement and absolute lateral drift at first fall or 60 s: mean ± SD across four evaluation-seed means (128 trials each).
Current humanoid reinforcement-learning policies excel at free-space motions but struggle with contact-rich tasks, as pure kinematic tracking cannot resolve the physical ambiguities of interacting with objects and uneven terrain. To address this, we introduce SceneBot, a unified motion-tracking framework capable of handling freespace locomotion, terrain traversal, and whole-body manipulation. SceneBot conditions a single policy on both reference motions and per-link contact labels, explicitly defining expected environmental interactions. To overcome the lack of annotated interaction data, we propose a hindsight scene reconstruction approach that infers scene-interaction graphs from retargeted human motion. Trained on 7.5 hours of this reconstructed, contact-rich data, SceneBot successfully generalizes to unseen motions and environments. Our results demonstrate that SceneBot is the first general framework to seamlessly unify free-space and contact-rich behaviors executing complex, long-horizon tasks like carrying a box upstairs and establishing contact conditioning as a powerful interface for humanoid control. All code and data will be open-sourced. More demos and information are available at: https://ericcsr.github.io/scenebot/
Sirui Chen, Shibo Zhao, Zhen Wu +3
Stanford, Amazon FAR United States · Amazon FAR United States · CMU, Amazon FAR United States
Humanoid-Object Interaction (HOI) is a fundamental capability for humanoid robots, yet it remains challenging due to the tight coupling between dynamic balance and stable interaction with diverse objects. Existing methods often require time-consuming task-specific policy training or rely on rigid trajectory replay, which limits their ability to accommodate novel interaction scenarios. In this work, we present \textit{GenHOI}, a simple yet effective framework that enables humanoid robots to perform diverse object-interaction tasks in a zero-shot manner by directly imitating a single generated video, without task-specific training or physical demonstration data. GenHOI first reconstructs the robot-object scene in simulation and renders a first-frame image, which, together with the language command, conditions the synthesis of a task-oriented interaction video. The generated video is then analyzed to identify interaction-relevant contact events and estimate hand-object contact regions, which are encoded as object-centric geometric constraints that convert visual interaction cues into physically grounded optimization priors. Guided by these priors, the reference motion recovered from the video is refined and smoothed to resolve the scale ambiguity inherent in 2D video generation, while adapting a single reference trajectory to unseen robot-object relative poses. The optimized trajectory is finally executed by a closed-loop tracking controller. We validate the proposed framework in extensive simulation and real-world experiments across diverse object-interaction tasks, including box grasping, asymmetric bimanual chair carrying, table lifting from below, and cylindrical-object enveloping.
Zhihai Bi, Qiang Zhang, Guoyang Zhao +8
The Hong Kong University of Science and Technology (Guangzhou) · Artificial General Intelligence Institute, University of Science and Technology of China · The University of Hong Kong +1
Generating physically plausible dynamic motions of human-object interaction (HOI) remains challenging, mainly due to existing HOI datasets limited to static interactions, and pretrained agents capable of either dynamic full-body motions without objects or static HOI motions. Recent works such as InsActor and CLoSD generate HOI motions in planning and execution stages, are yet limited to either static or short-term contacts e.g. striking. In this work, we propose a framework that fulfills dynamic and long-term interaction motions such as running while holding a table, by combining pretrained motion priors and imitation agents in planning and execution stages. In the planning stage, we augment HOI datasets with dynamic priors from a pretrained human motion diffusion model, followed by object trajectory generation. This plans dynamic HOI sequences. In the execution stage, a composer network blends actions of pretrained imitation agents specialized either for dynamic human motions or static HOI motions, enabling spatio-temporal composition of their complementary skills. Our method over relevant prior-arts consistently improves success rates while maintaining interaction for dynamic HOI tasks. Furthermore, blending pretrained experts with our composer achieves competitive performance in significantly reduced training time. Ablation studies validate the effectiveness of our augmentation and composer blending.