RobotAPO: Adversarial Physics Preference Optimization for Robotic Manipulation Video Generation
Organizations: Institute of Automation, Chinese Academy of Sciences · GigaAI · Beijing Institute of Technology
Abstract
Robotic manipulation videos are increasingly used as visual plans for embodied agents, but optimizing purely for visual plausibility often fails to capture the fragile physical manifold of real-world interactions. Even minor physics-violating errors at the interaction boundary, such as interpenetration or premature object motion, can completely invalidate the inferred timing and pose needed for downstream execution. Because standard supervised fine-tuning lacks the direct pressure to penalize these localized failures, we introduce AgiBot-PhysPref. This rigorously curated 10,000-sample preference dataset isolates condition-matched physics violations, turning the generator's own failure distribution into a foundational signal for physical consistency. Building upon this, we propose RobotAPO, an adversarial physics preference optimization framework operating in the continuous flow-matching denoising space. To prevent the policy from merely memorizing static curated failures, RobotAPO employs a lightweight adversarial counterfactual proposer that learns a condition-dependent, physical-failure-biased direction in denoising space. This encourages the model to explore and better respect the physical interaction boundary, all while maintaining a pure prompt-and-reference inference interface without requiring external structural conditioning. Comprehensive evaluations demonstrate that explicitly correcting these localized physics violations improves downstream robot execution from generated videos. On held-out AgiBot conditions, RobotAPO outperforms the strongest controlled internal baseline in physical consistency by 6.8% hard score and 10.0% soft score. Crucially, in real-robot replay, it translates these physical-consistency gains into a 37.4% relative improvement in task success over the strongest controlled internal baseline.
Figures & tables
| \toprule Property | Value |
|---|---|
| \midrule Generated candidates | |
| Human-reviewed candidates | |
| Final curated preference samples | |
| Task IDs | dual-arm AgiBot scenarios |
| Screening model | Qwen3-VL-30B-Instruct |
| Final label source | Human double annotation with disagreement review |
| \toprule Method | Size | Visual-H | Align-H | Phys-H | Visual-S | Align-S | Phys-S |
|---|---|---|---|---|---|---|---|
| \midrule \textcolor grayLVP ( Chen et al., 2025 ) | \textcolor gray14B | \textcolor gray3.647 | \textcolor gray3.787 | \textcolor gray3.473 | \textcolor gray3.030 | \textcolor gray3.252 | \textcolor gray3.363 |
| \midrule Wan2.2-TI2V | 5B | 3.287 | 3.493 | 2.967 | 2.812 | 3.064 | 2.891 |
| SFT baseline | 5B | 3.347 | 3.767 | 3.193 | 2.984 | 3.410 | 3.041 |
| Diffusion-DPO ( Wallace et al., 2024 ) | 5B | 3.340 | 3.760 | 3.253 | 3.120 | 3.484 | 3.172 |
| RealDPO ( Cheng et al., 2025 ) | 5B | 3.347 | 3.760 | 3.173 | 3.047 | 3.435 | 3.008 |
| \midrule \texorpdfstring RobotAPO RobotAPO (Ours) | 5B | 3.553 | 3.927 | 3.473 | 3.564 | 3.939 | 3.489 |
| \toprule Method | Size | Grasp Success (%) | Cond. Place Success (%) | Task Success (%) | Task Success 95% CI |
|---|---|---|---|---|---|
| \midrule \textcolor grayLVP | \textcolor gray14B | \textcolor gray72.5 | \textcolor gray86.5 | \textcolor gray62.7 | \textcolor gray[49.0, 74.7] |
| \midrule SFT baseline | 5B | 56.9 | 58.6 | 33.3 | [22.0, 47.0] |
| Diffusion-DPO | 5B | 64.7 | 69.7 | 45.1 | [32.3, 58.6] |
| RealDPO | 5B | 66.7 | 70.6 | 47.1 | [34.1, 60.5] |
| \texorpdfstring RobotAPO RobotAPO (Ours) | 5B | 78.4 | 82.5 | 64.7 | [51.0, 76.4] |
| \toprule Method | Visual-S | Align-S | Phys-S | Grasp Success (%) | Task Success (%) |
|---|---|---|---|---|---|
| \midrule SFT baseline | 2.984 | 3.410 | 3.041 | 56.9 | 33.3 |
| \texorpdfstring RobotAPO RobotAPO w/o CF/Adv | 3.194 | 3.614 | 3.134 | 68.6 | 49.0 |
| \texorpdfstring RobotAPO RobotAPO w/o Anchor/Reg | 3.000 | 3.326 | 3.079 | 60.8 | 35.3 |
| \texorpdfstring RobotAPO RobotAPO (Ours) | 3.564 | 3.939 | 3.489 | 78.4 | 64.7 |
| \toprule | Visual-H | Align-H | Phys-H | Visual-S | Align-S | Phys-S |
|---|---|---|---|---|---|---|
| \midrule | 3.407 | 3.793 | 3.300 | 3.284 | 3.704 | 3.270 |
| 3.460 | 3.853 | 3.373 | 3.392 | 3.784 | 3.351 | |
| 3.553 | 3.927 | 3.473 | 3.564 | 3.939 | 3.489 | |
| 3.433 | 3.827 | 3.420 | 3.365 | 3.758 | 3.421 |
| \toprule Method | Visual-H | Visual-S | Phys-H | Phys-S |
|---|---|---|---|---|
| \midrule SFT | 3.347 | 2.984 | 3.193 | 3.041 |
| Diffusion-DPO | 3.240 | 3.211 | 3.185 | 3.061 |
| RealDPO | 3.225 | 3.212 | 3.180 | 3.092 |
| \texorpdfstring RobotAPO RobotAPO (Ours) | 3.355 | 3.345 | 3.310 | 3.293 |
| \bottomrule |
| \toprule Diagnostic | Adversarial | Gaussian | Paired effect [95% CI] |
|---|---|---|---|
| \midrule Condition-preserving local physical failure | 37.5% | 20.3% | pp |
| Global corruption / condition change | 9.4% | 6.3% | pp |
| Unusable | 3.1% | 1.6% | pp |
| \bottomrule |
| \toprule Method | Visual-S | Phys-H | Phys-S | Gain [95% CI] |
|---|---|---|---|---|
| \midrule SFT | ||||
| Diffusion-DPO | ||||
| RealDPO | ||||
| \texorpdfstring RobotAPO RobotAPO | – | |||
| \bottomrule |
| \toprule Failure Type | Definition | Typical Evidence |
|---|---|---|
| \midrule Object interpenetration | Robot, object, or scene geometry passes through another entity. | Gripper penetrates object, or object sinks into table. |
| \addlinespace [2pt] Entity consistency error | Robot or object identity changes in appearance, shape, or existence over time. | Object flickers, disappears, changes shape, or changes identity across frames. |
| \addlinespace [2pt] Contact-causality error | Object motion or state transition is not causally supported by valid contact or manipulation, including premature motion before effective contact and action/state changes without effective contact. | Object moves before contact, state changes without effective contact, or gripper closes while the object slides without being grasped. |
| \addlinespace [2pt] Unsupported floating after release | An object remains suspended after losing support. | Released object stays in mid-air or does not settle naturally. |
| \addlinespace [2pt] Invalid interaction / other | No valid manipulation event under the task condition, or other visible failure not covered above. | Robot grasps air, never interacts with the target, or executes a misaligned motion that breaks the task. |
| \bottomrule |