IronMan: Information-Constrained Video-Action Learning for Robot Manipulation
Organizations: Shanghai Jiao Tong University · Joy Future Academy, JD
Abstract
Video Action Models (VAMs) couple visual dynamics modeling with action generation for robot manipulation. However, video representations are not naturally suited to action generation, as exposing the action policy to excessive visual detail can impair its generalization ability. Therefore, we introduce IronMan (Information-constRained videO-actioN learning for robot MANipulation), a robust video-action learning framework built on the information bottleneck principle. The core principle of this framework is to impose information constraints that suppress irrelevant visual information while preserving action-relevant dynamics cues. IronMan employs a dynamics-aware bottleneck that distills noisy, entangled one-step video features into compact world representations. Extensive simulation and real-world experiments demonstrate strong in-distribution (ID) performance and out-of-distribution (OOD) robustness while maintaining efficient inference. IronMan achieves success rates of 99.0% on LIBERO and 79.4% on RoboTwin clean2clean, outperforming all the evaluated baselines. Under OOD shifts, IronMan achieves a success rate of 79.1% on LIBERO-Plus, exceeding the strongest baseline by 10.4 percentage points. Project page: https://youngsoul0731.github.io/ironman-project-page/
Figures & tables
| Models | Em. PT. | LIBERO | RoboTwin c2c | ||||
|---|---|---|---|---|---|---|---|
| Spatial | Object | Goal | Long | Avg. | |||
| ✓ | 98.8 | 98.2 | 98.0 | 92.4 | 96.9 | 70.7 | |
| X-VLA | ✓ | 98.2 | 98.6 | 97.8 | 97.6 | 98.1 | 68.0 |
| Abot-M0 | ✓ | 98.8 | 99.8 | 99.0 | 96.6 | 98.6 | 57.4 |
| starVLA | ✗ | 98.7 | 99.7 | 98.6 | 94.2 | 97.8 | 46.5 |
| FastWAM | ✗ | 98.2 | 100.0 | 97.0 | 95.2 | 97.6 | 77.8 |
| Models | LIBERO-Plus | RoboTwin c2r | Latency (ms) | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Camera | Robot | Language | Light | Background | Noise | Layout | Overall | |||
| DiT4DiT | 59.7 | 61.1 | 88.4 | 90.8 | 33.7 | 60.0 | 83.6 | 68.7 | – | 180 |
| FastWAM | 16.8 | 43.7 | 68.1 | 77.2 | 51.5 | 36.5 | 59.5 | 49.2 | 1.9 | 190 |
| FastWAM-Joint | 41.3 | 65.5 | 91.0 | 86.4 | 53.8 | 56.2 | 79.2 | 67.4 | 5.2 | 580 |
| FastWAM-IDM | 37.7 | 67.2 | 90.3 | 92.7 | 54.2 | 56.6 | 79.2 | 67.9 | 9.3 | 810 |
| OpenWAM-Joint | 1.2 | 53.4 | 71.1 | 77.3 | 43.8 | 18.4 | 52.2 | 43.7 | 13.2 | 429 |
| DB arch. | IB obj. | Camera | Robot | Language | Light | Background | Noise | Layout | Overall |
|---|---|---|---|---|---|---|---|---|---|
| ✗ | ✗ | 49.8 | 77.1 | 93.2 | 87.0 | 63.3 | 61.3 | 82.9 | 73.4 |
| ✗ | ✓ | 56.0 | 73.6 | 90.8 | 83.1 | 65.5 | 61.7 | 82.0 | 73.2 |
| ✓ | ✗ | 54.7 | 75.9 | 93.4 | 85.2 | 63.9 | 64.2 | 82.8 | 74.3 |
| ✓ | ✓ | 64.4 | 75.9 | 92.7 | 93.4 | 64.4 | 78.3 | 83.7 | 79.1 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Score | Completion criterion |
|---|---|---|
| insert flowers | 0 | The first flower is not placed in the vase. |
| 3.33 | The first flower is placed in the vase. | |
| 6.67 | The second flower is placed in the vase. | |
| 10 | The third flower is placed in the vase. | |
| arrange fruits | 0 | The banana is not placed on the plate. |
| 2.5 | The banana is placed on the plate. |