Video Action Models (VAMs) couple visual dynamics modeling with action generation for robot manipulation. However, video representations are not naturally suited to action generation, as exposing the action policy to excessive visual detail can impair its generalization ability. Therefore, we introduce IronMan (Information-constRained videO-actioN learning for robot MANipulation), a robust video-action learning framework built on the information bottleneck principle. The core principle of this framework is to impose information constraints that suppress irrelevant visual information while preserving action-relevant dynamics cues. IronMan employs a dynamics-aware bottleneck that distills noisy, entangled one-step video features into compact world representations. Extensive simulation and real-world experiments demonstrate strong in-distribution (ID) performance and out-of-distribution (OOD) robustness while maintaining efficient inference. IronMan achieves success rates of 99.0% on LIBERO and 79.4% on RoboTwin clean2clean, outperforming all the evaluated baselines. Under OOD shifts, IronMan achieves a success rate of 79.1% on LIBERO-Plus, exceeding the strongest baseline by 10.4 percentage points. Project page: https://youngsoul0731.github.io/ironman-project-page/
Figures & tables
Figure 1: Motivation of IronMan. Existing video–action interfaces allow action policy to freely attend to irrelevant visual signals, increasing its sensitivity to visual details. IronMan instead uses an information-constrained bottleneck designed to retain the predictive dynamics needed for action while suppressing redundancy.
Figure 2: Overview of IronMan. A dynamics-aware bottleneck transforms one-step video features into compact world tokens using a spatio-temporal transformer and a Gaussian head. The action policy attends to world tokens and robot state tokens to generate actions; video decoding is optional and inactive during action inference.
Models
Em. PT.
LIBERO
RoboTwin c2c
Spatial
Object
Goal
Long
Avg.
π0.5
✓
98.8
98.2
98.0
92.4
96.9
70.7
X-VLA
✓
98.2
98.6
97.8
97.6
98.1
68.0
Abot-M0
✓
98.8
99.8
99.0
96.6
98.6
57.4
starVLA
✗
98.7
99.7
98.6
94.2
97.8
46.5
FastWAM
✗
98.2
100.0
97.0
95.2
97.6
77.8
Table 1: In-domain performance on LIBERO and RoboTwin clean2clean (c2c). Success rates are reported in percentages. Em. PT. denotes embodied pretraining. Bold and underlined values indicate the best and second-best results in each column, respectively.
Models
LIBERO-Plus
RoboTwin c2r
Latency (ms)
Camera
Robot
Language
Light
Background
Noise
Layout
Overall
DiT4DiT
59.7
61.1
88.4
90.8
33.7
60.0
83.6
68.7
–
180
FastWAM
16.8
43.7
68.1
77.2
51.5
36.5
59.5
49.2
1.9
190
FastWAM-Joint
41.3
65.5
91.0
86.4
53.8
56.2
79.2
67.4
5.2
580
FastWAM-IDM
37.7
67.2
90.3
92.7
54.2
56.6
79.2
67.9
9.3
810
OpenWAM-Joint
1.2
53.4
71.1
77.3
43.8
18.4
52.2
43.7
13.2
429
Table 2: Performance on LIBERO-Plus and RoboTwin clean2random (c2r), and inference latency. Success rates are reported in percentages. Bold and underlined success rates indicate the best and second-best results in each column, respectively.
DB arch.
IB obj.
Camera
Robot
Language
Light
Background
Noise
Layout
Overall
✗
✗
49.8
77.1
93.2
87.0
63.3
61.3
82.9
73.4
✗
✓
56.0
73.6
90.8
83.1
65.5
61.7
82.0
73.2
✓
✗
54.7
75.9
93.4
85.2
63.9
64.2
82.8
74.3
✓
✓
64.4
75.9
92.7
93.4
64.4
78.3
83.7
79.1
Table 3: Ablation study of the dynamics-aware bottleneck (DB) architecture and information bottleneck (IB) objective across seven types of perturbations in LIBERO-Plus. Bold values indicate the best result in each column.
Figure 3: Bottleneck mechanism analysis on LIBERO-Plus. Success rates (%) are shown for different bottleneck capacities Q×D and information bottleneck constraint strengths β .
Figure 4: Comparison of attention heatmaps for IronMan and the baseline without the dynamics-aware bottleneck under a background shift from LIBERO (ID) to LIBERO Plus (OOD).
Figure 5: Real-world performance under ID, Background, and Layout settings.
Figure 6: Comparison of the flower insertion task under background shift.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Table 4: Training configurations on LIBERO and RoboTwin Clean.
Figure 7: Attention visualization for grasping and lifting the cup.
Figure 8: Attention visualization for grasping and lifting the ketchup bottle.
Figure 9: Attention visualization for grasping and lifting the moka pot.
Figure 10: Attention visualization for grasping and lifting the soup can.
Task
Score
Completion criterion
insert flowers
0
The first flower is not placed in the vase.
3.33
The first flower is placed in the vase.
6.67
The second flower is placed in the vase.
10
The third flower is placed in the vase.
arrange fruits
0
The banana is not placed on the plate.
2.5
The banana is placed on the plate.
Appendix
Table 5: Sequential scoring criteria for the three real-world tasks. Entries are raw per-trial scores on a 0–10 scale; reported task scores are normalized to a 0–100 scale.
Figure 11: Real-world task demonstrations.
Figure 12: Action prediction error to visual perturbations. Box plots show the L1 error relative to each model’s clean reference under four categories of visual perturbation, with 40 values per category and model across four LIBERO tasks and 10 perturbation configurations per task.
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · Beijing Innovation Center of Humanoid Robotics