Active vision manipulation requires a policy to control both its camera and its end-effectors, yet camera motion determines which evidence remains visible within finite observation windows. Acquiring a new view can displace task-critical cues, while retaining a view forgoes potentially useful observations. We formulate this as an evidence-aware retain--acquire problem and present ActiveWAM, a unified world--action model that learns observation and manipulation jointly. To this end, we propose training-time inversion which constrains a frozen video prior by task-bearing source evidence and visible temporal changes, eliminating the need for test-time inversion or candidate ranking. At deployment, the policy generates bimanual and pan/tilt actions-including stay and reacquisition behaviors-from view-aware history, and updates context from newly measured RGB observations. Future-video prediction serves as a co-training signal, while action generation requires neither future-video decoding nor optimal viewpoint annotations. We introduce RoboTwin-AV, a 50-task benchmark with executable pan/tilt control and automatically generated demonstrations. ActiveWAM improves TAVIS out-of-distribution success by up to 17.0 percentage points over the strongest baselines, achieves 20.0 additional points over Fast-WAM on RoboTwin-AV, and outperforms it by 26.7 points on real-world physical kitchen tasks.
Figures & tables
Figure 1: (a) Fixed-vision WAM generates arm actions from static camera observations. (b) Active-vision policy controls head separately from manipulation. (c) ActiveWAM unifies head and arm generation in a common world-action model, using training-time history inversion to preserve task-relevant evidence while learning executable pan/tilt control. Task examples: simulation (RoboTwin-AV) shows object placement requiring spatial reframing; physical kitchen setup demonstrates composed multi-stage manipulation. Bar charts: ActiveWAM improves TAVIS success by 17.0pp on Head/GR1T2, RoboTwin-AV compound success by 20.0pp over Fast-WAM, and physical success by 26.7pp over Fast-WAM.
Figure 2: ActiveWAM overview. Active vision manipulation is treated as evidence-aware retain–acquire control. Training-time source-constrained inversion transforms observed histories, while shared WAM passes receive raw or transformed conditions with identical original future/action targets. View-aware history and current RGB condition encoders supply recent evidence and its camera context; the action branch determines executable arm and head control, with future-video prediction used as a co-training signal. Deployment retains raw inputs, one action-generation pass, and actual new measurements after a short executed prefix. It uses neither inversion nor candidate ranking, and requires no manually annotated optimal viewpoints or information-gain objective.
Figure 3: Representative head-view sequences with pan/tilt configurations: placing bread in a pan in RoboTwin-AV (top); pouring oil and egg mixture, then stirring on the real robot (bottom). ActiveWAM learns to expose task-critical spatial relations (bread-to-skillet alignment, egg-to-pan trajectory) through head motion while retaining previously acquired evidence within its history window.
Head / GR1T2
Head / Reachy2
Hands / GR1T2
Hands / Reachy2
Method
ID
S
P
ID
S
P
ID
S
P
ID
S
P
π0 [ 2 ]
47.1
27.9
1.9
38.8
22.1
11.2
70.1
49.0
14.6
76.7
51.0
38.2
π0.5 [ 18 ]
51.0 ±1.0
28.7 ±0.6
7.7 ±0.6
44.0 ±0.0
27.0 ±1.0
16.3 ±0.6
73.3 ±1.2
54.3 ±1.5
18.7 ±0.6
80.0 ±1.0
55.0 ±1.0
43.7 ±1.2
Fast-WAM
47.0 ±2.0
30.0 ±1.0
6.0 ±1.0
41.0 ±1.0
27.0 ±2.0
17.0 ±2.0
70.7 ±1.2
51.0 ±1.0
24.3 ±1.5
76.0 ±1.0
52.0 ±2.6
41.3 ±0.6
EasyWAM-Unified
51.3 ±1.5
30.0 ±2.0
11.7 ±1.5
46.0 ±1.0
27.3 ±1.2
20.7 ±1.2
68.0 ±2.6
53.3 ±0.6
23.7 ±0.6
76.3 ±0.6
52.7 ±0.6
39.7 ±2.9
ActiveWAM (ours)
63.7 ±2.1
47.0 ±1.0
24.7 ±0.6
52.7 ±1.5
44.0 ±2.6
31.3 ±1.5
75.7 ±1.2
64.0 ±0.0
36.3 ±0.6
79.7 ±3.1
63.3 ±0.6
54.7 ±1.5
Table 1: Success on TAVIS (%). The first row reproduces the official reported π0 results. Our evaluations retain native demonstrations and sensor access and evaluate a fixed checkpoint under three environment seeds, with 96 episodes/task/seed per condition. ID/S/P denote in-distribution, spatial OOD, and initial-pose OOD.
Method
Clean
App.
Pose
Both
ACT
31.3 ±1.5
8.7 ±1.2
20.0 ±1.0
3.7 ±0.6
Diffusion Policy
34.7 ±0.6
11.3 ±1.2
16.3 ±0.6
5.0 ±1.0
π0.5
63.0 ±1.0
42.3 ±1.2
43.0 ±1.7
34.0 ±1.0
Fast-WAM
76.3 ±1.2
41.7 ±1.5
56.0 ±1.0
33.3 ±1.2
EasyWAM-Unified
75.0 ±1.0
46.3 ±1.5
56.0 ±0.0
33.3 ±1.2
AV-ALOHA
39.3 ±1.2
25.7 ±0.6
26.3 ±0.6
19.3 ±1.5
Table 2: RoboTwin-AV SR (%): 50 tasks, three seeds, 100 episodes/task/seed/condition; shared sensor/action interfaces. App./Pose/Both: appearance/head-pose/compound shifts.
Variant
Clean
Compound
ActiveWAM (full)
80.3 ±2.1
53.3 ±1.5
w/o history
73.7 ±0.6
34.0 ±0.0
w/o inversion (raw history)
76.7 ±1.5
41.7 ±1.5
Same-Wan framewise inversion
79.7 ±0.6
43.5 ±1.0
w/o task preservation
78.7 ±0.6
41.0 ±1.0
w/o dynamic preservation
82.0 ±1.0
44.7 ±1.2
Table 3: RoboTwin-AV ablations (SR, %).
Head controller
Clean
Pose
Co-vis.
Travel
Locked
76.3 ±0.6
45.0 ±1.0
61.0 ±1.0
0.0
Object tracking
76.3 ±0.6
57.3 ±1.2
73.7 ±0.6
4.1 ±0.1
Relation look-at + stay
79.3 ±0.6
61.0 ±0.0
81.0 ±1.0
2.5 ±0.1
Joint learned (ours)
80.3 ±2.1
65.0 ±3.0
87.7 ±0.6
1.7 ±0.1
+ matched grounding
81.3 ±0.6
62.7 ±0.6
86.3 ±3.8
1.9 ±0.1
Table 4: Head control with inversion fixed. Clean and Pose give SR (%); Co-vis.: object–goal co-visibility (%); travel: rad/episode.
Variant
LIBERO
LIBERO-Plus
Raw history
98.4 ±0.8
68.2 ±0.1
Task-safe augmentation
98.5 ±0.2
72.5 ±0.1
Source-feature auxiliary
98.7 ±0.1
71.9 ±0.2
ActiveWAM
98.7 ±0.1
76.9 ±0.1
Table 5: Fixed-camera checks (SR, %). LIBERO-Plus weights all 10,030 variants by category counts.
Inversion
Head
Clean
Compound
Off
Locked
74.3 ±0.6
23.7 ±1.2
On
Locked
76.3 ±0.6
30.3 ±0.6
Off
Active
76.7 ±1.5
41.7 ±1.5
On
Active
80.3 ±2.1
53.3 ±1.5
Video loss
Inference
Clean
Compound
On
Joint
80.3 ±2.1
53.3 ±1.5
Table 6: RoboTwin-AV attribution/sensitivity (SR, %): trajectory-consistent factorial data and matched training/inference graphs. Coverage: accepted/attempted inversion windows.
Controller
SR (%)
Co-vis. (%)
Time (s)
Travel (rad)
Stale (%)
Locked head
45.0 ±1.0
61.0 ±1.0
10.0
0.0
1.2 ±0.2
Target centering
57.3 ±1.2
73.7 ±0.6
2.3 ±0.1
4.1 ±0.1
8.2 ±0.3
Relation + stay
61.0 ±0.0
81.0 ±1.0
1.9 ±0.1
2.5 ±0.1
5.1 ±0.1
Joint learned
62.7 ±0.6
86.3 ±3.8
1.6 ±0.1
1.9 ±0.1
4.2 ±0.3
No extra sense arm
62.0 ±1.7
88.0 ±1.0
1.6 ±0.0
1.9 ±0.1
4.2 ±0.2
Arms busy
54.7 ±0.6
76.3 ±1.2
2.1 ±0.1
2.4 ±0.1
6.2 ±0.2
Table 7: Controller diagnostics. RoboTwin-AV pose-shift rows use matched RGB grounding. Co-vis./stale are percentages; time is timeout-capped reacquisition (10 s). Arms-busy is evaluated on the corresponding TAVIS slice; the oracle is an upper bound.
Cucumber
Egg
Mixture
Method
S1
Total
S1
Total
S1
Total
π0.5 [ 18 ]
8
2
10
7
12
11
Fast-WAM
7
1
11
6
11
7
EasyWAM-Unified
7
1
10
7
12
9
Ours w/o inversion
10
4
12
7
14
11
ActiveWAM (full)
11
7
15
11
14
12
Table 8: Physical kitchen successes out of 20 attempts per task (60 per method and 300 across all five methods). Stage 1: first subtask only; Total: both subtasks completed in the same trial.