Organizations: The Hong Kong University of Science and Technology · The University of British Columbia · MMLab, The Chinese University of Hong Kong · Astribot · ZENBOT
Learning-based manipulation policies relying on RGB cameras often suffer from degraded observations under extreme exposure. Event cameras mitigate this degradation by asynchronously detecting pixel-level intensity changes to offer a high dynamic range. However, their observations heavily depend on camera placement, as fixed cameras miss static scene content while wrist-mounted camera motion causes previously visited regions to leave the field of view. To address these spatial-temporal limitations, we present ECHO (Event-augmented Context with Hindsight and Outlook), a wrist-only latent world action model that encodes wrist events into compact motion representations to provide temporal and spatial context for policy reasoning. Specifically, ECHO utilizes a pretrained event encoder to explain visual-feature changes between frames. Its hindsight module preserves the gripper trajectory with past event stream as addressable off-camera context. Concurrently, the outlook module introduces learnable event foresight queries supervised to anticipate the event window for future actions, enabling the policy to predict upcoming scene changes. Evaluated on wrist-only RLBench tasks, ECHO outperforms RGB and RGB+event baselines by 20.6 and 12.0 percentage points under normal lighting, and by 14.6 and 11.3 points under severe exposure drops, respectively, while also surpassing RGB references using a third-person camera. Real-world experiments with a wrist-mounted event camera validate that ECHO outperforms RGB-only and RGB+event baselines across multiple tasks under both nominal and severely dark lighting. Project page is at https://echo-wam.github.io/.
Figures & tables
Fig. 2: Overview of ECHO. (a) The event encoder is pretrained to represent the visual transition between consecutive RGB observations using frozen DINOv3 features and a latent world model. (b) During policy learning, wrist RGB, current event latent tokens Et , language, off-camera context Ct built from visited poses and current events, and foresight queries Ft supervised by future event latents form the vision-language prefix. A flow-matching action expert then predicts continuous action chunks from this prefix and proprioception.
Fig. 3: Event encoder pretraining. The event volume is encoded into an event latent that can (i) reconstruct the frozen RGB teacher’s transition through a warp and synthesis head, and (ii) separate true motion window from its reversed twin under a motion-direction contrast (MDC).
Fig. 4: Qualitative visualization of future predictions. Current event and RGB observations are shown with future ground truth and decoded predictions. For visualization only, lightweight ViT heads decode events from foresight queries and reconstruct RGB images from future teacher features predicted by the latent world model conditioned on current teacher features and foresight queries.
Fig. 5: Qualitative attention visualization of different ECHO components. Under large wrist-camera motion where action tokens attend heavily to background movement, memory tokens concentrate precisely at gripper–object contact points. Meanwhile, foresight queries cover broader surrounding regions to anticipate interaction dynamics.
Method
Event
Pretrained
Traj-ctx
Event-ctx
Foresight
laptop
toilet
fridge
umbrella
frame
plants
SR (%)
N ∣
D
N ∣
D
N ∣
D
N ∣
D
N ∣
D
N ∣
D
N ∣
D
π0 (Front only)
16 ∣
17
24 ∣
17
25 ∣
18
12 ∣
3
5 ∣
4
1 ∣
6
55.3 ( − 2.0) ∣
43.3 (+1.3)
π0 (Front + wrist)
19 ∣
9
24 ∣
17
20 ∣
20
15 ∣
5
11 ∣
3
8 ∣
2
64.7 (+7.4) ∣
37.3 ( − 4.7)
π0 (Wrist only)
18 ∣
17
24 ∣
12
20 ∣
20
4 ∣
5
3 ∣
1
4 ∣
3
48.7 ( − 8.6) ∣
38.7 ( − 3.3)
CogACT (Wrist only)
23 ∣
10
20 ∣
17
18 ∣
19
0 ∣
0
8 ∣
4
3 ∣
2
48.0 ( − 9.3) ∣
34.7 ( − 7.3)
π0 + event
✓
✓
24 ∣
20
24 ∣
14
22 ∣
17
5 ∣
9
3 ∣
0
8 ∣
3
57.3 (+0.0) ∣
42.0 (+0.0)
TABLE I: Success rate (%) on six RLBench tasks using the wrist-only camera and event protocol of Sec. V . Results are reported under normal lighting (left) and a −4 EV exposure shift (right).
Fig. 6: Qualitative reconstruction examples of the pretrained event encoder, on a simulation sequence and a real-robot sequence.
Normal Lighting
Severe Dark Lighting
Method
Perception
pick-place
bell-ring
pick-place
bell-ring
π0
wrist RGB
80 (+5)
15 ( − 15)
55 ( − 5)
5 ( − 10)
π0 + event
wrist RGB + event
75 (+0)
30 (+0)
60 (+0)
15 (+0)
ECHO
wrist RGB + event
90 (+15)
45 (+15)
80 (+20)
30 (+15)
TABLE II: Real-robot success rate (%) across tasks under normal and severe dark lighting ( 20 trials per cell). Red parenthesized deltas are relative to the π0 + event baseline.
Fig. 7: Real-robot setup for the tasks Ring the Bell Twice and Put the Toy Bee into the plate .
Fig. 8: Real-robot execution of ring the bell twice under normal and dark lighting. Wrist RGB (top) and events from the same interval (bottom) show contact between the gripper and bell.
Vision-Language-Action (VLA) models have become an important paradigm of embodied AI. However, existing VLA models typically assume well-lit and stable indoor settings, while real-world embodied manipulation may involve degraded RGB observations caused by illumination shifts, posing critical challenges for robust robotic manipulation. To address this gap, we propose \textbf{Event-VLA}, an event-enhanced VLA framework for generalizable manipulation across varying illumination conditions. We formulate VLA-based manipulation under degraded visibility as a practical robustness problem for RGB-centric policies, and introduce event streams as an illumination-robust, motion-sensitive complementary observation to improve robustness across visibility levels. Specifically, unlike conventional multimodal fusion that directly merges event features into the global semantic token space, Event-VLA injects event information through an action-query routing pathway. It uses learnable action queries to extract task-relevant semantics from the VLA reasoning process, and selectively aggregates event tokens via gated cross-attention to construct event-aware action representations. This design preserves the pretrained RGB-language semantic priors while effectively leveraging event information for robust action prediction. Experiments in simulation and real-world deployment show that Event-VLA maintains strong manipulation performance under normal lighting and improves success rates under low-light degradation and near-dark real-world settings.
Active vision manipulation requires a policy to control both its camera and its end-effectors, yet camera motion determines which evidence remains visible within finite observation windows. Acquiring a new view can displace task-critical cues, while retaining a view forgoes potentially useful observations. We formulate this as an evidence-aware retain--acquire problem and present ActiveWAM, a unified world--action model that learns observation and manipulation jointly. To this end, we propose training-time inversion which constrains a frozen video prior by task-bearing source evidence and visible temporal changes, eliminating the need for test-time inversion or candidate ranking. At deployment, the policy generates bimanual and pan/tilt actions-including stay and reacquisition behaviors-from view-aware history, and updates context from newly measured RGB observations. Future-video prediction serves as a co-training signal, while action generation requires neither future-video decoding nor optimal viewpoint annotations. We introduce RoboTwin-AV, a 50-task benchmark with executable pan/tilt control and automatically generated demonstrations. ActiveWAM improves TAVIS out-of-distribution success by up to 17.0 percentage points over the strongest baselines, achieves 20.0 additional points over Fast-WAM on RoboTwin-AV, and outperforms it by 26.7 points on real-world physical kitchen tasks.
Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this limitation, we present World-to-Wrist VLA (W2-VLA), a VLA model for fine-grained robot manipulation with task-conditioned future wrist modeling. Given current multi-view observations and a task instruction, W2-VLA contextualizes a set of latent modeling tokens as a compact interface between the vision-language model and the wrist predictor. Conditioned on this interface and the observed wrist history, the predictor forecasts future wrist latents, which are transformed into future-aware context for action prediction. In addition, we introduce W2-CoT, a synthesis pipeline that produces structured annotations describing manipulation progress, physical transition cues, and wrist-local evidence. These annotations provide auxiliary supervision that shapes the task-conditioned latent interface. Experiments on LIBERO, RoboTwin 2.0, and real-world manipulation tasks demonstrate improved fine-grained and contact-sensitive manipulation across both single-arm and bimanual settings, while maintaining action-generation rates above 80 Hz.
Yuhao Pan, Haosong Peng, Zhengshen Zhang +8
1The Hong Kong University of Science and Technology · 2National University of Singapore · 3Nanyang Technological University +4