cs.ROSep 29, 2026

RawVLA: Embodied Neural Image Signal Processor For Robotic Manipulation

Authors: Shuhong Liu, Heng Zhou, Lingfeng Qian, Yuhao Fang, Xianbao Hou, Qianyu Zhou, Lin Gu, Wei Sui, +2 more

Organizations: UTokyo · D-Robotics · TohokuU · NTU · HKUSTGZ

Abstract

Vision-language-action (VLA) models typically operate on RGB images produced by a fixed camera image signal processor (ISP), leaving the imaging pipeline outside the learning and evaluation loop. We systematically examine the consequences of this overlooked design choice across five fundamental ISP dimensions: gain, sensor noise, chromatic response, tonal response, and bit depth. Our analysis reveals that RAW-to-RGB processing materially shapes both action prediction and manipulation success, with different ISP dimensions exerting substantially different effects. Guided by these findings, we introduce RawVLA, a streaming neural ISP that adaptively renders RAW observations for frozen VLA policies while concentrating its capacity on the imaging factors relevant to embodied behavior. We further present RawVLA-Bench, a RAW-domain manipulation benchmark to expose image processing as an explicit evaluation variable across clean and adverse acquisition conditions. Experiments on RawVLA-Bench show that RawVLA preserves performance under standard conditions while substantially improving robustness under degraded imaging, establishing adaptive RAW processing as an effective interface between physical cameras and embodied policies.

Figures & tables

Appendix figures & tables25 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

    Jul 29, 2026Hengyi Xie, Chenfei Yao, Xianjin Wu +4Diffusion-Based Vision-Language-ActionsNvidia

  2. E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes

    Apr 6, 2026Jiajun Zhai, Hao Shi, Shangwei Guo +2Diffusion-Based Vision-Language-ActionsGeneralist Vision--Language--Action

  3. G3^3VLA: Geometric inductive bias for Vision-Language-Action Models

    Jun 23, 2026Yue Peng, Yongzhe Zhao, Artur Habuda +5Diffusion-Based Vision-Language-ActionsVisual Geometry Grounded Transformer