Fisheye-VLA: Decoupling Coverage and Acuity for Manipulation with a Single Fisheye Camera
Authors: Ziang Ren, Zike Yan, Raymond Zhang, Xuguo He, Zhongyu Li
Organizations: Hong Kong Embodied AI Lab · The Chinese University of Hong Kong, Hong Kong SAR, China · DeepCybo · University of Washington, Seattle, WA, USA
Manipulation requires both broad scene awareness and detailed local feedback, yet conventional camera rigs provide them through separate front and wrist cameras. We present Fisheye-VLA, a visual interface that brings these capabilities together using a single passive fisheye. A global view preserves the workspace, while local perspective crops direct detail toward the interaction. The key design question is where this local visual budget should go. We answer it through a controlled re-rendering study, comparing alternative crop directions on the same recorded observations. The study finds that end-effector-centered views capture most of the estimated benefit of a much larger candidate pool, motivating a compact allocation around both hands. Our interface uses calibrated end-effector projection and motion lead to track the crops, while a shared ray encoding preserves their spatial meaning as they move. Integrated with a pretrained VLA, it achieves 84% and 82% success in the two expanded tabletop regions, where some target placements extend beyond the front-camera coverage, and supports shelf and conveyor manipulation. Ablations show that local crops and their viewing directions become more important in the larger workspace regions. The results demonstrate that a single fisheye can support these manipulation tasks without physical wrist cameras.
Figures & tables
Figure 1: Fisheye-VLA uses one fixed camera (boxed) to supply a global view and two interaction-centered crops, each 224×224 . White outlines mark the 90∘ front-camera field of view. Rendering local detail from the wider fisheye image supports manipulation without wrist cameras.
Interface
Supply of local detail
Allocation requirement
View cost
View rule / wrist dependence
Wrist + front Jangir et al. [2022]
Physical wrist images
Camera mounting
Separate capture streams
Rig-specific wrist views
EyeRobot Kerr et al. [2025]
Foveated mechanical eye
Learned BC–RL gaze
Camera actuation
Eye-centered observations
GIAVA Chuang et al. [2026]
Foveated active-camera images
Human gaze + learned gaze
Actuation + tokenization
Gaze-conditioned image grid
SaPaVe Liu et al. [2026]
Actively redirected camera
Semantic camera learning
Camera actuation
Camera-conditioned policy
WristWorld Qian et al. [2025]
Generated wrist-view video
Geometry + video model
Reconstruction + generation
Synthesized wrist viewpoints
EgoMimic Kareer et al. [2024]
Head and robot wrist views
Human/robot data alignment
Separate capture streams
Shared head view; robot wrists
Table 1: Observation interfaces and view-construction requirements.
Figure 2: Observation construction and policy integration. (a) Current end-effector (EE) poses and a causal velocity estimate predict the viewing centers. Calibrated projection locates these centers in the captured fisheye image; inverse perspective sampling extracts and rectifies each local view in one step. A resized global view retains the scene. The three views enter the vision encoder, and their source rays supply a common geometric reference before action prediction. The standalone sphere illustrates the shared optical center; the grid inset uses an equidistant fisheye projection. (b) Candidate-view footprints for the offline placement instrument, selected from a shared anchor bank. These are measurement candidates, not the deployed policy inputs; online crops follow the EE poses in (a). (c) Larger views of handheld collection and the matched grippers show how tracked human poses supply the same rendering input. This interface uses pose-dependent local views while preserving global coverage.
Contrast
Round
Δ (mm)
95% CI
EE − random, one each
1
0.890
[0.520, 1.297]
Two EE − one EE
2
0.175
[0.007, 0.346]
All candidates − one EE
2
0.298
[ − 0.034, 0.626]
Object − random, one
2
− 0.148
[ − 0.386, 0.126]
Object − random, three
2
0.030
[ − 0.298, 0.360]
Table 2: Offline placement: paired gains with 95% CI.
Figure 4: Recorded execution sequences for the four task families. Columns show ordered stages from approach through transfer to completion; the same global/local rendering rule and multi-task policy serve every row. Shelf interaction requires keeping the elevated destination in context, and conveyor transfer requires following a moving object. The sequences illustrate reuse of the interface across different spatial and timing demands. Frame spacing indicates order, not equal elapsed time.
Configuration
R1
R2
R3
Avg.
Camera baselines
Global fisheye (w/o crops)
68
36
30
44.7
Front camera + two wrists
92
58
14
54.7
Front camera
44
38
4
28.7
Two wrist cameras
76
36
0
37.3
Fisheye + two wrists
70
74
78
74.0
Table 3: Task A: regional success and mean (%).
Task
Success (%)
A
Instructed fruit into box
79
B
Duck toy onto shelf
88
C
Mug from shelf into box
86
D
Conveyor → bin
72
D ′
Bin → conveyor
68
Table 4: Four task families; Task D is evaluated in both directions.
Current Vision-Language-Action (VLA) models often struggle with high-precision robotic manipulation. We attribute this limitation primarily to their visual attention being dispersed across task-irrelevant regions. To address this issue, we propose ActGaze, a training approach that guides VLA policies to gaze on task-relevant regions, much like humans gaze on critical visual cues while executing precise movements. Unlike prior methods that rely on external labels for gaze supervision, ActGaze derives spatial supervision directly from the VLA's own action objective by using counterfactual visual interventions to identify regions that are critical for action prediction. Extensive real-robot experiments on four high-precision robotic manipulation tasks demonstrate that ActGaze induces more focused visual attention on task-relevant regions and consistently outperforms the base VLA policy and other visual-grounding approaches.
Real-world robot deployment rarely maintains the training-stage camera setup, where cameras often experience repositioning or remounting depending on actual scenarios. Existing view-robust Vision-Language-Action (VLA) policies tolerate such camera variations only when the camera extrinsics are explicitly provided, making them fragile and hard to use especially when view robustness is critical. We argue that the policy should not be told where the camera is, but rather figure it out by itself. To this end, we introduce Camera-Centric VLA (CamVLA), a new VLA model that decouples manipulation controls from camera geometry by predicting (i) a camera-centric end-effector action expressed in the local camera frame, and (ii) a 6-DoF hand-eye matrix relating cameras to the robot base. A deterministic geometric transformation composes the two predictions into a robot base-frame action. This disentangles how I should move in pose-independent camera-centric action generation from where I am looking from in camera-perspective geometric grounding. The resulting policy is calibration-free, depth-free, and single-view, requiring only a single monocular RGB image as the visual observation and task instruction at deployment. Evaluations in both simulation and real-world robot data show that CamVLA consistently improves success rates across diverse unseen viewpoints. Project page: https://alibaba-damo-academy.github.io/CamVLA/.
We propose OC-VLA++, an extension of OC-VLA for viewpoint generalization under limited camera coverage. While OC-VLA grounds robot actions in the camera coordinate system to align action supervision with visual observations, camera-space grounding alone can still overfit to the few viewpoints observed during training. OC-VLA++ addresses this limitation by introducing geometry-guided paired-view supervision and an explicit cross-view action-equivariance objective. Given paired observations of the same manipulation scene from geometrically related viewpoints, the model is trained such that their camera-space predictions correspond to the same robot-frame action. This objective explicitly supervises how action predictions should transform across viewpoints, rather than relying solely on image-level augmentation. Experiments demonstrate substantial improvements in unseen-view generalization under limited camera coverage, with performance degrading more gracefully under increasing camera displacement. These results establish cross-view action equivariance as an effective complement to observation-centric action grounding for robust real-world deployment.