Fisheye-VLA: Decoupling Coverage and Acuity for Manipulation with a Single Fisheye Camera
Authors: Ziang Ren, Zike Yan, Raymond Zhang, Xuguo He, Zhongyu Li
Organizations: Hong Kong Embodied AI Lab · The Chinese University of Hong Kong, Hong Kong SAR, China · DeepCybo · University of Washington, Seattle, WA, USA
Manipulation requires both broad scene awareness and detailed local feedback, yet conventional camera rigs provide them through separate front and wrist cameras. We present Fisheye-VLA, a visual interface that brings these capabilities together using a single passive fisheye. A global view preserves the workspace, while local perspective crops direct detail toward the interaction. The key design question is where this local visual budget should go. We answer it through a controlled re-rendering study, comparing alternative crop directions on the same recorded observations. The study finds that end-effector-centered views capture most of the estimated benefit of a much larger candidate pool, motivating a compact allocation around both hands. Our interface uses calibrated end-effector projection and motion lead to track the crops, while a shared ray encoding preserves their spatial meaning as they move. Integrated with a pretrained VLA, it achieves 84% and 82% success in the two expanded tabletop regions, where some target placements extend beyond the front-camera coverage, and supports shelf and conveyor manipulation. Ablations show that local crops and their viewing directions become more important in the larger workspace regions. The results demonstrate that a single fisheye can support these manipulation tasks without physical wrist cameras.
Figures & tables
Figure 1: Fisheye-VLA uses one fixed camera (boxed) to supply a global view and two interaction-centered crops, each 224×224 . White outlines mark the 90∘ front-camera field of view. Rendering local detail from the wider fisheye image supports manipulation without wrist cameras.
Interface
Supply of local detail
Allocation requirement
View cost
View rule / wrist dependence
Wrist + front Jangir et al. [2022]
Physical wrist images
Camera mounting
Separate capture streams
Rig-specific wrist views
EyeRobot Kerr et al. [2025]
Foveated mechanical eye
Learned BC–RL gaze
Camera actuation
Eye-centered observations
GIAVA Chuang et al. [2026]
Foveated active-camera images
Human gaze + learned gaze
Actuation + tokenization
Gaze-conditioned image grid
SaPaVe Liu et al. [2026]
Actively redirected camera
Semantic camera learning
Camera actuation
Camera-conditioned policy
WristWorld Qian et al. [2025]
Generated wrist-view video
Geometry + video model
Reconstruction + generation
Synthesized wrist viewpoints
EgoMimic Kareer et al. [2024]
Head and robot wrist views
Human/robot data alignment
Separate capture streams
Shared head view; robot wrists
Table 1: Observation interfaces and view-construction requirements.
Figure 2: Observation construction and policy integration. (a) Current end-effector (EE) poses and a causal velocity estimate predict the viewing centers. Calibrated projection locates these centers in the captured fisheye image; inverse perspective sampling extracts and rectifies each local view in one step. A resized global view retains the scene. The three views enter the vision encoder, and their source rays supply a common geometric reference before action prediction. The standalone sphere illustrates the shared optical center; the grid inset uses an equidistant fisheye projection. (b) Candidate-view footprints for the offline placement instrument, selected from a shared anchor bank. These are measurement candidates, not the deployed policy inputs; online crops follow the EE poses in (a). (c) Larger views of handheld collection and the matched grippers show how tracked human poses supply the same rendering input. This interface uses pose-dependent local views while preserving global coverage.
Contrast
Round
Δ (mm)
95% CI
EE − random, one each
1
0.890
[0.520, 1.297]
Two EE − one EE
2
0.175
[0.007, 0.346]
All candidates − one EE
2
0.298
[ − 0.034, 0.626]
Object − random, one
2
− 0.148
[ − 0.386, 0.126]
Object − random, three
2
0.030
[ − 0.298, 0.360]
Table 2: Offline placement: paired gains with 95% CI.
Figure 4: Recorded execution sequences for the four task families. Columns show ordered stages from approach through transfer to completion; the same global/local rendering rule and multi-task policy serve every row. Shelf interaction requires keeping the elevated destination in context, and conveyor transfer requires following a moving object. The sequences illustrate reuse of the interface across different spatial and timing demands. Frame spacing indicates order, not equal elapsed time.
Configuration
R1
R2
R3
Avg.
Camera baselines
Global fisheye (w/o crops)
68
36
30
44.7
Front camera + two wrists
92
58
14
54.7
Front camera
44
38
4
28.7
Two wrist cameras
76
36
0
37.3
Fisheye + two wrists
70
74
78
74.0
Table 3: Task A: regional success and mean (%).
Task
Success (%)
A
Instructed fruit into box
79
B
Duck toy onto shelf
88
C
Mug from shelf into box
86
D
Conveyor → bin
72
D ′
Bin → conveyor
68
Table 4: Four task families; Task D is evaluated in both directions.