Many indoor environments in which robots operate, such as warehouses, offices, and hospitals, already have cameras installed. They observe parts of the building that the robot cannot see from where it stands, yet navigation policies, including recent vision-language-action (VLA) models, do not use them. We propose InfraVLA, an end-to-end method that adapts a pretrained navigation VLA to such static infrastructure views: a closed-circuit television (CCTV) encoder turns each external view into tokens of the input sequence. Because the views matter only at rare decision points, fine-tuning alone did not make the policy use them in our experiments; we therefore train in two stages, on demonstrations with upsampled counterfactual data and then on recovery data. We evaluate on two simulated warehouse tasks, finding an object named in the instruction and rerouting around blocked aisles, where the deciding information is often visible only to the infrastructure cameras. Tested in distribution, InfraVLA reached a success rate of 100% on both, against 34.0% and 73.6% for a baseline without CCTV input. On out-of-distribution test sets it reached 88.2% and 88.9%. On a real quadruped fine-tuned with under 10 minutes of demonstrations, the policy reached 83.3% against 29.2% for the on-board-only baseline.
Figures & tables
Fig. 2: Architecture. Green: our extension, the infrastructure images, encoded per view, pooled from 256 to 64 tokens, and inserted after the on-board tokens. Orange: the on-board image pipeline; blue: the task specification; both as in OmniVLA [ 2 ] .
Semantic (sim)
Spatial (sim)
Semantic (real)
CCTV cameras
4
4
2
Goal modality
language
2D pose
language
Success
1m to object
1m to goal
human-judged
Test rollouts
50
72
24
Trajectories (train / val.)
552 / 48
481 / 46
75 (all)
Training data (h)
3.03
6.89
0.16
TABLE I: Hyperparameters of the three experiments: data, training, recovery mining, snippet detection, and success criterion.
Fig. 3: Semantic task. Left: two rollouts of one scene in plan view, both instructed to reach the box (brown square) and both reaching it. From the black start dot, InfraVLA after stage 1 (gray) passes close to a shelf (dark gray) and after stage 2 (green) keeps its distance. Right: what the four ceiling cameras see in that same scene at the start; the box is top left.
Policy
SR (%) ↑
Clearance (m) ↑
Fine-tuned OmniVLA [ 2 ] (no CCTV)
34.0
0.67
InfraVLA, no upsampling
36.0
0.62
InfraVLA, upsampling (stage 1)
100.0
0.71
+ recovery (stage 2)
100.0
1.27
TABLE II: Semantic task, 50 rollouts. SR: success within 1m of the object. Clearance: mean over rollouts of the closest distance to a shelf (demonstrations: 1.47m ).
Object
Synonym 1
SR (%)
Synonym 2
SR (%)
Forklift
lift truck
91.7
vehicle
66.7
Barrel
canister
91.7
oil drum
75.0
Box
package
91.7
crate
83.3
Traffic cone
safety cone
100.0
witch’s hat
25.0
TABLE III: Semantic task, unseen instruction words: SR over 12 rollouts per word, for two synonyms of each training noun.
Evaluation set
Rollouts
SR (%) ↑
In-distribution
50
100.0
Held-out scene, other target
103
99.0
Held-out scene, forklift
34
88.2
TABLE IV: Semantic task, out-of-distribution tests. A separate policy trained through stage 2 with every scene that puts the forklift in the top-left aisle held out.
Fig. 4: Spatial task. Left: one test scenario in plan view. From the black start dot both policies drive the same route to the middle area; the on-board-only baseline (orange) then enters the aisle sealed by a forklift (yellow), while InfraVLA (green) takes the open aisle and reaches the goal (star). Right: the four ceiling views of the same scenario at the start; the forklift that seals the aisle is visible in the bottom-right view.
Policy
SR (%) ↑
Route (%) ↑
Fine-tuned OmniVLA [ 2 ] (no CCTV)
73.6
55.6
InfraVLA, no upsampling
76.4
51.4
InfraVLA, upsampling (stage 1)
77.8
44.4
+ CF snippets only, no recovery
75.0
52.8
+ recovery (stage 2), K=6
100.0
91.7
+ recovery (stage 2), K=12
100.0
93.1
TABLE V: Spatial task, 72 systematic scenarios. Route: share of rollouts through the same aisles as A*. + rows: the stage-1 policy trained 5k more steps on the data named (stage 2). CF: counterfactual.
Policy
Seen (%) ↑
Held-out (%) ↑
InfraVLA, stage 1
75.9
33.3
+ recovery (stage 2), K=30
100.0
88.9
TABLE VI: Spatial task, out-of-distribution test. A separate policy trained with two of the nine blocking patterns held out; recovery data from the seen patterns only. SR on the 54 seen and 18 held-out scenarios.
Fig. 5: The real setup, seen from the two fixed cameras at one instant of a demonstration. Left: CCTV 1 shows the traffic cone and the fire extinguisher. Right: CCTV 2 shows the other half of the area, with the robot and the plant. A logo on the robot has been masked for anonymity.
Policy
Overall
Plant
Fire ext.
Cone
Fine-tuned OmniVLA [ 2 ]
29.2
12.5
50.0
25.0
InfraVLA
83.3
100.0
62.5
87.5
TABLE VII: Real robot, 8 test arrangements with one rollout per object (24 rollouts). SR (%) ↑ of reaching the named object, judged by the operator, overall and per object.
Fig. 6: Obstacle avoidance preserved from pretraining, seen from CCTV 1. Top: instructed to reach the traffic cone, the policy walks around a blue bin. Bottom: instructed to reach the plant, it passes between two warning signs. The strip above each photograph holds four on-board frames in chronological order, each linked to the numbered disc marking roughly where it was taken; the disc positions are approximate. A logo in the background of two frames has been masked for anonymity.
Policy
SR (%) ↑
Clearance (m) ↑
InfraVLA, upsampling (stage 1)
100.0
0.71
16 CCTV tokens per view
28.0
0.77
Shared CCTV encoder
96.0
0.61
CCTV tokens first
94.0
0.65
Frozen on-board encoder
90.0
0.59
Frozen on-board encoder and LLM
26.0
0.61
TABLE VIII: Ablations of the architecture, on the semantic task with the stage-1 recipe, 50 rollouts each. First row: the stage-1 policy of Table II ; every other row changes one decision of Section III-B . SR and clearance as in Table II .