Physical active vision allows robots to change their viewpoint when task-relevant observations become unreliable, yet existing manipulation benchmarks provide limited support for studying how policies recover from occlusion during execution. We introduce BAVO-Bench (Bimanual Active Vision under Occlusion), a bimanual active-vision benchmark that systematically controls external visibility through Clean, Stage Occlusion, and Random-time Occlusion conditions, enabling evaluation of both manipulation performance and active visual recovery. Building on this setting, we present A-FAR (Active Future-Aware Recovery), an active-vision policy for joint viewpoint and manipulation control. A-FAR represents moving-camera observations in a unified robot-centric 3D frame and distills relational structure together with its future evolution from a pretrained 4D model, providing the policy with future-aware geometric guidance without requiring future observations at deployment. Experiments across multiple manipulation tasks show that A-FAR improves robustness to both structured and temporally shifted occlusions while maintaining strong performance under clean observations.
Figures & tables
Figure 1: Overview of our motivation, benchmark, and approach. (a) We study physical active vision as closed-loop recovery from external occlusion: using a single active camera as the sole visual sensor, the robot actively changes its viewpoint to recover task-relevant visibility and continue manipulation. (b) We introduce BAVO-Bench , a bimanual active-vision manipulation benchmark that systematically introduces controlled external occlusions to study viewpoint recovery, together with a VR teleoperation pipeline for collecting active-view demonstrations. (c) We present A-FAR , an active-vision policy that canonicalizes moving-view observations in robot-centric 3D and distills D4RT’s 4D relational priors, enabling stable, future-aware control under active viewpoint changes.
Figure 2: BAVO-Bench overview. Top: Five multi-stage bimanual manipulation tasks. Bottom: Representative rollouts under the two occlusion settings. Red-bordered frames mark occlusion events, and the timelines illustrate recovery and manipulation behaviors. Both settings occlude task-relevant targets; Stage Occlusion aligns interventions with predefined task transitions, whereas Random-time Occlusion samples intervention times independently of these transitions.
Figure 3: Overview of A-FAR. Current RGB-D observations are canonicalized into a robot-base point cloud and encoded as current point tokens Ht . A state-conditioned WorldQueryFormer predicts future relational tokens H^tF , which are fused with the current representation for ManiFlow action generation. During training, a frozen D4RT teacher queries the same source points at current and future target times to construct relational targets that supervise the current structure ( Lcur ) and its temporal evolution ( Levo ). The teacher branch and future frames are removed at deployment.
Clean
Stage Occlusion
Random-time Occlusion
Method
SR (%) ↑
TPS (%) ↑
SR (%) ↑
TPS (%) ↑
SR (%) ↑
TPS (%) ↑
ACT
10.3±1.2
46.2±6.8
24.7±2.3
58.1±1.7
15.7±3.2
46.6±0.4
π0.5
13.0
55.2
42.0
73.4
4.0
40.8
ManiFlow
45.7±1.5
70.5±6.4
60.7±3.1
81.5±3.2
15.0±3.5
50.5±2.7
A-FAR
51.3±4.2
73.9±5.4
63.7±2.1
83.9±0.9
18.7±1.2
51.9±1.4
Table 1: Main results on BAVO-Bench. Results are macro-averaged over five tasks. ACT, ManiFlow, and A-FAR report mean ± standard deviation over three training seeds; π0.5 is reported from one multi-task training run. Higher is better.
Task
ACT
π0.5
ManiFlow
A-FAR
Cube
16.49/0.00
32.81/0.00
26.25/2.18
25.88/3.44
Drawer
19.71/8.87
35.65/0.00
30.47/11.67
32.27/12.91
Slide
17.64/5.29
28.72/2.87
22.50/2.99
24.08/4.02
Stack
20.36/0.00
32.94/0.00
28.66/0.00
28.93/0.49
Sweep
21.49/0.71
41.99/4.20
34.97/5.25
37.04/8.04
Avg.
19.14/2.97
34.42/1.41
28.57/4.42
29.64/5.78
Table 2: Active-view analysis under Random-time Occlusion. Each entry reports CVG / Joint . CVG measures camera-attributable visibility recovery (pp), while Joint=SR×CVG/100 jointly reflects visibility recovery and task success. Higher is better.
Figure 4: Real-world deployment of A-FAR. Representative rollout under an externally introduced occlusion. From left to right, task-relevant visibility is disrupted during execution, the active camera changes viewpoint to recover the scene, and manipulation proceeds after the target region becomes observable again. The sequence demonstrates closed-loop active-view behavior on our physical manipulation platform.
Variant
C
S
R
Avg.
ManiFlow (base)
44
60
15
41.0
+ Future Evo.
47
64
16
42.3
+ Future Fusion
47
63
18
43.0
A-FAR (base)
50
66
19
44.7
ManiFlow (camera)
26
44
14
28.0
A-FAR (camera)
22
43
16
27.0
Table 3: Ablation study. SR (%) under Clean (C), Stage (S), and Random-time (R) Occlusion.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Objective
Ordered milestones
Cube Handoff
Handoff a cube from the right arm to the left arm and place it on the target plate.
M1 : the right gripper contacts the cube; M2 : the right arm lifts the cube; M3 : both grippers hold the cube during handoff; M4 : the left gripper takes over and the right gripper releases; M5 : the left arm places the cube on the target plate.
Block Stacking
Move the base block to the central workspace and stack the top block on it.
M1 : the right gripper contacts the base block; M2 : the base block is placed in the stacking region and released; M3 : the left gripper contacts the top block; M4 : the top block is lifted into the pre-stacking region; M5 : a stable stack is formed with both grippers released.
Drawer Stowing
Open the drawer, stow an object inside, and close the drawer.
M1 : the right gripper contacts the drawer handle; M2 : the drawer is opened; M3 : the left gripper contacts the object; M4 : the object is placed and released inside the drawer; M5 : the drawer is closed with the object remaining inside.
Slide Retrieval
Slide the object to expose its handle, then retrieve it with the left arm.
M1 : the right gripper contacts the blade; M2 : the handle is exposed by sliding; M3 : the right gripper releases the object; M4 : the left gripper contacts the exposed handle; M5 : the left arm lifts the object from the support surface.
Dustpan Sweeping
Position the dustpan and sweep the target object into it.
M1 : the left gripper contacts the dustpan handle; M2 : the dustpan reaches a valid sweeping configuration; M3 : the right gripper contacts the broom; M4 : the broom contacts the target object during sweeping; M5 : the target object reaches the dustpan entrance.
Appendix
Table 4: Task objectives and milestone progression in BAVO-Bench. Each task contains five ordered milestones, with M5 defining task success.
Task
Initial target
Transitions
Subsequent targets
Cube Handoff
cube
M2,M4
cube → plate
Block Stacking
base block
M2,M4
top block → stacking region
Drawer Stowing
drawer front
M2,M3
object → drawer front
Slide Retrieval
support platform
M1,M2
blade → handle
Dustpan Sweeping
dustpan
M2,M3
broom → target object
Appendix
Table 5: Stage Occlusion schedule. An initial occlusion is introduced at reset, followed by two stage-aligned target updates.
Figure 5: Rollout visualizations under Clean. Representative sequences for the five benchmark tasks without external visibility intervention. For each task, the top row shows the third-person scene view and the bottom row shows the active-camera view. Frames are ordered from left to right.
Figure 6: Rollout visualizations under Stage Occlusion. Representative sequences for the five benchmark tasks with stage-aligned occlusion events. For each task, the top row shows the third-person scene view and the bottom row shows the active-camera view. Frames are ordered from left to right. The sequences highlight how structured occlusions are introduced at predefined task transitions.
Figure 7: Rollout visualizations under Random-time Occlusion. Representative sequences for the five benchmark tasks with temporally unexpected occlusion events. For each task, the top row shows the third-person scene view and the bottom row shows the active-camera view. Frames are ordered from left to right. The sequences illustrate how visibility disruptions may occur during ongoing manipulation behaviors.
Configuration
Value
Point-cloud size
256
Teacher query points
256
Observation horizon
2
Action horizon
16
Point feature dimension
128
State feature dimension
64
Appendix
Table 6: Main architecture and training settings for A-FAR.
Task
Method
C
S
R
Cube Handoff
ACT
0.0±0.0
6.7±7.6
0.0±0.0
π0.5
5.0
30.0
0.0
ManiFlow
41.7±2.9
68.3±5.8
8.3±7.6
A-FAR
43.3±14.4
58.3±2.9
13.3±2.9
Drawer Stowing
ACT
50.0±5.0
56.7±2.9
45.0±8.7
π0.5
25.0
75.0
0.0
Appendix
Table 7: Task-level success rate (SR, %). Results are reported under Clean (C), Stage Occlusion (S), and Random-time Occlusion (R). ACT, ManiFlow, and A-FAR are reported as mean ± std over three training seeds; π0.5 is evaluated from a single training run. Best performance for each task and condition is shown in bold.
Task
Method
C
S
R
Cube Handoff
ACT
33.0±2.6
44.0±3.6
24.0±6.2
π0.5
45.0
68.0
31.0
ManiFlow
70.3±5.0
87.7±2.3
45.3±5.7
A-FAR
69.7±13.2
84.0±1.7
47.7±6.0
Drawer Stowing
ACT
78.7±3.1
80.7±2.1
77.3±4.7
π0.5
70.0
90.0
51.0
Appendix
Table 8: Task-level task progress score (TPS, %). Results are reported under Clean (C), Stage Occlusion (S), and Random-time Occlusion (R). ACT, ManiFlow, and A-FAR are reported as mean ± std over three training seeds; π0.5 is evaluated from a single training run. Best performance for each task and condition is shown in bold.
Figure 8: Real-world Experiement Setup At the home position, the manipulation arm’s gripper points toward the workspace, and the camera is oriented in the same direction.