We present PhysCaP, a Physics-Informed Code-as-Policy agent system for active perception in robotic manipulation. While vision-language-action policies excel at imitating demonstrations, they rely on passive observation and fail to infer latent physical properties critical for manipulation. PhysCaP augments code-as-policy frameworks with a physics-informed exploration layer that enables explicit information-seeking through interaction. Our method introduces training-free physical property extraction modules that estimate object mass and stiffness from robot proprioception without additional sensors. To balance exploration costs and the efficiency of information obtained, PhysCaP employs a multi-agent design: a Planner that decides when to explore and when to stop, and a Prioritizer that filters implausible interactions and ranks the remainder using a heuristic priority score, enabling efficient, targeted exploration. We evaluate PhysCaP on three real-world tabletop manipulation tasks and a simulated task in LIBERO. The results show that existing passive and naive interactive baselines either fail when physical properties are hidden or over-explore, whereas PhysCaP achieves comparable performance with fewer interactions and reduced execution time. Ablation studies further validate the effectiveness of the proposed physical property extraction modules. Project page: https://physcap.github.io
Figures & tables
Figure 1: PhysCaP augments Code-as-Policy (CaP) agents with physics-informed exploration, enabling inference of latent object properties ( e . g ., mass) via physical property extraction modules to solve manipulation tasks requiring hidden-state estimation, e . g ., removing an empty can.
Figure 2: Overview of PhysCaP, a Physics-Informed Code-as-Policy agent. Given a task requiring latent physical information, the Planner identifies what information is missing from the visual scene and proposes an initial exploration plan. The Prioritizer then refines this plan to improve interaction efficiency by filtering implausible actions and reducing redundant exploration. Finally, a code-generation agent produces executable programs that invoke the Physical Property Extraction (PhysX) modules ( get_mass , get_stiffness ) to actively measure the properties. The measurements update the system belief, which the Planner consults to decide whether to explore further or execute the task.
Figure 3: PhysCaP measures task-relevant latent physical properties beyond passive visual perception and completes tasks. PhysCaP decides which tools ( get_mass , get_stiffness ) to invoke, then synthesizes an informed code policy grounded in those measurements (left: Identify Empty Can; right: Pick Ripe Avocado).
Figure 4: Ground-truth physical properties of the relevant target objects are annotated in each task setup, with mass (g) and/or stiffness ( s∈{1,…,5} ) shown according to properties required by each tasks.
Method
Task 1: Identify Empty Can
Task 2: Pick Ripe Avocado
Task 3: Pack Grocery Bag
OI ( ↓ )
Time ( ↓ )
OI ( ↓ )
Time ( ↓ )
OI ( ↓ )
Time ( ↓ )
CaP+PhysX
4.00
396.84±15 s
4.00
515.61±9 s
9.17
1019.32±376 s
CaP+PhysX+Planner
3.75
368.40±62 s
4.13
563.29±76 s
7.80
853.14±54 s
PhysCaP-joint
2.00
222.11±2 s
2.71
384.39±111 s
6.50
682.3±139 s
PhysCaP (Ours)
2.00
256.57±62 s
2.00
300.47±53 s
5.50
572.79±61 s
Table 1: Task performance. PhysCaP achieves the lowest/tied interaction count across all tasks and the shortest execution time on two of three tasks. OI ( ↓ ) denotes the average number of physical object interactions during exploration including repeated interactions with the same object, while Time ( ↓ ) reports the total robot execution time.
Figure 5: Real-world task performance and exploratory efficiency. (a) PhysCaP matches CaP+PhysX on Empty Can and Ripe Avocado, and improves on the longer-horizon Pack Grocery, where CaP+PhysX can lose track of prior measurements, leading to remeasurement or failure. (b) Efficiency in physical exploratory interactions and total robot execution time; bubbles are grouped by task, with area indicating standard deviation. Lower-left is more efficient.
Figure 7
Figure 7: Validating the PhysX modules against known ground truth. (a) Across five objects spanning 13–963 g, PhysX tracks ground truth (GT) over the whole range, while VLM estimates—given the reference image (I), raw torque (T), and object name (O) in combination—are misled by appearance, e . g ., substantially overestimating the empty can. (b) Squeezing three avocados of differing ripeness: green lines are linear fits with shading for measurement variance, and dashed blue lines mark the stiffness-level boundaries set by the calibrated spring reference. The ripe avocado (Level 2) separates cleanly from both unripe ones (Levels 4–5).
Figure 8: Cross-VLM generalization. We test three VLM backbones on Identify Empty Can. Across all three, PhysCaP consistently outperforms the CaP+PhysX baseline on execution time (left) and object interactions (right), showing that our framework generalizes across VLM backbones.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Our experiment setup. The ZED 2i stereo camera is mounted directly behind the AgileX PiPER robot arm to provide a perspective that aligns with the robot’s view of the manipulation area. A supplementary Logitech C270 HD is placed to the side to capture external footage for recording purposes.
Figure 10: Visual progression of avocado ripening. As the fruit matures, its skin transitions from bright green to black. This darkening serves as a primary visual cue to assess whether the avocado is ripe enough to consume.
Figure 11: Find Blue Cube task. The cube is hidden under one of three cups, one of which is too small to hold it. All methods reach comparable success.
Figure 12: Qualitative results for the Identify Empty Can task. CaP drastically under-explores by blindly guessing which can is empty, as it inherently lacks the ability to infer hidden physical properties. The CaP+PhysX baseline over-explores by lifting and weighing every can in the scene. In contrast, our PhysCaP method efficiently weighs only the cans that present visual cues.
Figure 13: Qualitative results for the Pick Ripe Avocado task. The naive CaP baseline cannot detect hidden physical states and instead resorts to blind guessing, placing a random avocado on the tray. CaP+PhysX over-explores by squeezing every avocado before identifying the ripe one. In contrast, our PhysCaP method efficiently squeezes only the two visually dark avocados, finding the ripe one in the fewest steps.
Figure 14: Qualitative results for the Pack Grocery Bag task. CaP cannot reliably select a suitable foundation item without access to hidden physical properties, while CaP+PhysX over-explores by measuring mass and stiffness across the candidate cubes. In contrast, PhysCaP uses accumulated mass measurements to identify the heavier candidates and prioritizes stiffness measurements on this reduced set, allowing it to select the stiffer heavy candidate with fewer unnecessary interactions.
Figure 15: Qualitative results for the Find Blue Cube task. The CaP and CaP+PhysX methods exhaustively explore the scene by lifting all cups, whereas our PhysCaP method efficiently lifts only the visually hinted cups.
Figure 16: A quad-view rendering of the custom 3D cup asset utilized in the simulated experiments. The completely opaque design is engineered to reliably conceal target objects, necessitating the interactive exploration behaviors evaluated in our benchmarking tasks.
Figure 17: Qualitative results for the Identify Empty Can task in the LIBERO environment. The results demonstrate that the VLA baseline guesses blindly without physical feedback, while CaP+PhysX+Planner exhaustively over-explores by weighing every container. In contrast, our PhysCaP method efficiently targets and interactively weighs only the most probable candidates, completing the objective with superior accuracy and speed.
Figure 18: The object assets used in the Identify Empty Can task. The assets include two full cans, one half-full can, and the target empty can. This intentional variation in mass requires the robotic agent to interactively weigh the visually hinted containers to successfully identify the target.
Figure 19: The object assets used in the Pick Ripe Avocado task. The assets include two green avocados that are unripe and stiff, one dark avocado that’s not ripe enough and still a bit stiff, and another dark avocado that’s both ripe and soft. This intentional variation in stiffness requires the robotic agent to interactively test the stiffness of the visually hinted fruits to successfully identify the target.
Figure 20: The object assets used in the Pack Grocery Bag task. The assets include four colored boxes with varying levels of stiffness and mass. These variations require the robotic agent to interactively weigh and squeeze the blocks to identify the target successfully.
Figure 21: The object assets used in the Find Blue Cube task. The assets include the target blue cubes, large coffee cups designed to conceal the target, and a smaller cup acting as a visual distractor. This specific arrangement forces the robotic agent to employ interactive perception to successfully locate the hidden item among the containers.