PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration
Organizations: National Taiwan University · NVIDIA Research · National Yang Ming Chiao Tung University
Abstract
We present PhysCaP, a Physics-Informed Code-as-Policy agent system for active perception in robotic manipulation. While vision-language-action policies excel at imitating demonstrations, they rely on passive observation and fail to infer latent physical properties critical for manipulation. PhysCaP augments code-as-policy frameworks with a physics-informed exploration layer that enables explicit information-seeking through interaction. Our method introduces training-free physical property extraction modules that estimate object mass and stiffness from robot proprioception without additional sensors. To balance exploration costs and the efficiency of information obtained, PhysCaP employs a multi-agent design: a Planner that decides when to explore and when to stop, and a Prioritizer that filters implausible interactions and ranks the remainder using a heuristic priority score, enabling efficient, targeted exploration. We evaluate PhysCaP on three real-world tabletop manipulation tasks and a simulated task in LIBERO. The results show that existing passive and naive interactive baselines either fail when physical properties are hidden or over-explore, whereas PhysCaP achieves comparable performance with fewer interactions and reduced execution time. Ablation studies further validate the effectiveness of the proposed physical property extraction modules. Project page: https://physcap.github.io
Figures & tables
| Method | Task 1: Identify Empty Can | Task 2: Pick Ripe Avocado | Task 3: Pack Grocery Bag | |||
|---|---|---|---|---|---|---|
| OI ( ) | Time ( ) | OI ( ) | Time ( ) | OI ( ) | Time ( ) | |
| CaP+PhysX | s | s | s | |||
| CaP+PhysX+Planner | s | s | s | |||
| PhysCaP-joint | s | s | s | |||
| PhysCaP (Ours) | s | |||||
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.