FORM: Robot Manipulation through Direct Material Law Identification
Authors: Stepan Tretiakov, Ruihan Zhao, Cheng-Hsi Hsiao, Xingjian Li, Adam Thorpe, Hassan Iqbal, Sandeep Chinchali, Ufuk Topcu, +1 more
Organizations: Department of Mechanical Engineering, University of California, Berkeley, CA 94720, USA. · Chandra Family Department of Electrical and Computer Engineering · Maseeh Department of Civil, Architectural and Environmental Engineering · Oden Institute for Computational Engineering and Sciences · Department of Aerospace Engineering and Engineering Mechanics, The University of Texas at Austin, Austin, TX 78712, USA.
When interacting with an unfamiliar deformable material, a robot lacks prior knowledge of its physical properties and how it will respond to applied forces and motion. Rapid online identification is therefore essential for reliable manipulation. We present FORM (From Observed Response to Material laws), which identifies material properties from a single robot interaction and reuses the recovered model to plan manipulation under new actions and geometries. We use weak-form momentum balance to convert observed material motion and contact forces into linear equations in the unknown material parameters. These equations are assembled using the same material point method discretization as the forward simulator, so identification reduces to linear least-squares solves whose solutions can be used directly for prediction without refitting or conversion. Across four material classes, FORM reduces identification time from roughly 10--25 minutes for iterative baselines to 2--5 seconds, while maintaining competitive accuracy on new motions, initial conditions, and geometries. We demonstrate our approach in simulation and on hardware across four manipulation tasks: elastic rod insertion, golf putting with an elastic club, elastoplastic shaping, and target-volume pouring. In each task, the model identified from a single interaction is reused to plan new motions or manipulate a different geometry. FORM estimates elastic properties within 3.4% and elastoplastic properties within 2%, achieves 72.4--77.8% IoU in dough shaping, and keeps mean pouring error at 3.8 mL across target volumes of 60--160 mL.
Figures & tables
Fig. 1: The dynamics of deformable materials and fluids are often hard to observe from vision alone. Active physical interaction is required to estimate their latent properties and enable robust manipulation policies. (Top) Starting from the same initial shape, applying the identical robot action such as a 25 mm pinch or a 10 N press results in different deformations dictated by the stiffness and yield behavior. (Bottom) Similarly, applying the same pouring action at 10 degrees per second results in distinct flow behaviors governed by the varying viscosities of the fluids.
Fig. 2: FORM: material identification and manipulation. (a) Synthetic stereo tracks and motion reconstructed from real RGB-D observations. Colors indicate image displacement in stereo and initial height in RGB-D. The force trace is from the synthetic press. (b) Divergence-free test fields remove pressure. Green arrows denote test vectors, not motion. The weak balance identifies stiffness E and yield threshold Y . The stress–strain curve is schematic. (c) The recovered model and target are used to plan the robot motion in MPM, illustrated here in the simulated shaping task.
Law
Stress response
Parameters
Elastic and hyperelastic , with p=−κ(J−1) for the τ rows
neo-Hookean
τ=J2C10devbˉ
C10,κ
Mooney–Rivlin
τ=J2dev[C10bˉ+C01(Iˉ1bˉ−bˉ2)]
C10,C01,κ
Yeoh
τ=J2[i=1∑3iCi(Iˉ1−3)i−1]devbˉ
C1,C2,C3,κ
Gent
τ=J[1−(Iˉ1−3)/Jm]μdevbˉ
μ,Jm,κ
fixed-corotated
σ=J2μ(F−R)FT+λ(J−1)I
μ,λ
TABLE I: Constitutive models and parameters, with σ=−pI+τ . Stresses are Cauchy unless marked K for Kirchhoff.
Jell-O
Sand
Plasticine
Water
Method
Mean Loss
Solve Time (s)
Mean Loss
Solve Time (s)
Mean Loss
Solve Time (s)
Mean Loss
Solve Time (s)
NCLaw
5.2e-4
1513.2
1.6e-4
1305.6
1.2e-4
1300.2
5.7e-4
1306.8
Diff. Sys-Id
1.2e-8
715.8
6.3e-13
1132.8
1.6e-10
1028.4
4.0e-9
556.2
Ours (Fn Enc.)
3.5e-4
5.2
3.3e-8
0.1
2.9e-6
5.3
2.3e-6
4.8
Ours (Full)
5.1e-6
4.4
1.4e-8
2.4
5.2e-7
4.7
2.3e-6
2.2
Ours (No stress)
5.1e-6
3.5
2.0e-8
6.0
5.2e-7
6.5
2.3e-6
1.6
TABLE II: Reconstruction results for FORM and baselines. Variants use learned bases selected by material type ( Ours (Fn Enc.) ), full information ( Ours (Full) ), observations without stress ( Ours (No stress) ), or position-only observations ( Ours (Pos. only) ). Bold marks a lowest-loss non-oracle result and its solve time.
Fig. 3: Stiffness from a common bending probe transfers to rod insertion (top) and elastic-club putting (bottom). Blue and orange denote materials A and B. The pale outline marks the initial rod pose. Putting shows B and A at the end of the backswing, then A at contact and follow-through. Arcs measure projected grip-to-head inclination relative to horizontal; white balls and timestamps show recorded motion states. Colored balls overlay final positions from four separate executions, with color denoting club material. The circle marks the stopping target.
Fig. 4: Material identification and X shaping in simulation. Left: A’s identification press and final planned pinch. The arrow indicates plate motion; translucent cylinders show the simulated fingers, with the hand illustrating their pose. Right: target shape and matched and swapped results for A (blue) and B (orange), all at a common scale. Results are shown 1 s after finger withdrawal; dashed outlines mark the target.
Fig. 5: Hardware pressing and shaping of red Play-Doh, yellow butter slime, and gray plasticine (left to right). Top: recorded pressing. Middle: force-driven MPM predictions at matching times, using models identified from separate presses. Simulation views are horizontally mirrored for display. Bottom: reconstructed shapes after four open-loop pinches, rigidly aligned to the dashed target outlines without scaling.
Fig. 6: Glycerol pouring with an identified model. (a) Hardware and MPM at matching times during the 60∘ identification pour; this recording supplies both viscosity identification and effective contact calibration. Simulated liquid is colored blue for visibility. (b) Glycerol: frozen MPM predictions and hardware mean ± sample standard deviation over five trials per target. Water: one hardware trial per target, without error bars. (c) Repeat 2: target → measured receiver volume (mL), with readings rounded to the nearest mL.
Accurate physical parameter identification of manipulated objects is fundamental to advanced robotic manipulation and the construction of faithful digital twins. However, acquiring physically consistent inertial and frictional properties from real-world interactions remains challenging due to sensing noise, modeling errors, and limited prior knowledge. This paper presents RigPI, a systematic framework for identifying dynamic parameters of both unconstrained rigid bodies and multi-link rigid bodies during robot-object interaction. RigPI integrates vision-based semantic priors, force-torque measurements, and motion observations within a differentiable simulation pipeline. A vision-language model (VLM) provides informed initialization and a constrained search space, while gradient information from a differentiable physics simulator enables efficient and stable parameter refinement. The proposed two-stage optimization strategy alleviates sensitivity to noise and avoids physically implausible solutions. Extensive real-world experiments on objects with revolute and prismatic joints demonstrate that RigPI achieves accurate and stable parameter estimates, and successfully reproduces manipulation trajectories on a real robot with parameter-aware predictive validity. These results highlight the effectiveness and robustness of RigPI for real-world robotic system identification tasks.
Predicting how deformable objects evolve under robotic manipulation is a longstanding challenge. Existing approaches typically rely on per-object optimization to fit material parameters, which can be slow and cannot generalize, while end-to-end learned alternatives extrapolate poorly and often violate basic physical structure. We present PhysCoRe, a physics-corrected residual world model that couples a differentiable Material Point Method (MPM) simulator with two feed-forward neural networks. A material refinement module, Material from Motion (MfM), infers per-particle elasticity from visual observations, grounding the simulator in object-specific physics. A residual correction module, Residual from Dynamics (RfD), learns the discrepancy and predicts corrections to the simulator's internal dynamics, absorbing systematic biases that the analytical model cannot capture. This design also supports online material identification on novel objects. MfM adapts from limited interactions, and its predictive uncertainty steers further exploration toward the regions where its estimate is least confident. Experiments on real deformable-object manipulation sequences show that PhysCoRe outperforms state-of-the-art baselines in prediction accuracy, and that its predicted confidence forms a reliable distribution across the object's geometry, providing a natural signal for future confidence-guided exploration.
Shape control of deformable linear objects (DLOs) is challenging for imitation learning because deformation behavior varies with material properties such as stiffness and elasticity, so a single policy must generate different action sequences for different objects even when the goal shape is identical. We propose a diffusion policy conditioned on material labels that are estimated online during manipulation. A recurrent estimation network predicts the material label of the grasped object from the time series of multi-view images and robot joint states, and the predicted label conditions the diffusion policy at every inference step. We collected 480 real-robot demonstrations covering four DLO materials and three groove-placement tasks, and compared per-material specialist policies, a task-conditioned policy without material labels, a policy conditioned on ground-truth material labels, and the proposed policy. Conditioning on ground-truth material labels improved the average success rate from 45.8% to 60.0% over the task-only policy, and the proposed policy reached 60.8% without any prior material information, matching the policy given ground-truth labels. A post-hoc analysis shows that the estimator extracts material-related information from the manipulation observations and that the diffusion policy responds to the resulting conditioning signal, while the one pronounced failure case is associated with persistent confusion between two similar materials.
Graduate School of Natural Science and Technology, Kanazawa University, Kanazawa 920-1192, Japan · National Institute of Advanced Industrial Science and Technology (AIST), Tokyo 135-0064, Japan · Faculty of Frontier Engineering, Institute of Science and Engineering, Kanazawa University, Kanazawa 920-1192, Japan