Pretrained robotic policies can suffer substantial performance degradation under out-of-distribution (OOD) conditions encountered during deployment, motivating post-training through real-world interaction. However, reinforcement-learning (RL)-based post-training typically requires substantial environment interactions, a burden that is especially significant in manipulation, where each trial can be slow, costly, or destructive. Therefore, we present PEARS, a physics-prior-guided hybrid RL framework for sample-efficient online adaptation of pretrained policies with tactile feedback. After each episode, its physics-guided force reasoning (PFR) module uses physical priors encoded in a vision-language model (VLM) to diagnose failures from the visual outcome and tactile interaction history and update task-appropriate contact-force bounds. A high-frequency hybrid force-position controller then enforces these bounds during contact. Complementarily, tactile-conditioned diffusion steering reinforcement learning adjusts the latent noise of the frozen flow-matching policy to correct errors in free-space motion and contact timing without updating the base model. In simulation, PEARS improves success rates by 12.4-37.4 percentage points over the strongest per-task baselines. PEARS also reduces the number of interaction episodes required for a certain success threshold by up to 53.2% relative to the fastest baseline. In real-world experiments, PEARS achieves success rates of 95% on Whiteboard Erasing and 90% on Pipette Liquid Aspiration. These results show that combining the PFR module with policy steering can accelerate adaptation while reducing costly interactions. The project website is available at https://song-kun.github.io/pears.
Figures & tables
Fig. 2: Stage-aware force control during whiteboard wiping.
Algorithm 1 PEARS
Method
Thin Sheet Transfer
Bottle Cap Twisting
Fragile Fruit Picking
SR (%) Mean ± Std. ↑
nAUC (%) ↑
E80↓
Force failure (%) ↓
SR (%) Mean ± Std. ↑
nAUC (%) ↑
E80↓
Force failure (%) ↓
SR (%) Mean ± Std. ↑
nAUC (%) ↑
E80↓
Force failure (%) ↓
Base Policy
73.7±0.6
73.7
/
5.0
43.0±1.0
43.0
/
46.0
23.7±1.2
23.7
/
52.3
Residual RL [ 6 ]
75.7±0.6
76.0
12
6.7
68.0±20.3
68.0
47
20.3
37.3±2.1
23.2
/
44.7
AWR [ 33 ]
67.7±2.5
68.2
37
7.7
21.0±13.1
20.3
/
61.7
15.7±1.2
16.4
/
55.7
DPPO [ 34 ]
4.7±3.5
2.5
/
5.7
32.7±1.5
31.7
/
55.3
22.3±2.3
16.3
/
48.0
DSRL [ 7 ]
82.3±4.5
72.2
29
9.7
57.7±1.5
57.9
/
31.3
40.3±4.2
25.3
/
36.3
TABLE I: Online adaptation and final-policy performance comparison. “/” denotes an unreached E80 .
Fig. 3: Online adaptation success across three simulation tasks. Curves show mean trailing-10 success rates across three 100-episode runs, using complete windows ending at episodes 10–100. Shaded bands denote mean ± one sample standard deviation (SD)
Fig. 4: Example of force and torque intervals predicted by PFR during online interaction. a) Thin Sheet Transfer. b) Bottle Cap Twisting.
SR (%)
DSRL only
w/o DSRL
Fixed Range
Rand Search
Qwen3
PEARS
Sheet
82.3±4.5
77.0±1.0
84.7±3.8
88.3±6.0
93.0±1.7
94.7±1.5
Cap
57.7±1.5
84.7±1.5
62.3±2.5
77.0±5.6
86.7±0.6
87.7±1.5
Fruit
40.3±4.2
61.7±2.1
44.0±2.6
40.3±2.5
67.0±2.6
77.7±1.2
TABLE II: Ablation of PFR, DSRL, force-range selection, and the VLM backbone used in PFR across three simulation tasks.
Task
Base Policy
DSRL [ 7 ]
PEARS
Whiteboard Erasing
11 / 20
15 / 20
19 / 20
Pipette Liquid Aspiration
8 / 20
12 / 20
18 / 20
TABLE III: Real-world performance (successful trials / total trials).
School of Computing and Data Science, The University of Hong Kong, Hong Kong SAR. · Shanghai Qizhi Institute, Shanghai, China. · Shanghai Jiao Tong University, Shanghai, China. +2