Diffusion and VLA policies for manipulation are often deployed through downstream impedance controllers. The stiffness and damping gains of these controllers affect task success, yet are commonly inherited from data collection rather than selected for the deployed policy. Although empirical gain sweeps can improve performance, they require repeated evaluation rollouts. We introduce PHRetune, an offline method that derives controller gains for a frozen policy without evaluation rollouts or gain search. Our approach learns a port-Hamiltonian model from demonstrations to estimate the effort and energy associated with the policy's predicted actions. The policy is applied to recorded demonstration observations, and its predictions are assessed against demonstration-derived effort and energy budgets. From this comparison, we derive a single gain scale in closed form, adjusting the downstream controller while preserving the policy and its action representation. The gains are fixed before evaluation, without requiring a prior manipulator model, task rewards, policy retraining, or additional runtime computation. Across LIBERO suites, PHRetune improves Diffusion Policy success by up to 9.4 percentage points, with the derived gains achieving the highest observed success rates in empirical gain sweeps. On all four real-world manipulation tasks, PHRetuned Diffusion Policy outperforms the nominal policy, alternative gain-tuning methods, and a policy-retraining baseline. The same procedure improves success with SmolVLA and OpenVLA-OFT on every task, while reducing acceleration and jerk for both VLA backbones.
Figures & tables
Backbone
Model
Goal
Obj.
Spa.
Diffusion Policy
Inv- Σ
0.229
0.232
0.244
DiffTune
1.123
1.140
1.131
PHRetune
0.699
0.461
0.531
Table 1 : Simulation tuned gains. We report tuned β scaling for PHRetune, Inv- Σ , and DiffTune. Robosuite defaults kp=150 with kd coupled via a fixed 1 damping ratio. The equivalent β values for Inv- Σ and DiffTune are calculated after their respective derivation.
Demos
Model
Goal
Obj.
Spa.
Total
20
Diffusion Policy
84.6
76.7
74.3
78.5
Inv- Σ
77.1
77.5
81.7
78.8
DiffTune
80.9
71.5
67.9
73.4
MtG-SM2M
86.1
66.7
80.9
77.9
PHRetune
89.0
86.4
88.3
87.9
50
Diffusion Policy
92.2
86.1
88.7
89.0
Table 2 : Diffusion Policy success rate. We report the single-task success rate on LIBERO-goal, -object, and -spatial subsets trained on 20 and 50 demos. PHRetuned Diffusion Policy achieves the highest success rate compared to baseline methods. Notably, PHRetuned Diffusion Policy trained on 20 demos approaches the performance of vanilla Diffusion Policy trained on the full 50.
Figure 1 : LIBERO gain sweep. We report the success rate over three seeds for the (a) goal, (b) object, and (c) spatial partition of LIBERO by uniformly sweeping the gain scale. The success rate for PHRetune and baselines are highlighted. MtG is omitted from the curve and β values for Inv- Σ and DiffTune are calculated post-hoc using derived gain values.
Mode
Goal
Obj.
Spa.
Total
Close Loop
β∗
0.646
0.460
0.546
88.0
SR
89.7
86.3
88.1
Shadow Mode
β∗
0.699
0.461
0.531
87.9
SR
89.0
86.4
88.3
Table 3 : Gain derivation in shadow mode and in closed loop. Energy-based gain scale β∗ and success rate of the 20-demonstration policies per LIBERO suite. Shadow mode scores the frozen policy’s predicted chunks on the demonstration frames alone. Close Loop scores the chunks it executes over simulator rollouts. Both routes recover nearly the same gains and overall success rate.
Figure 2 : Real-world manipulation tasks. (a) Placing a hat on a stand, (b) transferring salt-and-pepper mills into a tray, (c) placing a camera in its fitted holder, and (d) grasping an eraser and wiping a whiteboard.
Task
Task Prompt
Hat
Place the hat on the stand
S&P
Put the salt and pepper shakers in the box
Camera
Place the camera into the slot in the box
Eraser
Pick up the eraser and erase the whiteboard
Table 4 : Real-world task prompts. We report the task prompt for all four tasks. Real-world diffusion policy receives no language conditioning; SmolVLA receives language conditioning for single-task policies; and OpenVLA-OFT receives language conditioning for all four tasks.
Backbone
Model
Hat
S&P
Camera
Eraser
Diffusion Policy
Inv- Σ
0.217
0.178
0.163
0.166
DiffTune
1.109
1.068
1.098
1.104
PHRetune
0.696
0.461
0.707
0.859
SmolVLA
PHRetune
0.667
0.460
0.513
0.630
OpenVLA-OFT
PHRetune
0.731
0.492
0.476
0.570
Table 5 : Real-world tuned gains. We report the derived β∗ for Diffusion Policy, SmolVLA, and OpenVLA-OFT.
Model
Hat
S&P
Camera
Eraser
Overall
Tele-Op
1.00
0.87
0.87
0.91
0.91
Diffusion Policy
0.6
0.4
0.5
0.5
0.50
Inv- Σ
0.2
0.1
0.0
0.0
0.08
DiffTune
0.5
0.3
0.5
0.2
0.38
MtG-SM2M
0.6
0.4
0.4
0.4
0.45
PHRetune
0.8
0.6
0.7
0.8
0.73
Table 6 : Real-world manipulation success rate with Diffusion Policy. We report per-task success rate and their mean across the four tasks. PHRetuned Diffusion Policy outperforms the nominal policy and baseline methods on every task. Tele-Op provides a human teleoperation reference.
Model
Hat
S&P
Camera
Eraser
Overall
SmolVLA
0.7
0.5
0.4
0.6
0.58
DiffTune
0.7
0.3
0.3
0.5
0.45
PHRetune
0.9
0.7
0.7
0.9
0.80
Table 7 : Real-world manipulation success rate with SmolVLA. We report per-task success rate across the four tasks. PHRetuned SmolVLA improves success over the nominal policy and DiffTune on every task.
Metric
SmolVLA
PHRetune
Tele-Op
EE Jerk RMS
22.4±4.8
11.7±4.3(↓47.7%)
9.3±1.2
EE Jerk Peak
77.2±17.4
39.3±20.1(↓49.1%)
31.5±9.0
EE Accel. RMS
0.93±0.23
0.52±0.17(↓44.1%)
0.43±0.07
EE Accel. Peak
3.09±0.58
1.65±0.58(↓46.6%)
1.39±0.48
Joint Accel. RMS
7.91±3.77
3.34±1.74(↓57.8%)
2.30±0.34
Time to Grasp
3.66±0.68
3.66±0.15
2.99±0.48
Table 8 : Real-world motion dynamics with VLAs. We report end-effector and joint motion metrics, time to grasp, controller gain, and success rate for nominal and PHRetuned SmolVLA, PHRetuned OpenVLA-OFT over ten inference rollouts of the hat task. PHRetune reduces acceleration and jerk toward teleoperation levels with little increase to the time to grasp. Failures do not contribute to the reported metrics. End effector (EE) jerk is measured in m s -3 , EE accel. in m s -2 , joint accel. in rad s -2 , and time to grasp in seconds.
Model
Hat
S&P
Camera
Eraser
Overall
OpenVLA-OFT
0.8
0.4
0.5
0.7
0.60
DiffTune
0.8
0.3
0.4
0.5
0.50
PHRetune
1.0
0.6
0.6
0.9
0.78
Table 9 : Real-world manipulation success rate with OpenVLA-OFT. We report per-task success rate across the four tasks. PHRetuned OpenVLA-OFT improves success over the nominal policy and DiffTune on every task.
School of Electronic and Information Engineering Tongji University Shanghai 200092, China · Shanghai Research Institute for Intelligent Autonomous Systems Tongji University Shanghai 200092, China