Diffusion and VLA policies for manipulation are often deployed through downstream impedance controllers. The stiffness and damping gains of these controllers affect task success, yet are commonly inherited from data collection rather than selected for the deployed policy. Although empirical gain sweeps can improve performance, they require repeated evaluation rollouts. We introduce PHRetune, an offline method that derives controller gains for a frozen policy without evaluation rollouts or gain search. Our approach learns a port-Hamiltonian model from demonstrations to estimate the effort and energy associated with the policy's predicted actions. The policy is applied to recorded demonstration observations, and its predictions are assessed against demonstration-derived effort and energy budgets. From this comparison, we derive a single gain scale in closed form, adjusting the downstream controller while preserving the policy and its action representation. The gains are fixed before evaluation, without requiring a prior manipulator model, task rewards, policy retraining, or additional runtime computation. Across LIBERO suites, PHRetune improves Diffusion Policy success by up to 9.4 percentage points, with the derived gains achieving the highest observed success rates in empirical gain sweeps. On all four real-world manipulation tasks, PHRetuned Diffusion Policy outperforms the nominal policy, alternative gain-tuning methods, and a policy-retraining baseline. The same procedure improves success with SmolVLA and OpenVLA-OFT on every task, while reducing acceleration and jerk for both VLA backbones.
Figures & tables
Backbone
Model
Goal
Obj.
Spa.
Diffusion Policy
Inv- Σ
0.229
0.232
0.244
DiffTune
1.123
1.140
1.131
PHRetune
0.699
0.461
0.531
Table 1 : Simulation tuned gains. We report tuned β scaling for PHRetune, Inv- Σ , and DiffTune. Robosuite defaults kp=150 with kd coupled via a fixed 1 damping ratio. The equivalent β values for Inv- Σ and DiffTune are calculated after their respective derivation.
Demos
Model
Goal
Obj.
Spa.
Total
20
Diffusion Policy
84.6
76.7
74.3
78.5
Inv- Σ
77.1
77.5
81.7
78.8
DiffTune
80.9
71.5
67.9
73.4
MtG-SM2M
86.1
66.7
80.9
77.9
PHRetune
89.0
86.4
88.3
87.9
50
Diffusion Policy
92.2
86.1
88.7
89.0
Table 2 : Diffusion Policy success rate. We report the single-task success rate on LIBERO-goal, -object, and -spatial subsets trained on 20 and 50 demos. PHRetuned Diffusion Policy achieves the highest success rate compared to baseline methods. Notably, PHRetuned Diffusion Policy trained on 20 demos approaches the performance of vanilla Diffusion Policy trained on the full 50.
Figure 1 : LIBERO gain sweep. We report the success rate over three seeds for the (a) goal, (b) object, and (c) spatial partition of LIBERO by uniformly sweeping the gain scale. The success rate for PHRetune and baselines are highlighted. MtG is omitted from the curve and β values for Inv- Σ and DiffTune are calculated post-hoc using derived gain values.
Mode
Goal
Obj.
Spa.
Total
Close Loop
β∗
0.646
0.460
0.546
88.0
SR
89.7
86.3
88.1
Shadow Mode
β∗
0.699
0.461
0.531
87.9
SR
89.0
86.4
88.3
Table 3 : Gain derivation in shadow mode and in closed loop. Energy-based gain scale β∗ and success rate of the 20-demonstration policies per LIBERO suite. Shadow mode scores the frozen policy’s predicted chunks on the demonstration frames alone. Close Loop scores the chunks it executes over simulator rollouts. Both routes recover nearly the same gains and overall success rate.
Figure 2 : Real-world manipulation tasks. (a) Placing a hat on a stand, (b) transferring salt-and-pepper mills into a tray, (c) placing a camera in its fitted holder, and (d) grasping an eraser and wiping a whiteboard.
Task
Task Prompt
Hat
Place the hat on the stand
S&P
Put the salt and pepper shakers in the box
Camera
Place the camera into the slot in the box
Eraser
Pick up the eraser and erase the whiteboard
Table 4 : Real-world task prompts. We report the task prompt for all four tasks. Real-world diffusion policy receives no language conditioning; SmolVLA receives language conditioning for single-task policies; and OpenVLA-OFT receives language conditioning for all four tasks.
Backbone
Model
Hat
S&P
Camera
Eraser
Diffusion Policy
Inv- Σ
0.217
0.178
0.163
0.166
DiffTune
1.109
1.068
1.098
1.104
PHRetune
0.696
0.461
0.707
0.859
SmolVLA
PHRetune
0.667
0.460
0.513
0.630
OpenVLA-OFT
PHRetune
0.731
0.492
0.476
0.570
Table 5 : Real-world tuned gains. We report the derived β∗ for Diffusion Policy, SmolVLA, and OpenVLA-OFT.
Model
Hat
S&P
Camera
Eraser
Overall
Tele-Op
1.00
0.87
0.87
0.91
0.91
Diffusion Policy
0.6
0.4
0.5
0.5
0.50
Inv- Σ
0.2
0.1
0.0
0.0
0.08
DiffTune
0.5
0.3
0.5
0.2
0.38
MtG-SM2M
0.6
0.4
0.4
0.4
0.45
PHRetune
0.8
0.6
0.7
0.8
0.73
Table 6 : Real-world manipulation success rate with Diffusion Policy. We report per-task success rate and their mean across the four tasks. PHRetuned Diffusion Policy outperforms the nominal policy and baseline methods on every task. Tele-Op provides a human teleoperation reference.
Model
Hat
S&P
Camera
Eraser
Overall
SmolVLA
0.7
0.5
0.4
0.6
0.58
DiffTune
0.7
0.3
0.3
0.5
0.45
PHRetune
0.9
0.7
0.7
0.9
0.80
Table 7 : Real-world manipulation success rate with SmolVLA. We report per-task success rate across the four tasks. PHRetuned SmolVLA improves success over the nominal policy and DiffTune on every task.
Metric
SmolVLA
PHRetune
Tele-Op
EE Jerk RMS
22.4±4.8
11.7±4.3(↓47.7%)
9.3±1.2
EE Jerk Peak
77.2±17.4
39.3±20.1(↓49.1%)
31.5±9.0
EE Accel. RMS
0.93±0.23
0.52±0.17(↓44.1%)
0.43±0.07
EE Accel. Peak
3.09±0.58
1.65±0.58(↓46.6%)
1.39±0.48
Joint Accel. RMS
7.91±3.77
3.34±1.74(↓57.8%)
2.30±0.34
Time to Grasp
3.66±0.68
3.66±0.15
2.99±0.48
Table 8 : Real-world motion dynamics with VLAs. We report end-effector and joint motion metrics, time to grasp, controller gain, and success rate for nominal and PHRetuned SmolVLA, PHRetuned OpenVLA-OFT over ten inference rollouts of the hat task. PHRetune reduces acceleration and jerk toward teleoperation levels with little increase to the time to grasp. Failures do not contribute to the reported metrics. End effector (EE) jerk is measured in m s -3 , EE accel. in m s -2 , joint accel. in rad s -2 , and time to grasp in seconds.
Model
Hat
S&P
Camera
Eraser
Overall
OpenVLA-OFT
0.8
0.4
0.5
0.7
0.60
DiffTune
0.8
0.3
0.4
0.5
0.50
PHRetune
1.0
0.6
0.6
0.9
0.78
Table 9 : Real-world manipulation success rate with OpenVLA-OFT. We report per-task success rate across the four tasks. PHRetuned OpenVLA-OFT improves success over the nominal policy and DiffTune on every task.
World models built on recurrent state space architectures enable efficient latent imagination, yet remain physically unstructured, producing dynamics that violate conservation and dissipative principles. We introduce a unified Port-Hamiltonian framework that remedies this through three synergistic mechanisms. First, we embed implicit physical priors into recurrent transitions by modeling projected latent evolution as action controlled energy routing governed by flow and dissipation, biasing the projected PH phase space toward a more compact and physically structured representation. Second, we develop a kinematics aware energy world model that estimates the Hamiltonian and power balance from proprioceptive observations, providing an explicit physical signal for thermodynamic reasoning. Third, leveraging these energy gradients, we establish an energy guided Actor-Critic that uses Lagrangian multipliers to regularize policy optimization toward lower energy and smoother control. Across visual control benchmarks, this paradigm not only attains superior asymptotic returns but also elevates internal simulator fidelity by establishing a tighter, lower variance alignment between imagined and real rewards, all while reducing latent phase space volume by 4.18-8.41%, energy consumption by up to 7.80%, and mean squared jerk by up to 9.38%.
Xueyu Luan, Chenwei Shi
School of Electronic and Information Engineering Tongji University Shanghai 200092, China · Shanghai Research Institute for Intelligent Autonomous Systems Tongji University Shanghai 200092, China
We develop a physics-informed learning framework for energy-shaping control of port-Hamiltonian (pH) systems from trajectory data. The proposed approach co-learns a pH system model and an optimal energy-balancing passivity-based controller (EB-PBC) through alternating optimization with policy-aware data collection. At each iteration, the system model is refined using trajectory data collected under the current control policy, and the controller is re-optimized on the updated model. Both components are parameterized by neural networks that embed the pH dynamics and EB-PBC structure, ensuring interpretability in terms of energy interactions. The learned controller renders the closed-loop system inherently passive and provably stable, and exploits passive plant dynamics without canceling the natural potential. A dissipation regularization enforces strict energy decay during training, thereby enhancing robustness to sim-to-real gaps. The proposed framework is validated on state-regulation and swing-up tasks for planar and torsional pendulum systems.
Ankur Kamboj, Biswadip Dey, Vaibhav Srivastava
Electrical and Computer Engineering, Michigan State University, East Lansing, MI 48824 · Meta Reality Labs, Redmond, WA 98052
Reinforcement learning (RL) is a powerful and convenient tool to modernize controller design. In this work, we study the zero-shot transfer of RL-based control policies from simulation to hardware for cart-pole swing-up and stabilization. The two policies are trained independently, and the handoff is implemented in Simulink via switching logic. We apply a first-order action smoothing filter to prevent hardware damage from high-frequency oscillatory actuation. Pairing this bandwidth-aware filtering with sensitivity-guided domain randomization (DR) and a simple linear curriculum learning (CL) schedule, we obtain a swing-up policy that in all of our experiments injects sufficient energy for handoff into the stabilizer's region of attraction. The stabilization policy rejects disturbances within the tested range, and the swing-up policy can re-engage after larger perturbations and restores the pendulum to the inverted position.
Nikki Xu, Hien Tran
Department of Mathematics, Box 8205, NC State University, Raleigh, NC 27695