Generative action models based on diffusion and flow matching have been increasingly adopted in vision-language-action (VLA) policies for their ability to capture diverse behaviors, including multiple valid action sequences under the same observation and instruction. Their iterative sampling procedures, however, require repeated network evaluations to generate each action chunk, increasing inference latency in closed-loop control. We propose ESP (Energy-Score Policy), a teacher-free approach that maps policy context and noise directly to an action chunk in a single network evaluation. ESP trains the action head with the energy score rather than mean squared error. Whereas squared-error regression targets the conditional mean, the energy score is strictly proper: its expected value is uniquely minimized by the target distribution. This provides a principled objective for learning multimodal action distributions without iterative sampling, with exact recovery at the population optimum when the model can represent the target distribution. Experiments on both simulation and real-world manipulation tasks demonstrate competitive task success with substantially lower action-generation latency than the flow matching baseline. These results support direct distributional learning as an efficient alternative to iterative generative robot policies.
Figures & tables
Fig. 1: Energy-Score Policy (ESP) for one-step multimodal action generation. Flow matching can generate multimodal actions but requires integration of a velocity field over multiple network evaluations (NFE >1 ). Mean squared error (MSE) regression is single-step (NFE =1 ) but collapses to the conditional mean and cannot represent multimodality. Trained with the strictly proper energy score, ESP generates multimodal actions in one network evaluation (NFE =1 ).
Fig. 2: ESP training and one-step action generation. The generator maps a condition c and independent prior samples zi∼pz directly to action chunks xi=fθ(c,zi) . Training evaluates all K outputs with the energy score, whereas inference draws one z and requires one network evaluation.
Fig. 3: TwoBranch simulation results. The target is y=sc+0.08ϵ with equiprobable signs. The one-step group uses one network evaluation; the multi-step reference evaluates the same trained flow model with 10 Euler steps. Match ↑ is the percentage of predictions within 3× noise standard deviations of either branch center at the same input c . Results use one seed and illustrate behavior rather than establishing a discriminating benchmark.
Fig. 4: Suite-wise π0.5 closed-loop success on LIBERO. Left: mean success rates in percent, with sample SD error bars. Right: the same results as success fractions (mean ± SD). Statistics use three training seeds; Average first averages the four suites within each seed.
Path
ESP, NFE =1
XM, NFE =1
Flow, NFE =10
Action expert
3.35
3.31
28.39
Cached end to end
4.46
4.22
29.14
Uncached end to end
59.01
58.78
82.54
TABLE I: Policy inference latency (ms)
Fig. 5: Fixed-initialization end-effector XY trajectories. Rows show Tasks 0, 7, and 8 from the six-task experiment. Columns show the scene, ESP and XM at NFE =1 , flow at NFE =10 , and demonstrations. Scene labels show objects A, B and placement goals GA,GB , and obstacles visible in the scene are unlabeled. Each policy subfigure shows all 32 rollouts. The success/failure bar shows the success rate across rollouts with different policy-noise sequences. Successful trajectories progress from purple to yellow over time, and failures are gray. The black dot marks the shared rollout start. Task 8 success additionally requires a non-spatial goal condition (the stove-on predicate), which is not visible in the XY paths.
ESP (Ours)
ESP (Ours), expanded
FFN width
4096
73728
Spatial
0.955±0.036
0.953±0.019
Object
0.970±0.013
0.985±0.009
Goal
0.933±0.019
0.953±0.008
Long10
0.813±0.023
0.857±0.012
Average
0.918±0.014
0.937±0.004
TABLE II: Action-expert capacity ablation within ESP (success rate mean ± SD)
Fig. 6: Franka environment, tasks, language instructions, and physical rollout outcomes. The environment figure shows the base and wrist cameras, task figures show representative drawer and pick & place examples from the physical policy rollouts. The bottom two strips show observations over time from successful Drawer Closing (upper) and Pick & Place (lower) rollouts. The table reports successful trials / total trials.
Fig. 7: Training progress on LIBERO- Spatial ( K=8 ). Success rate over (a) expert-action exposure and (b) training time for checkpoints from the same runs as the main π0.5 results. Although flow matching processes optimizer steps 3.81× faster, ESP at 15k steps (4.78h) achieves 93.0% success, comparable to 93.5% for flow matching at 60k steps (4.98h).