cs.ROJun 17, 2026

DF-ExpEnse: Diffusion Filtered Exploration for Sample Efficient Finetuning

Authors: Calvin LuoChen SunShuran Song

Organizations: 1Stanford University · 2Brown University.

Abstract

A natural recipe for intelligent robotic decision-making is initializing from pretrained generative control policies, which have summarized offline experience, and adapting them to self-collected online experience. We present DF-ExpEnse, an exploration technique that improves the quality of online experience collection, thus increasing finetuning sample-efficiency. DF-ExpEnse leverages the multimodal modeling capabilities of the generative control policy to create an expressive and tractably evaluatable candidate set. It then utilizes an ensemble of critics to identify the action that best balances quality with high exploration interest. In fleet settings, DF-ExpEnse further enables cross-agent communication to facilitate collaborative exploration as a group. DF-ExpEnse can be seamlessly integrated with existing strategies that finetune pretrained generative control policies via reinforcement learning. We experimentally validate consistent sample-efficiency benefits through DF-ExpEnse across a variety of manipulation and locomotion tasks, compared to default finetuning and alternative action selection schemes. Project can be found at https://df-expense.github.io.

Explore similar work

CardsList
  1. OGPO: Sample Efficient Full-Finetuning of Generative Control Policies

    May 4, 2026Sarvesh Patil, Mitsuhiko Nakamoto, Manan Agarwal +14Off-Policy Evaluation