stat.MLSep 27, 2026

The Statistical Benefits of Multiple Responses for Learning from Demonstrations

Authors: Chandramauli Chakraborty, Cong Ma

Organizations: Department of Statistics, University of Chicago

Abstract

Many generative systems return multiple candidate responses and are evaluated according to the best one. Recent work shows that, when demonstrations are optimal, pass@kk can reduce the sample complexity of learning from demonstrations by a logarithmic factor in kk. We ask what happens when the demonstrator is not assumed to be optimal. We find that multiple responses provide a qualitatively stronger benefit in this setting. In a finite reward-class model with no reward feedback, moving from pass@11 to any pass@kk with k≥2k\ge2 changes the worst-case dependence on target accuracy from 1/ε21/\varepsilon^2 to 1/ε1/\varepsilon, uniformly over demonstrator quality. Under standard evaluation, where an unknown reward is fixed before training, increasing kk provides an additional and distinct benefit: the optimal dependence on a reward class of size NN improves from log⁡N\log N to log⁡N/log⁡k\log N/\log k. We further show that these two effects can be separated. Under robust evaluation, where one learned policy must compete with the demonstrator simultaneously for every reward in the class, the fast 1/ε1/\varepsilon dependence persists, while the 1/log⁡k1/\log k improvement can disappear. We establish matching upper and lower bounds in the corresponding regimes and give a greedy multiplicative-weights learner achieving the upper bounds without any assumption on demonstrator quality.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Theoretical Foundations of max⁡\max@kk Reinforcement Learning

    Jul 20, 2026Riccardo Poiani, Martino Bernasconi, Andrea CelliTop-KTheory

  2. Inverse Reinforcement Learning without an Optimal Demonstrator: A Feasible Reward Set Approach

    May 29, 2026Kihyun Kim, Shripad Deshmukh, Nikos Vlassis +1Offline Reinforcement LearningInverse Design