Mulligan: Performance-Guided Data Collection for Efficient On-Robot Learning
Organizations: Stanford University
Abstract
Learning from human demonstrations is a reliable way to teach robots new tasks, but the gains from each additional demonstration shrink as the policy improves. Continued improvement can instead come from supervised deployment, where an operator places the objects and intervenes when the policy fails. We ask how to maximize improvement from a fixed budget of supervised episodes on high-precision manipulation tasks with wide ranges of object placements. We observe that failures can concentrate in a small subset of initial states, so uniform collection spends much of the operator's time on states the policy already handles. Mulligan makes the initial-state distribution a decision, starting each round's episodes at observed failures and untried states. To further improve data efficiency, we augment interactive imitation learning with a value function trained on all data, including failures that imitation discards. Across three real-world tasks evaluated on 2,550 held-out, blinded episodes and two simulated tasks, Mulligan outperforms uniform initial-state sampling at matched collection budgets, and combined with value-based action selection, HiL-IDQL+Mulligan, improves final real-task success by 10-34 percentage points. With operator interventions, the human-robot team completes 98% of collection episodes, remaining productive while the policy learns. Videos, code, and data are available at https://mulligan.page/.
Figures & tables
Appendix figures & tables54 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Arm | Success | Contrast | [CI], |
|---|---|---|---|---|
| Insert Marker R5 | HG-DAgger | ( ) | — | — |
| HG-DAgger+ Mulligan | ( ) | vs. base | pp , | |
| HiL-IDQL+ Mulligan | ( ) | vs. base | pp , | |
| vs. actor | pp , | |||
| Thread Nut R5 | HG-DAgger | ( ) | — | — |
| HG-DAgger+ Mulligan | ( ) | vs. base | pp , |
| Task | Paired | HG-DAgger | HG-DAgger+ Mulligan | Paired [ CI] | Exact McNemar |
|---|---|---|---|---|---|
| Insert Marker | ( ) | ( ) | pp | ||
| Thread Nut | ( ) | ( ) | pp | ||
| Route Cable | ( ) | ( ) | pp |
| Round | Training episodes / arm | Full successes | Mean clip score |
|---|---|---|---|
| R0 | |||
| R1 | |||
| R2 | |||
| R3 | |||
| R4 | |||
| R5 |
| Contrast | R0 | R1 | R2 | R3 |
|---|---|---|---|---|
| Square-Narrow | ||||
| HG-DAgger+ Mulligan vs. HG-DAgger | ||||
| HiL-IDQL+ Mulligan vs. HiL-IDQL | ||||
| HiL-IDQL+ Mulligan vs. HG-DAgger | ||||
| Square-Broad | ||||
| HG-DAgger+ Mulligan vs. HG-DAgger | ||||
| Collection strategy / actor data | Grid SR (%) | [ CI] |
|---|---|---|
| Uniform, no-CF, HO (reference) | — | |
| Uniform, no-CF, auto successes | ||
| Sobol coverage, no-CF, HO | ||
| Sobol coverage, with-CF, HO | ||
| Performance-guided ( ), no-CF, HO | ||
| Performance-guided ( ), with-CF, HO |
| Round | Arm | Fresh / CF | Collection success | Human share |
|---|---|---|---|---|
| R1 | HG-DAgger | / | ||
| HG-DAgger+ Mulligan | / | |||
| R2 | HG-DAgger | / | ||
| HG-DAgger+ Mulligan | / | |||
| R3 | HG-DAgger | / | ||
| HG-DAgger+ Mulligan | / |
| HG-DAgger | HiL-IDQL+ Mulligan | ||||||
|---|---|---|---|---|---|---|---|
| Task | Round | Demos | Corrections | Autonomous | Demos | Corrections | Autonomous |
| Insert Marker | R0 | — | — | — | — | ||
| R1 | — | ||||||
| R2 | — | ||||||
| R3 | — | ||||||
| R4 | — | ||||||
| Task | Objective | Median demo |
| Real world | ||
| Insert Marker | Grasp a pen from the table and insert it into a holder whose mouth leaves mm of clearance. | s |
| Thread Nut | Grasp a square nut and thread it onto a fixed-diameter peg. | s |
| Route Cable | Seat a rope into two clips in succession. | s |
| Simulation | ||
| Square-Narrow | Insert a square nut onto a fixed peg. This is the Robomimic Square task (MimicGen Square_D0 ). | s |
| Task | DoF | Position range ( , cm) | Yaw |
|---|---|---|---|
| Insert Marker | Pen , holder | ||
| Thread Nut | Nut , peg | ||
| Route Cable | Rope free end in only, each clip | clips | |
| Square-Narrow | Nut , peg fixed | ||
| Square-Broad | Nut , peg |
| Task | Success | Headline metric | Cap |
|---|---|---|---|
| Insert Marker | Marker released and fully seated (S7) | success rate | ( s) |
| Thread Nut | Nut released and fully seated (S7) | success rate | ( s) |
| Route Cable | Both clips seated (S10) | clip success rate | ( s) |
| Square-Narrow | Nut on the peg (Robosuite success check) | success rate | ( s) |
| Square-Broad | Nut on the peg (Robosuite success check) | success rate | ( s) |
| Rung | Insert Marker | Thread Nut |
|---|---|---|
| S0 | No useful approach toward the marker. | No useful approach toward the nut. |
| S1 | Approach or grasp attempt without a usable grasp (missed close, or an end pinch that cannot be inserted). | Reach or grasp attempt without a grasp that can lift and carry the nut. |
| S2 | Marker acquired with a usable grasp and lifted clear of the table. | Nut grasped and lifted clear of the table, even if later lost. |
| S3 | Marker brought to the holder area; pressing the holder body without the tip engaging the hole stays S3. | Held nut reaches the peg region in an insertion-directed pose; hole not verified on the peg. |
| S4 | Marker tip engages the hole opening but is not fully seated. | Nut’s hole visibly engages or aligns over the peg top; not seated at the base. |
| S5 | Gripper released a partial insertion that the holder retains. | Gripper released a partial insertion; the nut hangs on the peg above the base. |
| Rung | Route Cable |
|---|---|
| S0 | No useful approach. |
| S1 | Approach or contact without a usable grasp. |
| S2 | A grasp sufficient to route. |
| S3–S5 | Reach, align, and engage the first clip. |
| S6 | One clip seated. |
| S7–S9 | Reach, align, and engage the remaining clip. |
| Task | Rounds | episodes | frames | frames | Critic set (episodes / frames) | Collected episodes | |
|---|---|---|---|---|---|---|---|
| Insert Marker | R0–R5 | / | |||||
| Thread Nut | R0–R5 | / | |||||
| Route Cable | R0–R5 | / | |||||
| Square-Narrow | R0–R3 | / | |||||
| Square-Broad | R0–R3 | / |
| Task | R1 | R2 | R3 | R4 | R5 |
|---|---|---|---|---|---|
| Insert Marker | |||||
| Thread Nut | |||||
| Route Cable |
| Task | Fit | Runs | Median h | Range h | Sum GPU-h |
|---|---|---|---|---|---|
| Square-Narrow | Actor + scalar critic | 80 | 1.90 | 1.33–5.17 | 166.9 |
| Square-Narrow | DIVL critic | 80 | 0.65 | 0.53–1.17 | 56.9 |
| Square-Broad | Actor + scalar critic | 90 | 3.01 | 1.74–6.53 | 298.8 |
| Square-Broad | DIVL critic | 90 | 0.97 | 0.83–1.90 | 91.3 |
| Insert Marker | Actor | 12 | 3.16 | 1.82–7.85 | 46.3 |
| Insert Marker | Critic | 4 | 5.29 | 1.61–10.30 | 22.5 |
| Hyperparameter | Simulation | Real world |
| Architecture | ||
| Backbone | 1-D conditional U-Net over the action chunk | |
| U-Net channel widths | ||
| Kernel size / group-norm groups | / | |
| Diffusion-step conditioning | -D embedding, FiLM scale modulation | |
| Observation encoder | none (state vector) | ResNet18, one per camera |
| Hyperparameter | Simulation | Real world |
| Objective | ||
| Value head | categorical (DIVL) | categorical (DIVL) |
| Expectile / base quantile | ||
| Discount | ( Thread Nut ), ( Insert Marker , Route Cable ) | |
| Target-network rate | (Polyak) | |
| -ensemble | networks, min aggregation for both the TD label and reranking | |
| Hyperparameter | Simulation | Real world |
| Network and optimization | ||
| and trunk | MLP with layer norm | |
| input | observation features concatenated with the flattened -step chunk | |
| Visual features | — | the deployed actor’s encoder, frozen; features precomputed over augmented views per frame |
| Proprioception dropout | ||
| Extra critic head | none | state-conditioned FiLM rescaling of the critic’s action input ( Sec. G.3 ), used in the round-5 Insert Marker and Thread Nut critics and in all Route Cable critics |
| Setting | Value |
|---|---|
| Candidates , simulation | |
| Candidates , Insert Marker | at R2–R4, at R5 |
| Candidates , Thread Nut | at R3–R4, at R5 |
| Candidates , Route Cable | at R3–R5 |
| Scored quantity | , min over the -network ensemble |
| Selection rule | argmax |
| Setting | Simulation | Real world |
| Sobol draw | ||
| Sampled dimensions | ( Square-Narrow : nut , yaw), ( Square-Broad ) | ( Insert Marker , Thread Nut ), ( Route Cable ) |
| Remaining dimensions | Square-Narrow nut : drawn independently uniform | — |
| Scrambling | linear matrix scrambling with a digital shift (SciPy), always on | |
| Pool size | ( Square-Narrow ), before rejection ( Square-Broad ) | ( Insert Marker , Thread Nut ), ( Route Cable ) |
| Rejection | Square-Broad : nut–peg center distance above cm | Route Cable : clip separation at least cm |
| Setting | Simulation | Real world |
|---|---|---|
| Platform | Robosuite with MimicGen assets | Franka Panda on DROID |
| Control rate | Hz | Hz |
| Controller | operational-space pose, delta, base frame | Cartesian velocity |
| Per-step action limits | cm, rad | — |
| Cameras | none; state observations | ZED stereo, HD720 captured at fps |
| Teleoperation | SpaceMouse, with operator-triggered interventions | |
| Cohort | DP IQL | RECAP | (pp) | Paired CI |
|---|---|---|---|---|
| Square-Narrow R2 | ||||
| Square-Narrow R3 | ||||
| Square-Broad R3 |
| Operator protocol | Env steps | Epi- sodes | Operator hours | Human frames | Takeover episodes | Training success | Policy-only successes | Final eval | Peak eval |
|---|---|---|---|---|---|---|---|---|---|
| Square-Narrow | |||||||||
| Step 0, dense | – k | ||||||||
| Step 0, sparse bursts | – k | ||||||||
| Step 0, long bouts | – k | ||||||||
| Buffer, then none | – k | – | |||||||
| After takeoff | – k | ||||||||
| Task | All-data | Human-only | at | at (same actor) | Rerank gain all-data | Rerank gain human-only |
|---|---|---|---|---|---|---|
| Square-Narrow | ||||||
| Square-Broad |
| Condition | Seeds | Final SR (%) |
|---|---|---|
| Baseline, reward , clipped targets | 5 | |
| Shifted reward , no target clipping | 5 | |
| Intervention frames treated as terminal | 5 | |
| Actor update cadence | 5 | |
| Actor update cadence | 5 | |
| Flat (uniform) critic sampling | 3 |
| Task | Cells | Mean (pp) | Cell range (pp) | Cells favoring DIVL / scalar |
|---|---|---|---|---|
| Square-Narrow | / | |||
| Square-Broad | / | |||
| Both tasks | / |
| Regime (scalar-IQL base SR) | scalar IQL | divl-fixed | divl-adaptive |
|---|---|---|---|
| Low SR ( ) | |||
| High SR ( ) | |||
| High SR ( ) |
| Critic sampler | scalar IQL (control) | divl-fixed | divl-adaptive |
|---|---|---|---|
| straddled | ( ) | ( ) | |
| balanced_outcome | ( ) | ( ) | |
| flat | ( ) | ( ) |
| Critic ensemble | ensemble | Final SR (%) | |
|---|---|---|---|
| - min | — | ||
| REDQ-2 | |||
| Single random critic |
| Method | Task | Varies | DoF | Position range | Yaw | Source |
| HiL-SERL [ 20 ] | IKEA side panel | arm | 3 | cm in , | Tab. 6 | |
| HiL-SERL [ 20 ] | Cable clipping | object | 3 | cm in , | Tab. 5 | |
| ConRFT [ 22 ] | Pick banana | object | 2 | cm in , | none | Tab. V |
| ConRFT [ 22 ] | Insert wheel | object | 2 | cm in , | none | Tab. XI |
| RL-100 [ 25 ] | Box folding | object | 3 | cm | unstated ‡ | Suppl. |
| RL-100 [ 25 ] | Orange juicing | object | 3 | cm tray | unstated ‡ | Suppl. |
| Method | Task | Demos | Human corrections | Autonomous experience | Robot time | Source |
| HiL-SERL [ 20 ] | IKEA side panel | rate curve only † | / transitions ‡ (p1/p2) | / h | Tab. 1, 6 | |
| HiL-SERL [ 20 ] | Cable clipping | rate curve only † | transitions ‡ | h | Tab. 1, 5 | |
| ConRFT [ 22 ] | Pick banana | rate curve only † | – episodes ∗ | min | Tab. I, V | |
| ConRFT [ 22 ] | Insert wheel | rate curve only † | – episodes ∗ | min | Tab. I, XI | |
| RL-100 [ 25 ] | Box folding | ( h) | none § | offline + online epi. | h | Tab. S3 |
| RL-100 [ 25 ] | Orange juicing | ( h) ¶ | none § | offline + online epi. ¶ | h ¶ | Tab. S3 |