VLM-Guided Experience Replay
Organizations: Technion · ForSight Robotics · Nvidia Research
Abstract
Recent advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) have enabled powerful semantic and multimodal reasoning capabilities, creating new opportunities to enhance sample efficiency, high-level planning, and interpretability in reinforcement learning (RL). While prior work has integrated LLMs and VLMs into various components of RL, the replay buffer, a core component for storing and reusing experiences, remains unexplored. We propose addressing this gap by leveraging VLMs to guide the prioritization of experiences in the replay buffer. Our key idea is to use a frozen, pre-trained VLM as an automated evaluator to identify and prioritize promising sub-trajectories from the agent's experiences. Across scenarios, including game-playing and robotics, spanning both discrete and continuous domains, agents trained with our proposed prioritization method achieve 15-57% higher average success rates and improve sample efficiency by 35-55% compared to previous approaches. Project page: https://esharony.me/projects/vlm-rb/
Figures & tables
| vs. UER | vs. PER | ||||
|---|---|---|---|---|---|
| Algorithm | Level | Performance ( ) | Sample Efficiency ( ) | Performance ( ) | Sample Efficiency ( ) |
| DQN+IQN (DoorKey) | 8 8 | 0% (1.00/1.00) | 22% (139K/178K) | 0% (1.00/1.00) | 26% (139K/187K) |
| 12 12 | 61% (1.00/0.62) | 57% (277K/640K) | 22% (1.00/0.82) | 37% (292K/467K) | |
| 16 16 | 283% (0.92/0.24) | 53% (420K/893K) | 92% (0.92/0.48) | 51% (305K/622K) | |
| SAC+TD3 (Scene) | 3 | 0% (1.00/1.00) | 40% (213K/354K) | 0% (1.00/1.00) | 20% (213K/266K) |
| 4 | 22% (1.00/0.82) | 62% (295K/776K) | 2% (1.00/0.98) | 38% (394K/634K) | |
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| Hyperparameter | TD3 | SAC |
|---|---|---|
| Network Architecture | ||
| Hidden Dimensions | [512, 512, 512] | |
| Q-Network Layer Norm | True | True |
| Optimization | ||
| Critic Learning Rate | ||
| Actor Learning Rate | ||
| Hyperparameter | DQN | IQN |
|---|---|---|
| Optimization | ||
| Learning Rate | ||
| Batch Size | 128 | |
| Discount Factor ( ) | 0.95 | |
| Target Update Frequency | 1,000 steps | |
| Target Update Rate ( ) | 1.0 (Hard Update) | |
| Priority | DoorKey-16x16 | Scene-5 |
|---|---|---|
| UER | 0.20 (920K, never) | 0.44 (900K, never) |
| PER (TD-only) | 0.76 (330K, 850K) | 0.62 (570K, 970K) |
| VLM-only | 1.00 (310K, 410K) | 0.88 (470K, 670K) |
| 1.00 (310K, 410K) | 0.94 (320K, 440K) |
| Edge | 0.25 | 0.30 | 0.35 | 0.40 | 0.45 | 0.50 | 0.55 | 0.60 | 0.65 |
|---|---|---|---|---|---|---|---|---|---|
| Bin share (%) | 1.45 | 13.95 | 10.30 | 33.05 | 11.30 | 10.50 | 16.00 | 3.20 | 0.25 |
| Balanced accuracy | 0.500 | 0.509 | 0.597 | 0.662 | 0.871 | 0.938 | 0.927 | 0.583 | 0.506 |
| Score form | ASR@1M | Steps to 0.8 |
|---|---|---|
| (2/3) | ||
| Threshold-free ( ) |
| Improvement | |||
| Base. | Perf. ( ) | Sample Eff. ( ) | |
| Pure priority from start | UER | N/A (0.00/0.24) | N/A (1000K/916K) |
| PER | N/A (0.00/0.80) | N/A (1000K/592K) | |
| 0.25 | UER | +233.3% (0.80/0.24) | +40.6% (544K/916K) |
| PER | +0.0% (0.80/0.80) | -0.3% (594K/592K) | |
| 0.50 | UER | +316.7% (1.00/0.24) | +62.9% (340K/916K) |
| Transition class | Buffer share | Frequency / uniform |
|---|---|---|
| VLM-positive | ||
| Successful-episode subset | ||
| VLM-negative | ||
| No VLM score yet |
| Model | Load (GiB) | Peak (GiB) | Time (s) | FPS |
|---|---|---|---|---|
| 1B | 2.86 | 3.77 | 0.46 | 69.27 |
| 3B | 6.56 (+130%) | 8.16 (+116%) | 0.77 (+66%) | 41.75 (-40%) |
| 8B | 18.25 (+539%) | 20.34 (+439%) | 2.15 (+366%) | 14.88 (-79%) |
| Model | Bal. acc. | AUC | Agreement | ||
|---|---|---|---|---|---|
| Perception-LM-1B | 0.50 | 0.94 | 0.98 | 100% | 1.00 |
| Gemma-4-E4B | 0.12 | 0.88 | 0.92 | 86% | 0.70 |
| Qwen3-VL-2B | 0.89 | 0.85 | 0.89 | 81% | 0.73 |
| Model | Calibration | ASR@1M | |
|---|---|---|---|
| Perception-LM-1B | Raw score | None | |
| Gemma-4-E4B | Raw score | None | |
| Gemma-4-E4B | Offline threshold |
| Ground-truth definition | Precision | Recall | FPR | FNR |
|---|---|---|---|---|
| Strict success | 0.69 | 0.99 | 0.12 | 0.01 |
| Milestone acquisition | 0.92 | 0.62 | 0.04 | 0.38 |
| 0.00 | 0.10 | 0.25 | 0.50 | 0.75 | |
|---|---|---|---|---|---|
| ASR@1M | |||||
| vs. UER (0.20) | |||||
| Worst deficit |
| Prompt | Base | Paraph. A | Paraph. B | Terse | Under. | AUC |
|---|---|---|---|---|---|---|
| Base | – | 0.989 | ||||
| Paraphrase A | 0.909 / 0.892 | – | 0.989 | |||
| Paraphrase B | 0.945 / 0.956 | 0.953 / 0.911 | – | 0.987 | ||
| Terse | 0.300 / 0.931 | 0.210 / 0.857 | 0.254 / 0.933 | – | 0.983 | |
| Underspecified | 0.701 / 0.881 | 0.791 / 0.813 | 0.747 / 0.909 | 0.000 / 0.858 | – | 0.955 |
| Hardware | PER | VLM-RB | Rel. Speed |
|---|---|---|---|
| NVIDIA A100 | 111 | 97 | |
| NVIDIA A40 | 92 | 81 | |
| NVIDIA A4000 | 76 | 67 |
| Performance ( ) | Sample Efficiency ( ) | Wall-Clock Saving ( ) | |||
|---|---|---|---|---|---|
| Alg. | Task | Baseline | Best ASR | Steps to Base. Best | Time vs. Baseline |
| SAC | Scene-3 | UER | +0.0% (1.00/1.00) | +50.7% (224K/454K) | +44.7% |
| PER | +0.0% (1.00/1.00) | +18.8% (224K/276K) | +9.1% | ||
| Scene-4 | UER | +4.2% (1.00/0.96) | +54.2% (378K/826K) | +48.7% | |
| PER | +4.2% (1.00/0.96) | +52.8% (378K/800K) | +47.1% | ||
| Scene-5 | UER | +113.6% (0.94/0.44) | +60.0% (366K/914K) | +55.2% |
| Performance ( ) | Sample Efficiency ( ) | |||
|---|---|---|---|---|
| Env Type | Agg. Algorithms | Baseline | Mean Best ASR | Mean Steps to Base Peak |
| Scene | (SAC + TD3) | UER | +32.6% (0.96/0.73) | +55.2% (291K/650K) |
| PER | +15.1% (0.96/0.84) | +34.6% (372K/568K) | ||
| DoorKey | (DQN + IQN) | UER | +57.0% (0.97/0.62) | +51.1% (279K/570K) |
| PER | +27.0% (0.97/0.77) | +42.3% (245K/425K) |