cs.LGFeb 2, 2026

VLM-Guided Experience Replay

Authors: Elad Sharony, Tom Jurgenson, Orr Krupnik, Dotan Di Castro, Shie Mannor

Organizations: Technion · ForSight Robotics · Nvidia Research

Abstract

Recent advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) have enabled powerful semantic and multimodal reasoning capabilities, creating new opportunities to enhance sample efficiency, high-level planning, and interpretability in reinforcement learning (RL). While prior work has integrated LLMs and VLMs into various components of RL, the replay buffer, a core component for storing and reusing experiences, remains unexplored. We propose addressing this gap by leveraging VLMs to guide the prioritization of experiences in the replay buffer. Our key idea is to use a frozen, pre-trained VLM as an automated evaluator to identify and prioritize promising sub-trajectories from the agent's experiences. Across scenarios, including game-playing and robotics, spanning both discrete and continuous domains, agents trained with our proposed prioritization method achieve 15-57% higher average success rates and improve sample efficiency by 35-55% compared to previous approaches. Project page: https://esharony.me/projects/vlm-rb/

Figures & tables

Appendix figures & tables22 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning

    Apr 18, 2026Weiyu Ma, Yongcheng Zeng, Yan Song +4Large Language Model Reinforcement LearningPrioritized Experience Replay

  2. Rollout-Level Advantage-Prioritized Experience Replay for GRPO

    Jun 3, 2026Gyeongtae Yoo, Sanghyeok Park, Soohyuk Jang +2Parallel RolloutsReplay