cs.CLSep 29, 2026

ER-JEPA: Experience Replay Improves Joint-Embedding Predictive Learning in Language Models

Authors: Jingnan Pu, Zi-En Fan, Feng Lian

Organizations: School of Automation Science and Technology, Xi’an Jiaotong University No. 28, West Xianning Road, Xi’an, Shaanxi 710049, China

Abstract

Large language models (LLMs) excel at token-level generation but may learn undesirable abstract semantics and lack comprehensive perception. LLM-JEPA mitigates this by aligning different views of the same underlying knowledge via a joint-embedding predictive architecture (JEPA). However, strong alignment does not necessarily lead to accurate, stable predictions. To address this, we propose ER-JEPA, which adds an episodic replay path to LLM-JEPA. ER-JEPA stores training pairs in a memory. At each step, it stores and retrieves relevant data to provide additional supervision. This enables learning from both the current batch and stored training pairs, providing additional supervision for token prediction and representation alignment. Experiments across multiple datasets (NL-RX, GSM8K, Spider, and NQ-Open) demonstrate that ER-JEPA consistently outperforms LLM-JEPA.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. The JEPA Paradox in Language: The Geometry of Linguistic Alternatives

    Jul 26, 2026Anh Trac Duc Dinh, Khang Nhat Hoang VoJoint-Embedding Predictive ArchitecturesLatent Prediction