Large language models (LLMs) excel at token-level generation but may learn undesirable abstract semantics and lack comprehensive perception. LLM-JEPA mitigates this by aligning different views of the same underlying knowledge via a joint-embedding predictive architecture (JEPA). However, strong alignment does not necessarily lead to accurate, stable predictions. To address this, we propose ER-JEPA, which adds an episodic replay path to LLM-JEPA. ER-JEPA stores training pairs in a memory. At each step, it stores and retrieves relevant data to provide additional supervision. This enables learning from both the current batch and stored training pairs, providing additional supervision for token prediction and representation alignment. Experiments across multiple datasets (NL-RX, GSM8K, Spider, and NQ-Open) demonstrate that ER-JEPA consistently outperforms LLM-JEPA.
Figures & tables
Figure 1 : ER-JEPA outperforms LLM-JEPA across tasks.
Figure 2 : Comparison of baseline training, LLM-JEPA, and ER-JEPA. (a) The baseline learns through next token prediction. (b) LLM-JEPA adds prediction in the embedding space to learn abstract relations between inputs and targets. (c) ER-JEPA adds a replay path to LLM-JEPA. It stores past training examples and replays their tokens, providing additional supervision for both token prediction and representation alignment. Past experience thus continues to shape learning as the model changes, helping it correct errors and reinforce learned knowledge for more reliable predictions.
Figure 3 : LLM-JEPA can achieve strong alignment yet still make errors and lose previously correct predictions. For LLM-JEPA, the JEPA loss converges in (a). Yet in (b) and (c), errors remain even when cosine similarities between inputs and targets are close to one. Some correct predictions also become incorrect as training continues. For ER-JEPA, the JEPA loss also converges in (d). In (e) and (f), more predictions are correct despite slightly lower cosine similarities. At each checkpoint, ER-JEPA consistently has more correct answers than LLM-JEPA. These results suggest that replay supervision helps resolve errors that alignment alone leaves unresolved. This figure uses Llama-3.2-1B on 2000 SYNTH test examples with seed 82. Heatmap rows are sorted independently within each method by final-checkpoint correctness and prediction history.
Figure 4 : ER-JEPA improves accuracy at matched training compute and reduces overfitting. (a) All three replay policies outperform LLM-JEPA at every evaluated budget. These gains support the benefit of introducing the replay path. (b) Baseline loses accuracy after an early increase. LLM-JEPA delays this decline but still loses accuracy later. Content and uniform replay maintain their gains, while hard replay shows only a small final decline. These trends suggest that replay mitigates overfitting. They complement the mean curves in Figure 1 (right). Results use Llama-3.2-1B on SYNTH at six compute budgets from 40.05 to 240.31 PFLOPs.
Figure 5
Baseline error category
Baseline errors
Method
Corrected
Correction rate (%)
Over-generation
718.4±112.5
LLM-JEPA
378.8±64.6
53.07±7.92
ER-JEPA
610.8±126.1
84.51±4.70
Under-generation
2.6±1.5
LLM-JEPA
0.6±0.5
24.00±25.10
ER-JEPA
1.2±0.8
54.67±44.07
Same-length mismatch
168.4±10.9
LLM-JEPA
34.8±5.8
20.58±2.35
ER-JEPA
42.4±7.0
25.07±2.62
Table 1 : ER-JEPA achieves higher mean correction rates across all three Baseline error categories. Results use Llama-3.2-1B on SYNTH at 240.309 PFLOPs. Values are means ± one sample standard deviation over five seeds. Correction rates are computed separately for each seed before averaging.
Figure 7 : ER-JEPA achieves higher accuracy and corrects more errors during training. (a) Test accuracy at comparable training compute for SFT, LLM-JEPA, and ER-JEPA. (b) The fraction of wrong predictions that become correct at the next checkpoint. (c) The fraction of correct predictions that become wrong at the next checkpoint. Compared with LLM-JEPA, ER-JEPA improves error correction across all five seeds, while the reduction in lost correct predictions is smaller and varies across seeds. Results use Llama-3.2-1B on the same 2,000 SYNTH test examples.
Figure 8
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 10 : ER-JEPA predictions on examples that LLM-JEPA predicts incorrectly. Rows indicate LLM-JEPA error categories, and columns indicate ER-JEPA prediction categories for the same input and seed. The row total n counts LLM-JEPA errors pooled across five seeds. Each cell shows its count divided by n , expressed as a percentage to one decimal place. Results use Llama-3.2-1B on 2,000 SYNTH test examples per seed at a matched training compute budget of 240.309 PFLOPs.
Subset
Classification rule
Meaning
Complement
Contains the operator ~ .
Excludes strings matched by a subexpression.
Intersection
Contains the operator & .
Requires multiple patterns to hold simultaneously.
Alternation
Contains the operator | .
Allows a choice between alternative patterns.
Quantifier
Contains * , + , ? , or a brace quantifier such as {m,n} .
Specifies repetition counts or optionality.
Boundary / anchor
Contains \b , \B , ^ , or $ .
Constrains matching positions relative to word boundaries or string endpoints.
Character class
Contains a bracketed character class, such as [a-z] .
Specifies a set of allowed characters.
Appendix
Table 2 : Definitions of the six regex structural subsets. Each subset contains examples whose target regular expressions include the specified feature.
Seed
Baseline errors
LLM-JEPA corrected
Correction rate (%)
ER-JEPA corrected
Correction rate (%)
(a) Over-generation
4
901
448
49.72
807
89.57
23
697
435
62.41
583
83.64
37
705
375
53.19
608
86.24
82
591
344
58.21
455
76.99
84
698
292
41.83
601
86.10
Appendix
Table 3 : Correction counts and rates for each seed on the same Baseline errors. Both methods use the same inputs and seed within each row. Corrected subsets may overlap between methods. The counts and rates in Table 1 are summarized across these five seeds.
Figure 11 : Example predictions from Baseline, LLM-JEPA, and ER-JEPA on the SYNTH dataset.
Figure 12 : Example predictions from Baseline, LLM-JEPA, and ER-JEPA on the NQ-OPEN dataset. Correct and incorrect answers are shown in dark green and dark red, respectively. For each prediction, only the first semicolon-separated answer is displayed.
Figure 13 : Raw accuracy at six shared target PFLOPs for seeds 4, 23, 37, and 82. Each panel shows one seed. Seed 84 is shown in Figure 4 (b). The panels do not show means or error bars.
Model / target
Regular expression
Input: lines not having the string “dog” followed by a number, 3 or more times
Ground truth
~((dog.[0-9].){3,})
LLLM
~((dog.[0-9].){3,})
LLLM−JEPA
~((dog.[0-9].){3,})
LER−JEPA
~((dog.[0-9].){3,})
Input: lines containing ending with a vowel, zero or more times
Appendix
Table 4 : Some regular expressions generated by Llama-3.2-1B-Instruct after fine-tuning with LLLM , LLLM−JEPA , and LER−JEPA losses. Color code: wrong , extra , missing .
Method
Accuracy (%) ↑
Min
Max
ER-JEPA (Sparse Retrieval)
86.37±0.35
85.75
86.60
Dense Retrieval
86.21±0.37
85.85
86.70
Appendix
Table 5 : Sparse and dense memory addressing for content replay. All other training settings are fixed. Results report mean accuracy, standard deviation, minimum, and maximum over five seeds.
Memory capacity M
Accuracy (%) ↑
PFLOPs
Time (min)
Tokens (M)
10
84.22±0.75
181.536±0.211
10.48±0.07
8.581±0.166
102
84.32±0.65
219.605±0.533
11.62±0.22
9.427±0.175
103
84.76±0.52
219.766±0.439
16.74±1.33
10.969±0.097
104
84.98±0.35
219.667±0.393
22.23±0.17
11.252±0.010
Appendix
Table 6 : Ablation on episodic memory capacity M for Meta-Llama-3.2-1B-Instruct on NL-RX-SYNTH. Values at epoch 4 are mean ± standard deviation over five seeds. The observed training costs are not compute matched.
Figure 14 : Accuracy across four training epochs for four episodic memory capacities. Points show means over five seeds. Error bars show one sample standard deviation. These runs are not compute matched.
Figure 15 : Ablation on the JEPA hyperparameters λ and k . Each entry reports accuracy (%), with the best result in bold.
Center for AI Research (CAIR), VinUniversity, Hanoi, Vietnam · Faculty of Computer Science and Engineering Ho Chi Minh City University of Technology (HCMUT), VNU-HCM Ho Chi Minh City, Vietnam · Mohamed bin Zayed University of Artificial Intelligence Abu Dhabi, United Arab Emirates