Prior audio-visual self-supervised learning methods rely on mechanisms such as EMA target encoders, prediction heads, reconstruction decoders, and contrastive losses. We introduce LeAVJEPA, the first audio-visual encoder trained under LeJEPA's collapse-free objective. A single early-fusion Vision Transformer processes audio, video, and joint audio-video inputs. Modality dropout treats a missing modality as another view of the same event, making cross-modal alignment implicit in the objective. The model aligns global embeddings with modality-specific local embeddings, and SIGReg prevents representational collapse. A controlled ablation identifies modality dropout as the key mechanism for audio-visual alignment. Despite the architectural simplicity, LeAVJEPA reaches 36.0 mAP on AudioSet-20K and 91.3% accuracy on ESC-50 under frozen evaluation. After fine-tuning, it reaches 61.1% accuracy on VGGSound, and its embeddings support zero-shot audio-visual retrieval.
Figures & tables
Method
Single shared encoder
No raw-input recon.
No contr. negatives
No decoder
No teacher / EMA
No predictor
Training objective
AV-MAE ( Georgescu et al., 2023 )
✗
✗
✓
✗
✓
✓
Recon.
CAV-MAE ( Gong et al., 2023 )
✗
✗
✗
✗
✓
✓
Recon. + contr.
MAViL ( Huang et al., 2023 )
✗
✗
✗
✗
✗
✓
Raw/latent recon. + contr.
CAV-MAE Sync ( Araujo et al., 2025 )
✗
✗
✗
✗
✓
✓
Recon. + contr.
MJEPA ( Teotia et al., 2026 )
✓
✓
✓
✓
✗
✗
Latent pred.
LeAVJEPA (ours)
✓
✓
✓
✓
✓
✓
Invariance + SIGReg
Table 1: Comparison of audio-visual SSL methods. LeAVJEPA is the only method that requires none of the listed components.
Figure 2: Attention finds the sound source without localization supervision. Audio-to-video attention lands on the object producing the sound, and video-to-audio attention follows its harmonic and temporal structure. Neither behaviour is explicitly supervised; both emerge from matching single-modality views to a joint centre.
Figure 3: Semantic classes cluster, and a clip’s views stay close. t-SNE of [CLS] embeddings, where each clip contributes an audio-only ( ), a video-only ( ), and a joint ( ) point. Within each cluster, the three shapes mix: audio-only and video-only views of a class land together even though they share no input. The zoom shows the guitar sample, whose joint embedding ( zg ) lies between its audio ( z^a ) and video ( z^v ) embeddings, as the invariance loss (bottom) encourages.
Figure 4: Modality-specific views improve cross-modal learning in LeAVJEPA. Left: joint local views remain near the audio-only baselines (dashed lines). Random modality dropout improves probe accuracy, and making every local view audio-only or video-only yields the largest gain. Right: either loss term alone fails to learn useful features.
AudioSet-20K (mAP)
ESC-50
Method
Params
Pre-train data
A
V
A-V
Acc.
Separate encoders
CAV-MAE ( Gong et al., 2023 )
170M
IN+AS
19.38
18.14
34.59
77.5
MAViL ( Huang et al., 2023 )
170M
IN+AS
30.00
–
–
90.8
EquiAV ( Kim et al., 2024 )
170M
IN+AS
34.25
18.60
38.60
93.2
CAV-MAE Sync ( Araujo et al., 2025 )
170M
IN+AS
21.66
16.20
28.50
89.2
XKD ( Sarkar & Etemad, 2024 )
170M
AS
–
–
–
93.6
Table 2: Frozen transfer with an attentive probe. AudioSet-20K mAP for audio-only (A), video-only (V), and audio-video (A-V) inputs, and ESC-50 accuracy. Baseline frozen results and parameter counts are as reported by Teotia et al. (2026) . Best result per column in bold. ∗ Not reported in the MJEPA paper; provided by the authors through personal correspondence.
Audio → video
Video → audio
Method
R@1
R@5
R@10
R@1
R@5
R@10
CAV-MAE ( Gong et al., 2023 )
15.1
34.0
43.0
18.8
39.5
50.1
CAV-MAE Sync ( Araujo et al., 2025 )
27.9
52.4
62.2
35.2
58.3
67.6
EquiAV ( Kim et al., 2024 )
29.6
53.7
63.1
30.1
53.3
62.9
MJEPA ViT-L † ( Teotia et al., 2026 )
29.3
53.7
61.8
30.6
53.4
61.3
LeAVJEPA ViT-B (ours)
10.5
28.9
36.6
11.4
28.2
36.8
Table 3: AudioSet cross-modal retrieval on the CAV-MAE balanced evaluation subset (1,725 clips; 1,722 available for our evaluation), recall in percent. Baseline rows are as reported in the respective papers and all train an explicit cross-modal alignment term or predictor; LeAVJEPA rows are zero-shot cosine retrieval unless marked. Best baseline result per column in bold. † MJEPA retrieves through its jointly trained cross-modal predictor with an L1 nearest-neighbour search rather than direct embedding similarity. ‡ Contrastive alignment head trained on frozen features; not zero-shot.
Method
Params
Init.
Pretraining
Top-1
AV-MAE ( Georgescu et al., 2023 )
172M
–
VGGS
64.2
AV-MAE ( Georgescu et al., 2023 )
611M
–
VGGS
65.0
AVSiam ( Lin & Bertasius, 2024 )
100M
IN-21K SL
AS-2M
64.9
AVSiam ( Lin & Bertasius, 2024 )
332M
IN-21K SL
AS-2M
67.1
CAV-MAE ( Gong et al., 2023 )
164M
IN-SSL
AS-2M
65.5
MAViL ( Huang et al., 2023 )
172M
IN-SSL
AS-2M
67.1
Table 4: Audio-visual classification on VGGSound (top-1, end-to-end fine-tuning). Published baselines all use separate audio and video encoders, except AVSiam, and are listed above the rule; our AudioSet-pretrained shared-encoder ViT-B (87M) and ViT-L (304M) below it. Best result per block in bold.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1: LeAVJEPA early-fusion encoder. Video tubelets and mel-spectrogram patches are embedded, summed with factorized positional and modality-type embeddings, concatenated with a [CLS] token, and processed by a 12 -layer ViT-Base. The same encoder is used for every view, including the modality-dropout local views, where the absent modality contributes no tokens and the sequence is correspondingly shorter.
Setting
ViT-B
ViT-L
Data
Dataset
AudioSet-2M (unbalanced), ∼ 1.9M clips
Clip length
8 s (from a 10 s clip, random offset)
Architecture
Encoder
12L, 768d, 12h
24L, 1024d, 16h
Parameters
87M
304M
Appendix
Table A1: AudioSet-2M pretraining configurations. Settings spanning both columns are shared.
Setting
ViT-B
ViT-L
Probes on top of pretrained backbone
Linear classifier
LayerNorm + linear on [CLS] (309 classes)
Attentive classifier
1 query, 12-head cross-attention
1 query, 16-head cross-attention
Optimization
Optimizer
AdamW
Head learning rate
2×10−4
Appendix
Table A2: VGGSound fine-tuning configuration for the AudioSet-pretrained ViT-B and ViT-L backbones. Settings spanning both columns are shared.
Figure A2: Embedding standard deviation over training. SIGReg drives the per-dimension std toward 1.0 (the isotropic Gaussian target).
Figure A3: Pretraining loss components (VGGSound-only run, 50 epochs).
Figure A4: Online probing curves on frozen backbone features during pretraining.
Figure A5: AudioSet pretraining loss components over 57 epochs of AudioSet-2M.
Figure A6: Fine-tuning curves for end-to-end fine-tuning of the pretrained LeAVJEPA backbone.
Figure A7: Semantic structure of the embedding space. t-SNE of the [CLS] embeddings of the fine-tuned LeAVJEPA ViT-B backbone for N=11,143 VGGSound training clips drawn from 60 classes grouped into six semantic families (colours). Clips cluster by family with finer per-class sub-structure of the same colour; the residual overlap falls mainly between acoustically related families (animal calls vs. human voice).
Figure A8: Feature PCA of video patch tokens for two VGGSound instrument classes. A single PCA is fit jointly over the last-layer video patch tokens of four clips per class (rows); its top three components are mapped to RGB and overlaid on four frames per clip (columns). The instrument takes a consistent colour across clips and frames, distinct from the player and the background.
Pretraining views
Linear
Attentive
Invariance
Embed. std
Audio-only and video-only local (LeAVJEPA)
43.9 / 72.7
47.1 / 74.9
0.042
0.989
Random modality dropout ( p=0.5 )
38.1 / 67.7
41.3 / 71.2
0.060
0.966
Both modalities + masking
23.1 / 49.2
34.1 / 63.1
0.018
0.982
Both modalities in local views
13.7 / 33.5
24.7 / 51.2
0.004
0.984
No local views ( K=0 )
10.3 / 26.9
23.2 / 48.2
0.003
0.985
Audio-only pretraining
14.1 / 35.1
22.5 / 48.5
0.005
0.984
Appendix
Table A3: View ablation on VGGSound at epoch 30. Probe entries are online top-1/top-5 test accuracy; results are single runs.
Figure A9: Architectural ablation: shared vs. dual encoder on VGGSound pretraining. Linear-/attentive-probe top-1 accuracy, embedding standard deviation, invariance loss, SIGReg loss, and total LeJEPA loss are plotted against pretraining epoch. Curves are clipped to the shorter of the two runs (53 epochs).
Figure A10: LeJEPA hyperparameter sensitivity: SIGReg weight λ on VGGSound pretraining, sweeping λ∈{0.03,0.05,0.10} with λ=0.05 as the baseline used elsewhere in the paper. Probe accuracies are from the online probes on the training split. λ=0.03 and 0.05 behave almost identically, whereas λ=0.10 leaves both probes far lower and the invariance loss far higher. Curves are clipped to the shortest run ( λ=0.10 , 21 epochs).
Figure A11: LeJEPA hyperparameter sensitivity: number of local views K on VGGSound pretraining, sweeping K∈{2,4,6,8} with K=2 as the paper’s default. Curves are clipped to the shortest run ( K=8 , 7 epochs).
Figure A12: Removing SIGReg ( λ=0 ) vs. the baseline ( λ=0.05 ) on VGGSound pretraining. Without the SIGReg term, the invariance loss collapses to zero and the embedding standard deviation drops several orders of magnitude below the baseline, indicating representation collapse.
Figure A13: Removing the invariance loss ( λ=1 ) vs. the baseline ( λ=0.05 ) on VGGSound pretraining. SIGReg alone drives the embedding standard deviation to its target, but the invariance loss never decreases, indicating that the model never learns to align global and local views without the invariance term.
Figure A14: Removing modality dropout and tube/freq-time masking vs. the baseline on VGGSound pretraining. Without partial-view perturbations the invariance loss collapses to near zero, indicating that alignment between global and local views becomes trivial when both modalities are always present and unmasked.