LeAVJEPA: A Minimalist Architecture for Audio-Visual Self-Supervised Learning
Organizations: ELLIS Institute Finland Aalto University
Abstract
Prior audio-visual self-supervised learning methods rely on mechanisms such as EMA target encoders, prediction heads, reconstruction decoders, and contrastive losses. We introduce LeAVJEPA, the first audio-visual encoder trained under LeJEPA's collapse-free objective. A single early-fusion Vision Transformer processes audio, video, and joint audio-video inputs. Modality dropout treats a missing modality as another view of the same event, making cross-modal alignment implicit in the objective. The model aligns global embeddings with modality-specific local embeddings, and SIGReg prevents representational collapse. A controlled ablation identifies modality dropout as the key mechanism for audio-visual alignment. Despite the architectural simplicity, LeAVJEPA reaches 36.0 mAP on AudioSet-20K and 91.3% accuracy on ESC-50 under frozen evaluation. After fine-tuning, it reaches 61.1% accuracy on VGGSound, and its embeddings support zero-shot audio-visual retrieval.
Figures & tables
| Method | Single shared encoder | No raw-input recon. | No contr. negatives | No decoder | No teacher / EMA | No predictor | Training objective |
|---|---|---|---|---|---|---|---|
| AV-MAE ( Georgescu et al., 2023 ) | ✗ | ✗ | ✓ | ✗ | ✓ | ✓ | Recon. |
| CAV-MAE ( Gong et al., 2023 ) | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ | Recon. + contr. |
| MAViL ( Huang et al., 2023 ) | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | Raw/latent recon. + contr. |
| CAV-MAE Sync ( Araujo et al., 2025 ) | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ | Recon. + contr. |
| MJEPA ( Teotia et al., 2026 ) | ✓ | ✓ | ✓ | ✓ | ✗ | ✗ | Latent pred. |
| LeAVJEPA (ours) | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | Invariance + SIGReg |
| AudioSet-20K (mAP) | ESC-50 | ||||||
| Method | Params | Pre-train data | A | V | A-V | Acc. | |
| Separate encoders | CAV-MAE ( Gong et al., 2023 ) | 170M | IN+AS | 19.38 | 18.14 | 34.59 | 77.5 |
| MAViL ( Huang et al., 2023 ) | 170M | IN+AS | 30.00 | – | – | 90.8 | |
| EquiAV ( Kim et al., 2024 ) | 170M | IN+AS | 34.25 | 18.60 | 38.60 | 93.2 | |
| CAV-MAE Sync ( Araujo et al., 2025 ) | 170M | IN+AS | 21.66 | 16.20 | 28.50 | 89.2 | |
| XKD ( Sarkar & Etemad, 2024 ) | 170M | AS | – | – | – | 93.6 | |
| Audio video | Video audio | |||||
| Method | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 |
| CAV-MAE ( Gong et al., 2023 ) | 15.1 | 34.0 | 43.0 | 18.8 | 39.5 | 50.1 |
| CAV-MAE Sync ( Araujo et al., 2025 ) | 27.9 | 52.4 | 62.2 | 35.2 | 58.3 | 67.6 |
| EquiAV ( Kim et al., 2024 ) | 29.6 | 53.7 | 63.1 | 30.1 | 53.3 | 62.9 |
| MJEPA ViT-L † ( Teotia et al., 2026 ) | 29.3 | 53.7 | 61.8 | 30.6 | 53.4 | 61.3 |
| LeAVJEPA ViT-B (ours) | 10.5 | 28.9 | 36.6 | 11.4 | 28.2 | 36.8 |
| Method | Params | Init. | Pretraining | Top-1 |
| AV-MAE ( Georgescu et al., 2023 ) | 172M | – | VGGS | 64.2 |
| AV-MAE ( Georgescu et al., 2023 ) | 611M | – | VGGS | 65.0 |
| AVSiam ( Lin & Bertasius, 2024 ) | 100M | IN-21K SL | AS-2M | 64.9 |
| AVSiam ( Lin & Bertasius, 2024 ) | 332M | IN-21K SL | AS-2M | 67.1 |
| CAV-MAE ( Gong et al., 2023 ) | 164M | IN-SSL | AS-2M | 65.5 |
| MAViL ( Huang et al., 2023 ) | 172M | IN-SSL | AS-2M | 67.1 |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | ViT-B | ViT-L |
| Data | ||
| Dataset | AudioSet-2M (unbalanced), 1.9M clips | |
| Clip length | 8 s (from a 10 s clip, random offset) | |
| Architecture | ||
| Encoder | 12L, 768d, 12h | 24L, 1024d, 16h |
| Parameters | 87M | 304M |
| Setting | ViT-B | ViT-L |
| Probes on top of pretrained backbone | ||
| Linear classifier | LayerNorm + linear on [CLS] (309 classes) | |
| Attentive classifier | 1 query, 12-head cross-attention | 1 query, 16-head cross-attention |
| Optimization | ||
| Optimizer | AdamW | |
| Head learning rate | ||
| Pretraining views | Linear | Attentive | Invariance | Embed. std |
|---|---|---|---|---|
| Audio-only and video-only local (LeAVJEPA) | 43.9 / 72.7 | 47.1 / 74.9 | 0.042 | 0.989 |
| Random modality dropout ( ) | 38.1 / 67.7 | 41.3 / 71.2 | 0.060 | 0.966 |
| Both modalities + masking | 23.1 / 49.2 | 34.1 / 63.1 | 0.018 | 0.982 |
| Both modalities in local views | 13.7 / 33.5 | 24.7 / 51.2 | 0.004 | 0.984 |
| No local views ( ) | 10.3 / 26.9 | 23.2 / 48.2 | 0.003 | 0.985 |
| Audio-only pretraining | 14.1 / 35.1 | 22.5 / 48.5 | 0.005 | 0.984 |