MacJEPA: Missingness-Robust Audio-Visual Recognition from Untrimmed Egocentric Videos
Authors: Souptik Sen, Zahra Ahmadi
Organizations: Peter L. Reichertz Institute for Medical Informatics, Hannover Medical School, Germany · Lower Saxony Center for Artificial Intelligence and Causal Methods in Medicine (CAIMed), Hannover, Germany
Audio-visual models improve egocentric action recognition by exploiting complementary cues, yet typically assume that both streams remain available at inference. Existing missing-modality methods operate on trimmed, single-event clips in which a stream is entirely present or absent, whereas real sensors fail and recover within long, untrimmed observations. We redefine egocentric modality missingness as temporally localized sensor outages within untrimmed, multi-event observations, with whole-clip absence as the limiting case. We introduce \textbf{MacJEPA}, a missing-modality-robust \textbf{Ma}sked-\textbf{c}ontext query \textbf{JEPA} that recognizes visual actions and acoustic events from supplied interval queries over audio-visual context. Window-local modality dropout simulates these sensor outages during training. MacJEPA further repurposes masking in JEPA from a self-supervised pretext into a supervised robustness objective, aligning masked and clean latent representations of both multimodal content tokens and the task-conditioned queries. All objectives are optimized jointly with recognition in a single stage, requiring no test-time adaptation. Across Epic-Kitchens-100 and Epic-Sounds, a single checkpoint remains competitive under complete input and consistently surpasses published missing-modality baselines when either the dominant or auxiliary stream is removed. MacJEPA thus unifies strong full-input recognition with temporal missing-modality robustness in a single model operating on untrimmed multi-event videos.
Figures & tables
Figure 1: Local sensor outages across overlapping multimodal events in a 19.2-second untrimmed observation.
Figure 2: MacJEPA training framework. Window-local modality dropout corrupts the online multimodal sequence with typed, timed masks, while content and query JEPA align it with clean EMA-teacher targets under joint task supervision. Two windows and one query per stream are shown for clarity.
Ablation
Epic-Kitchens action, V drop
Epic-Sounds sound, A drop
0
25
50
75
100
0
25
50
75
100
Unimodal reference
13.0
13.0
13.0
13.0
13.0
41.4
41.4
41.4
41.4
41.4
H1: Encoder + CE
53.6
51.6
45.7
27.8
4.1
56.5
55.1
53.5
45.2
26.3
H2: H1 + window-local modality dropout
52.9
51.7
48.7
44.0
16.3
55.5
54.4
53.8
51.0
46.7
H3: H2 + content JEPA
53.6
53.1
50.8
45.2
18.2
56.1
55.5
55.0
51.5
48.4
H4: H3 + query JEPA ( MacJEPA )
54.0
53.7
52.2
47.3
19.5
58.1
56.0
54.3
52.1
50.0
Table 1: Controlled component analysis under dominant-modality removal. Top-1 accuracy (%) for Epic-Kitchens action and Epic-Sounds sound reported. Follows nested per-window temporal missingness protocol.
Figure 3: Query attention over visual (left) and audio (right) windows before (above) and after (below) local outages. Hatched regions are missing, boxes compare clean and corrupted predictions against ground truth.
Audio-Visual Benchmarks
xp
Verb
Noun
Action
TBN
224p
66.0
47.2
36.7
MMT
224p
64.0
57.3
42.8
MBT
224p
64.8
58.0
43.4
MTCN
336p
70.7
62.1
49.6
M&M
420p
72.0
66.3
53.6
TIM
224p
77.1
67.2
57.5
Table 2: Full-modality validation top-1 accuracy (%). MacJEPA uses a single saved checkpoint for both Epic-Kitchens and Epic-Sounds. Best results are bold, and second-best results are underlined.
Benchmarks
Missing rate (%)
0
25
50
75
100
Unimodal (A only)
40.0
40.0
40.0
40.0
40.0
MiDl-LTA
63.7
58.4
52.4
46.7
41.4
MMT
63.4
59.0
54.6
50.0
45.3
MacJEPA (ours)
74.0
67.8
61.6
55.6
49.5
Table 3: Dominant-modality robustness. Top-1 accuracy (%) for Epic-Kitchens verb under missing video and Epic-Sounds sound under missing audio. Follows observation-level missingness protocol.
Benchmarks
Missing rate (%)
0
25
50
75
100
Unimodal (V only)
63.2
63.2
63.2
63.2
63.2
MiDl-LTA
63.7
61.7
59.5
57.3
55.3
MacJEPA (ours)
74.0
73.5
73.2
72.9
72.5
Table 4: Auxiliary-modality robustness. Top-1 accuracy (%) for Epic-Kitchens verb under missing audio and Epic-Sounds sound under missing video. Follows observation-level missingness protocol.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Training setting
Epic-Kitchens action, V drop
Epic-Sounds sound, A drop
0
25
50
75
100
0
25
50
75
100
Unimodal reference
13.0
13.0
13.0
13.0
13.0
41.4
41.4
41.4
41.4
41.4
π=0.3
54.1
53.9
53.6
49.7
13.7
57.4
56.7
55.4
53.1
47.8
π=0.5
53.9
53.9
53.2
48.8
16.9
57.7
57.2
55.9
54.1
50.1
π=0.9
53.6
53.2
51.8
46.5
21.5
58.1
57.2
55.8
53.6
50.1
π=0.7 (MacJEPA)
54.0
53.7
52.2
47.3
19.5
58.1
56.0
54.3
52.1
50.0
Appendix
Table 5: Effect of the training-time modality-drop probability π . We report top-1 accuracy (%) across per-window missing rates for Epic-Kitchens action under video removal and Epic-Sounds sound under audio removal. π=0.7 is the MacJEPA setting. Follows nested per-window temporal missingness protocol.
Window setting
Epic-Kitchens action, V drop
Epic-Sounds sound, A drop
0
25
50
75
100
0
25
50
75
100
Unimodal reference
13.0
13.0
13.0
13.0
13.0
41.4
41.4
41.4
41.4
41.4
w=2 ( 18×1.07 s)
53.6
53.6
53.1
50.9
14.8
57.7
56.9
56.1
54.1
48.6
w=3 ( 12×1.60 s)
53.5
53.5
52.4
48.6
16.9
57.1
56.6
55.7
54.1
49.3
w=6 ( 6×3.20 s)
53.7
53.3
51.4
44.5
22.3
57.9
57.3
55.9
53.7
50.0
w=4 ( 9×2.13 s; MacJEPA)
54.0
53.7
52.2
47.3
19.5
58.1
56.0
54.3
52.1
50.0
Appendix
Table 6: Effect of outage granularity w . We report top-1 accuracy (%) across per-window missing rates for Epic-Kitchens action under video removal and Epic-Sounds sound under audio removal. w=4 is the MacJEPA setting. Follows nested per-window temporal missingness protocol.
Observation setting
Epic-Kitchens action, V drop
Epic-Sounds sound, A drop
0
25
50
75
100
0
25
50
75
100
Unimodal reference
13.0
13.0
13.0
13.0
13.0
41.4
41.4
41.4
41.4
41.4
T=20 ( 10.67 s, K=5 )
53.7
53.6
52.7
47.5
22.0
57.7
57.0
55.9
53.6
49.5
T=28 ( 14.93 s, K=7 )
53.7
53.6
52.9
47.7
19.6
58.1
57.6
56.1
53.9
49.9
T=44 ( 23.47 s, K=11 )
53.6
53.3
52.5
47.7
18.4
56.6
56.3
55.0
53.0
49.2
T=36 ( 19.20 s, K=9 ; MacJEPA)
54.0
53.7
52.2
47.3
19.5
58.1
56.0
54.3
52.1
50.0
Appendix
Table 7: Effect of observation length T . We report top-1 accuracy (%) across per-window missing rates for Epic-Kitchens action under video removal and Epic-Sounds sound under audio removal. T=36 is the MacJEPA setting. Follows nested per-window temporal missingness protocol.
Encoder setting
Epic-Kitchens action, V drop
Epic-Sounds sound, A drop
0
25
50
75
100
0
25
50
75
100
Unimodal reference
13.0
13.0
13.0
13.0
13.0
41.4
41.4
41.4
41.4
41.4
Lenc=2 (21.2M)
53.4
53.2
52.1
46.7
18.2
57.6
57.0
56.2
54.0
48.8
Lenc=4 (38.0M)
53.6
53.3
51.9
47.5
19.3
57.2
56.6
55.4
53.6
50.0
Lenc=8 (71.6M)
53.7
53.5
52.4
47.9
18.3
57.8
57.0
56.0
53.8
50.5
Lenc=6 (54.8M; MacJEPA)
54.0
53.7
52.2
47.3
19.5
58.1
56.0
54.3
52.1
50.0
Appendix
Table 8: Effect of encoder depth Lenc . We report top-1 accuracy (%) across per-window missing rates for Epic-Kitchens action under video removal and Epic-Sounds sound under audio removal. Total parameter counts are shown in parentheses. Follows nested per-window temporal missingness protocol.
Visual features
Video missing rate (%)
0
25
50
75
100
Unimodal reference (A only)
13.0
13.0
13.0
13.0
13.0
Omnivore only ( Dv=1024 )
52.7
52.4
51.2
45.3
19.6
VideoMAE only ( Dv=1024 )
52.5
51.7
49.6
43.4
18.3
Omnivore + VideoMAE ( Dv=2048 ; MacJEPA)
54.0
53.7
52.2
47.3
19.5
Appendix
Table 9: Robustness across frozen visual feature sources. We report Epic-Kitchens action top-1 accuracy (%) under per-window video removal. Omnivore and VideoMAE are concatenated in the MacJEPA setting. Follows nested per-window temporal missingness protocol.
Self-supervised learning from large-scale video data has emerged as a dominant paradigm for visual representation learning. Since audio and visual streams naturally co-occur in video data, extending this success to jointly learn from both modalities is a natural next step, yet it remains challenging. Existing audio-visual self-supervised methods rely on modality-specific encoders and complex combinations of contrastive or reconstruction objectives, limiting cross-modal synergy and scalability. Joint Embedding Predictive Architectures (JEPAs) offer a simple, modality-agnostic alternative, but have to date been applied primarily to individual modalities. We introduce MJEPA, a joint-embedding predictive architecture for audio-visual learning that uses a single, unified encoder for both modalities. Our approach uses only a single predictive objective, applied both within and across modalities. We show that cross-modal prediction is critical: without it, a shared encoder degrades below unimodal baselines; with it, each modality's representation benefits from the other. Our frozen ViT-g model outperforms the best prior frozen baseline by over 6.8 mAP on AudioSet-20K, surpasses fully finetuned models on ESC-50 and FSD50K, and is competitive on video benchmarks despite using 10x less video data.
Prior audio-visual self-supervised learning methods rely on mechanisms such as EMA target encoders, prediction heads, reconstruction decoders, and contrastive losses. We introduce LeAVJEPA, the first audio-visual encoder trained under LeJEPA's collapse-free objective. A single early-fusion Vision Transformer processes audio, video, and joint audio-video inputs. Modality dropout treats a missing modality as another view of the same event, making cross-modal alignment implicit in the objective. The model aligns global embeddings with modality-specific local embeddings, and SIGReg prevents representational collapse. A controlled ablation identifies modality dropout as the key mechanism for audio-visual alignment. Despite the architectural simplicity, LeAVJEPA reaches 36.0 mAP on AudioSet-20K and 91.3% accuracy on ESC-50 under frozen evaluation. After fine-tuning, it reaches 61.1% accuracy on VGGSound, and its embeddings support zero-shot audio-visual retrieval.
Omnimodal large language models enable unified audio video understanding, but long joint token sequences make inference costly, and existing benchmarks do not fully isolate audio visual association in noisy user generated videos. We introduce UGC-AVQA, a public UGC benchmark with 1,000 videos and 4,816 QA pairs, where an audio removal test ensures that benchmark questions require both acoustic and visual evidence. To reduce inference cost, we propose OMAC, a training free plug in compression method that preserves salient visual memory and temporally grounded audio anchors. To further make compact models robust to compressed inputs, we introduce O-MARC, a compression distillation framework for learning with memory compressed multimodal contexts. On Qwen2.5-Omni-3B, O-MARC improves the average score across four benchmarks to 45.8, outperforming full token inference at 44.1 and OmniZip at 41.0. OMAC also keeps inference efficient, reducing latency by 34.6% (1.53× speedup) and memory by 34.7% compared with full token inference.
Peiran Wu, Yunze Liu, Chi-Hao Wu +2
University of Bristol · Memories.ai Research · University of Central Florida