MacJEPA: Missingness-Robust Audio-Visual Recognition from Untrimmed Egocentric Videos
Authors: Souptik Sen, Zahra Ahmadi
Organizations: Peter L. Reichertz Institute for Medical Informatics, Hannover Medical School, Germany · Lower Saxony Center for Artificial Intelligence and Causal Methods in Medicine (CAIMed), Hannover, Germany
Audio-visual models improve egocentric action recognition by exploiting complementary cues, yet typically assume that both streams remain available at inference. Existing missing-modality methods operate on trimmed, single-event clips in which a stream is entirely present or absent, whereas real sensors fail and recover within long, untrimmed observations. We redefine egocentric modality missingness as temporally localized sensor outages within untrimmed, multi-event observations, with whole-clip absence as the limiting case. We introduce \textbf{MacJEPA}, a missing-modality-robust \textbf{Ma}sked-\textbf{c}ontext query \textbf{JEPA} that recognizes visual actions and acoustic events from supplied interval queries over audio-visual context. Window-local modality dropout simulates these sensor outages during training. MacJEPA further repurposes masking in JEPA from a self-supervised pretext into a supervised robustness objective, aligning masked and clean latent representations of both multimodal content tokens and the task-conditioned queries. All objectives are optimized jointly with recognition in a single stage, requiring no test-time adaptation. Across Epic-Kitchens-100 and Epic-Sounds, a single checkpoint remains competitive under complete input and consistently surpasses published missing-modality baselines when either the dominant or auxiliary stream is removed. MacJEPA thus unifies strong full-input recognition with temporal missing-modality robustness in a single model operating on untrimmed multi-event videos.
Figures & tables
Figure 1: Local sensor outages across overlapping multimodal events in a 19.2-second untrimmed observation.
Figure 2: MacJEPA training framework. Window-local modality dropout corrupts the online multimodal sequence with typed, timed masks, while content and query JEPA align it with clean EMA-teacher targets under joint task supervision. Two windows and one query per stream are shown for clarity.
Ablation
Epic-Kitchens action, V drop
Epic-Sounds sound, A drop
0
25
50
75
100
0
25
50
75
100
Unimodal reference
13.0
13.0
13.0
13.0
13.0
41.4
41.4
41.4
41.4
41.4
H1: Encoder + CE
53.6
51.6
45.7
27.8
4.1
56.5
55.1
53.5
45.2
26.3
H2: H1 + window-local modality dropout
52.9
51.7
48.7
44.0
16.3
55.5
54.4
53.8
51.0
46.7
H3: H2 + content JEPA
53.6
53.1
50.8
45.2
18.2
56.1
55.5
55.0
51.5
48.4
H4: H3 + query JEPA ( MacJEPA )
54.0
53.7
52.2
47.3
19.5
58.1
56.0
54.3
52.1
50.0
Table 1: Controlled component analysis under dominant-modality removal. Top-1 accuracy (%) for Epic-Kitchens action and Epic-Sounds sound reported. Follows nested per-window temporal missingness protocol.
Figure 3: Query attention over visual (left) and audio (right) windows before (above) and after (below) local outages. Hatched regions are missing, boxes compare clean and corrupted predictions against ground truth.
Audio-Visual Benchmarks
xp
Verb
Noun
Action
TBN
224p
66.0
47.2
36.7
MMT
224p
64.0
57.3
42.8
MBT
224p
64.8
58.0
43.4
MTCN
336p
70.7
62.1
49.6
M&M
420p
72.0
66.3
53.6
TIM
224p
77.1
67.2
57.5
Table 2: Full-modality validation top-1 accuracy (%). MacJEPA uses a single saved checkpoint for both Epic-Kitchens and Epic-Sounds. Best results are bold, and second-best results are underlined.
Benchmarks
Missing rate (%)
0
25
50
75
100
Unimodal (A only)
40.0
40.0
40.0
40.0
40.0
MiDl-LTA
63.7
58.4
52.4
46.7
41.4
MMT
63.4
59.0
54.6
50.0
45.3
MacJEPA (ours)
74.0
67.8
61.6
55.6
49.5
Table 3: Dominant-modality robustness. Top-1 accuracy (%) for Epic-Kitchens verb under missing video and Epic-Sounds sound under missing audio. Follows observation-level missingness protocol.
Benchmarks
Missing rate (%)
0
25
50
75
100
Unimodal (V only)
63.2
63.2
63.2
63.2
63.2
MiDl-LTA
63.7
61.7
59.5
57.3
55.3
MacJEPA (ours)
74.0
73.5
73.2
72.9
72.5
Table 4: Auxiliary-modality robustness. Top-1 accuracy (%) for Epic-Kitchens verb under missing audio and Epic-Sounds sound under missing video. Follows observation-level missingness protocol.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Training setting
Epic-Kitchens action, V drop
Epic-Sounds sound, A drop
0
25
50
75
100
0
25
50
75
100
Unimodal reference
13.0
13.0
13.0
13.0
13.0
41.4
41.4
41.4
41.4
41.4
π=0.3
54.1
53.9
53.6
49.7
13.7
57.4
56.7
55.4
53.1
47.8
π=0.5
53.9
53.9
53.2
48.8
16.9
57.7
57.2
55.9
54.1
50.1
π=0.9
53.6
53.2
51.8
46.5
21.5
58.1
57.2
55.8
53.6
50.1
π=0.7 (MacJEPA)
54.0
53.7
52.2
47.3
19.5
58.1
56.0
54.3
52.1
50.0
Appendix
Table 5: Effect of the training-time modality-drop probability π . We report top-1 accuracy (%) across per-window missing rates for Epic-Kitchens action under video removal and Epic-Sounds sound under audio removal. π=0.7 is the MacJEPA setting. Follows nested per-window temporal missingness protocol.
Window setting
Epic-Kitchens action, V drop
Epic-Sounds sound, A drop
0
25
50
75
100
0
25
50
75
100
Unimodal reference
13.0
13.0
13.0
13.0
13.0
41.4
41.4
41.4
41.4
41.4
w=2 ( 18×1.07 s)
53.6
53.6
53.1
50.9
14.8
57.7
56.9
56.1
54.1
48.6
w=3 ( 12×1.60 s)
53.5
53.5
52.4
48.6
16.9
57.1
56.6
55.7
54.1
49.3
w=6 ( 6×3.20 s)
53.7
53.3
51.4
44.5
22.3
57.9
57.3
55.9
53.7
50.0
w=4 ( 9×2.13 s; MacJEPA)
54.0
53.7
52.2
47.3
19.5
58.1
56.0
54.3
52.1
50.0
Appendix
Table 6: Effect of outage granularity w . We report top-1 accuracy (%) across per-window missing rates for Epic-Kitchens action under video removal and Epic-Sounds sound under audio removal. w=4 is the MacJEPA setting. Follows nested per-window temporal missingness protocol.
Observation setting
Epic-Kitchens action, V drop
Epic-Sounds sound, A drop
0
25
50
75
100
0
25
50
75
100
Unimodal reference
13.0
13.0
13.0
13.0
13.0
41.4
41.4
41.4
41.4
41.4
T=20 ( 10.67 s, K=5 )
53.7
53.6
52.7
47.5
22.0
57.7
57.0
55.9
53.6
49.5
T=28 ( 14.93 s, K=7 )
53.7
53.6
52.9
47.7
19.6
58.1
57.6
56.1
53.9
49.9
T=44 ( 23.47 s, K=11 )
53.6
53.3
52.5
47.7
18.4
56.6
56.3
55.0
53.0
49.2
T=36 ( 19.20 s, K=9 ; MacJEPA)
54.0
53.7
52.2
47.3
19.5
58.1
56.0
54.3
52.1
50.0
Appendix
Table 7: Effect of observation length T . We report top-1 accuracy (%) across per-window missing rates for Epic-Kitchens action under video removal and Epic-Sounds sound under audio removal. T=36 is the MacJEPA setting. Follows nested per-window temporal missingness protocol.
Encoder setting
Epic-Kitchens action, V drop
Epic-Sounds sound, A drop
0
25
50
75
100
0
25
50
75
100
Unimodal reference
13.0
13.0
13.0
13.0
13.0
41.4
41.4
41.4
41.4
41.4
Lenc=2 (21.2M)
53.4
53.2
52.1
46.7
18.2
57.6
57.0
56.2
54.0
48.8
Lenc=4 (38.0M)
53.6
53.3
51.9
47.5
19.3
57.2
56.6
55.4
53.6
50.0
Lenc=8 (71.6M)
53.7
53.5
52.4
47.9
18.3
57.8
57.0
56.0
53.8
50.5
Lenc=6 (54.8M; MacJEPA)
54.0
53.7
52.2
47.3
19.5
58.1
56.0
54.3
52.1
50.0
Appendix
Table 8: Effect of encoder depth Lenc . We report top-1 accuracy (%) across per-window missing rates for Epic-Kitchens action under video removal and Epic-Sounds sound under audio removal. Total parameter counts are shown in parentheses. Follows nested per-window temporal missingness protocol.
Visual features
Video missing rate (%)
0
25
50
75
100
Unimodal reference (A only)
13.0
13.0
13.0
13.0
13.0
Omnivore only ( Dv=1024 )
52.7
52.4
51.2
45.3
19.6
VideoMAE only ( Dv=1024 )
52.5
51.7
49.6
43.4
18.3
Omnivore + VideoMAE ( Dv=2048 ; MacJEPA)
54.0
53.7
52.2
47.3
19.5
Appendix
Table 9: Robustness across frozen visual feature sources. We report Epic-Kitchens action top-1 accuracy (%) under per-window video removal. Omnivore and VideoMAE are concatenated in the MacJEPA setting. Follows nested per-window temporal missingness protocol.