Multimodal systems often encode every available input, even when a subset suffices for prediction. Adaptive acquisition can reduce this cost by using predictions from incrementally fused evidence to decide which modality to encode next and when to stop. However, sequential fusion makes these predictions order-dependent, so decisions based on them may need to distinguish factorially many histories of the same acquired set. We introduce SemARC, which couples a Sequential Modality Aggregator (SeMA) with an Adaptive Runtime Controller (ARC) and uses acquired evidence to select each modality before its encoder runs. SeMA executes only selected encoder and fusion branches, updates a fixed-size state, and predicts after each acquisition without recomputing earlier branches. We supervise every acquisition prefix under randomized modality subsets and orders to encourage consistent predictions across acquisition orders. ARC combines a set-dependent marginal-utility prior with residual fitted-Q learning to select the next available modality or stop, without inspecting unacquired inputs or retaining acquisition order. Across six multimodal classification datasets and eleven baselines, SemARC achieves 3.2% higher macro-F1 and 61.4% lower total inference GFLOPs on average relative to each dataset's most accurate baseline. End-to-end latency falls by 44.0% across GPU and CPU and by 47.2% on Android INT8 relative to the fastest measured baseline, on average. Under varying runtime modality missingness, SemARC still skips available modalities, matching or exceeding the best baseline macro-F1 in 21 of 24 conditions with 14.8% lower total GFLOPs on average. SemARC thus offers a practical path toward efficient multimodal inference across heterogeneous devices.
Figures & tables
Figure 1: Select before processing to avoid its cost. (a) Monolithic fusion processes every modality. (b) Post-encoding selection pays for unused representations. (c) SeMARC couples ARC and SeMA to acquire and incrementally fuse modalities or stop, executing only selected branches. (d) On Android (INT8), SeMARC achieves lower latency and energy and higher F1 (annotated) than the shown baselines.
Figure 2: Modality contributions vary across samples. Class-wise Shapley contributions on IEMOCAP motivate sample-adaptive acquisition.
Figure 3: SeMARC training. (a) SeMA learns prefix predictions ps under modality dropout and randomized orders, with classification supervision and same-forward distillation from detached pn , the prediction after all n retained modalities. (b) 1. Group frozen-backbone ordered-prefix predictions by acquired set S to estimate class-conditional marginal gains TS . 2. Combine the fixed cost-aware prior Q0 and trainable residual uθ to score acquisition and stopping actions using acquired evidence and availability. 3. Fit cached-action values to Double-DQN targets, updating only the residual; STOP uses its terminal reward.
Figure 4: Macro-F1 deviation from the mean across acquisition orders of each fixed modality subset.
Type
Method
IEMOCAP
MM-Fi
CMI
CZU-MHAD
EAV
UTD-MHAD
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
Monolithic
MAESTRO
.51 ±.04
6696 (.23)
.18 ±.09
18098 (.28)
.47 ±.02
0.28 (.28)
.25 ±.16
6091 (.30)
.53 ±.02
6230 (.15)
.70 ±.02
7306 (.19)
ShaSpec
.61 ±.03
6697 (1.51)
.66 ±.08
18099 (1.68)
.53 ±.01
2.14 (2.14)
.59 ±.06
6093 (2.17)
.60 ±.05
6231 (.97)
.71 ±.13
7307 (1.29)
MBT
.64 ±.02
6696 (.41)
.77 ±.19
18098 (.46)
.55 ±.01
0.55 (.55)
.82 ±.11
6091 (.56)
.61 ±.03
6230 (.26)
.81 ±.06
7306 (.35)
DecAlign
.48 ±.01
6707 (11.41)
.28 ±.19
18106 (8.45)
.56 ±.02
23.33 (23.33)
.36 ±.10
6114 (23.37)
.62 ±.02
6232 (1.77)
.75 ±.08
7310 (4.21)
Fixed-seq.
MultiModN
.58 ±.06
6696 (.76)
.40 ±.06
18098 (.84)
.45 ±.01
1.07 (1.07)
.47 ±.09
6092 (1.08)
.65 ±.04
6230 (.49)
.66 ±.10
7306 (.64)
Table 1: Macro-F1 and mean total GFLOPs (GF) over three folds. Parentheses show post-encoder fusion cost ( GFf ). Bold and underlining mark the best and second-best values per metric. Complete results: Appendix E.1 .
Figure 5: Energy–latency trade-offs on IEMOCAP across GPU, CPU, and Android INT8. Lower-left is preferred.
Figure 6: IEMOCAP macro-F1 versus post-extraction GFLOPs ( GFf ) under 0–40% test-time modality missingness. Upper-left is preferred.
Configuration
F1 ↑
t (ms) ↓
SeMARC
.77
520.27
MultiModN + ARC
.53
1078.90
SeMA + EAMA
.74
463.50
SeMA + Joint-AFA
.75
920.32
SeMA + GMS
.73
772.08
Table 2: Mean macro-F1 and GPU latency across six datasets.
Figure 7: Relative macro-F1 changes for selected ablations on IEMOCAP under full modality availability. Complete results: Appendix I .
Figure 10
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Seed
E
Ew/Er
pmax
LR
Batch
Le/Lf
T
Folds
IEMOCAP
239
40
10/20
0.4
10−4
32
3/3
100
0–2
MM-Fi
239
100
20/40
0.2
10−4
16
2/4
128
0–2
CMI
239
40
10/20
0.4
10−4
32
3/3
128
0–2
CZU-MHAD
239
100
20/40
0.2
10−4
32
3/3
128
0–2
EAV
239
100
20/40
0.2
10−4
32
3/3
128
0–2
UTD-MHAD
9
100
20/40
0.4
10−4
32
3/3
128
0–2
Appendix
Table 3: SeMA training configuration per dataset ( E : epochs; Ew/Er : modality-dropout warm-up and ramp epochs; pmax : final modality-dropout probability; Le/Lf : encoder and fusion blocks per branch; T : sequence length).
Dataset
M
Fold
Seed
λtrain
IEMOCAP
6
0 / 1 / 2
1 / 1 / 0
0.5 / 0.5 / 0.5
MM-Fi
5
0 / 1 / 2
2 / 0 / 0
0.5 / 0.5 / 1.0
CMI
7
0 / 1 / 2
1 / 2 / 0
0.5 / 0.5 / 0.5
CZU-MHAD
7
0 / 1 / 2
0 / 1 / 2
0.5 / 0.5 / 0.5
EAV
3
0 / 1 / 2
2 / 1 / 2
0.5 / 0.5 / 0.5
UTD-MHAD
4
3 / 4 / 5
0 / 1 / 2
0.5 / 0.5 / 0.5
Appendix
Table 4: Per-fold settings of the deployed uniform-cost ARC . The prior is priced at λtrain during fitting and at inference on every fold.
Modality
Sampling rate
Variates
video (DINOv3 ViT-L/16)
29.97 fps, 720 × 480
1024
audio (XLSR-1B)
16 kHz, mono
1280
text (BERT-base)
–
768
mocap_hand
120 Hz
18
mocap_head
120 Hz
6
mocap_rotated
120 Hz
165
Appendix
Table 5: IEMOCAP modalities (left) and cross-session splits (right). Output classes: neutral, angry, happy, sad.
TI IWR6843AOP, 30 Hz; N×5 (x,y,z,Doppler,intensity)
128
lidar (PointNet)
Ouster OS1-32, 20 Hz; N×3
128
depth (DINOv3)
RealSense D435, 30 fps, 640 × 480 16-bit
1024
Appendix
Table 6: MM-Fi modalities and cross-subject splits. All modalities share the 10 Hz unified frame index (297 frames per 30 s sequence). Output classes: 27 actions (stretching, chest expansion, twists, mark time, limb extension, lunges, squat, raising and waving hands, picking up, throwing, kicking, body extension, jumping, bowing).
Modality
Input / rate
Variates
depth (DINOv3)
100 sampled frames
1024
skeleton
25 joints
100
sensor_head
≈ 338 Hz, IMUs 0–1
14
sensor_torso
≈ 338 Hz, IMUs 2–3
14
sensor_rarm
≈ 338 Hz, IMUs 4–5
14
sensor_larm
≈ 338 Hz, IMUs 6–7
14
Appendix
Table 7: CZU-MHAD input representations (left) and cross-subject splits (right). Each paired-IMU input has 2×(3+3+1)=14 features, including timestamps; all streams use T=128 . Output classes: right/left high wave, right/left horizontal wave, hammer, grasp, draw fork (right/left), draw circle (right/left), forward kick (right/left), side kick (right/left), clap, bend down, wave up and down, sur place, body turn (left/right), lateral movement (left/right).
Table 9: EAV: three independent cross-subject splits. Validation and test IDs are listed explicitly; the remaining subjects among IDs 1–42 form the training set. Each split contains 29/6/7 train/validation/test subjects. Test subjects can recur across splits.
Modality
Sensor
Variates
imu
3-axis accelerometer + quaternion
7
thm
5 thermopiles
5
tof_1 – tof_5
five 8×8 time-of-flight grids
64 each
Appendix
Table 10: CMI modalities (left) and cross-subject splits over participant groups A, B, C of 27 each (right). Output classes: above ear pull hair, cheek pinch skin, eyebrow pull hair, eyelash pull hair, forehead pull hairline, forehead scratch, neck pinch skin, neck scratch (BFRB-like); drink from bottle/cup, feel around in tray, glasses on/off, pinch knee/leg skin, pull air toward face, scratch knee/leg, text on phone, wave hello, write name in air, write name on leg.
ID
Gesture
Type
Total
Share (%)
Fold 0
Fold 1
Fold 2
0
Above ear – pull hair
BFRB
638
7.8
216
210
212
1
Cheek – pinch skin
BFRB
637
7.8
215
210
212
2
Drink from bottle/cup
non-BFRB
161
2.0
54
53
54
3
Eyebrow – pull hair
BFRB
638
7.8
216
210
212
4
Eyelash – pull hair
BFRB
640
7.9
216
212
212
5
Feel around in tray and pull out an object
non-BFRB
161
2.0
54
53
54
Appendix
Table 11: CMI class distribution after deduplication by sequence ID. BFRB denotes body-focused repetitive behaviour. IDs follow lexicographic gesture-name order. The three test-fold counts sum to the full class count; percentages are rounded independently.
Modality
Sampling rate
Variates
skeleton
Kinect, 20 joints × (x,y,z)
60
inertial
50 Hz, 3-axis accel + 3-axis gyro
6
depth (DINOv3)
Kinect, 320 × 240
1024
rgb (DINOv3)
15 fps, 640 × 480
1024
Appendix
Table 12: UTD-MHAD modalities (left) and 4/2/2 cross-subject splits (right). Output classes: swipe left, swipe right, wave, clap, throw, arm cross, basketball shoot, draw X, draw circle (CW), draw circle (CCW), draw triangle, bowling, boxing, baseball swing, tennis swing, arm curl, tennis serve, push, knock, catch, pickup and throw, jog, walk, sit to stand, stand to sit, lunge, squat.
Type
Method
Fold 0
Fold 1
Fold 2
F1 ↑
MU ↓
GF(GFf)↓
F1 ↑
MU ↓
GF(GFf)↓
F1 ↑
MU ↓
GF(GFf)↓
Monolithic
MAESTRO
.472
1.0
6696 (.23)
.559
1.0
6696 (.23)
.496
1.0
6696 (.23)
ShaSpec
.623
1.0
6697 (1.51)
.636
1.0
6697 (1.51)
.584
1.0
6697 (1.51)
MBT
.656
1.0
6696 (.41)
.627
1.0
6696 (.41)
.626
1.0
6696 (.41)
DecAlign
.470
1.0
6707 (11.41)
.484
1.0
6707 (11.41)
.486
1.0
6707 (11.41)
Fixed-seq.
MultiModN
.514
1.0
6696 (.76)
.602
1.0
6696 (.76)
.632
1.0
6696 (.76)
Appendix
Table 13: IEMOCAP , per fold, full modality availability: macro-F1 ( ↑ ), modality usage MU ( ↓ ) and end-to-end GFLOPs ( ↓ ; fusion cost GFf in parentheses). Fold means and sample standard deviations of these values are the cells of Table 1 ; F1 is shown to three decimals. Per column, bold and underlining mark the best and second-best distinct displayed values (F1 highest; MU, GF and GFf lowest, GF compared on the displayed total); ties share markings.
Type
Method
Fold 0
Fold 1
Fold 2
F1 ↑
MU ↓
GF(GFf)↓
F1 ↑
MU ↓
GF(GFf)↓
F1 ↑
MU ↓
GF(GFf)↓
Monolithic
MAESTRO
.232
1.0
18098 (.28)
.229
1.0
18098 (.28)
.067
1.0
18098 (.28)
ShaSpec
.713
1.0
18099 (1.68)
.691
1.0
18099 (1.68)
.570
1.0
18099 (1.68)
MBT
.824
1.0
18098 (.46)
.929
1.0
18098 (.46)
.557
1.0
18098 (.46)
DecAlign
.120
1.0
18106 (8.45)
.492
1.0
18106 (8.45)
.236
1.0
18106 (8.45)
Fixed-seq.
MultiModN
.429
1.0
18098 (.84)
.437
1.0
18098 (.84)
.323
1.0
18098 (.84)
Appendix
Table 14: MM-Fi , per fold, full modality availability: macro-F1 ( ↑ ), modality usage MU ( ↓ ) and end-to-end GFLOPs ( ↓ ; fusion cost GFf in parentheses). Fold means and sample standard deviations of these values are the cells of Table 1 ; F1 is shown to three decimals. Per column, bold and underlining mark the best and second-best distinct displayed values (F1 highest; MU, GF and GFf lowest, GF compared on the displayed total); ties share markings.
Type
Method
Fold 0
Fold 1
Fold 2
F1 ↑
MU ↓
GF(GFf)↓
F1 ↑
MU ↓
GF(GFf)↓
F1 ↑
MU ↓
GF(GFf)↓
Monolithic
MAESTRO
.456
1.0
0.28 (.28)
.487
1.0
0.28 (.28)
.464
1.0
0.28 (.28)
ShaSpec
.544
1.0
2.14 (2.14)
.516
1.0
2.14 (2.14)
.539
1.0
2.14 (2.14)
MBT
.561
1.0
0.55 (.55)
.543
1.0
0.55 (.55)
.546
1.0
0.55 (.55)
DecAlign
.565
1.0
23.33 (23.33)
.543
1.0
23.33 (23.33)
.575
1.0
23.33 (23.33)
Fixed-seq.
MultiModN
.445
1.0
1.07 (1.07)
.452
1.0
1.07 (1.07)
.441
1.0
1.07 (1.07)
Appendix
Table 15: CMI , per fold, full modality availability: macro-F1 ( ↑ ), modality usage MU ( ↓ ) and end-to-end GFLOPs ( ↓ ; fusion cost GFf in parentheses). Fold means and sample standard deviations of these values are the cells of Table 1 ; F1 is shown to three decimals. Per column, bold and underlining mark the best and second-best distinct displayed values (F1 highest; MU, GF and GFf lowest, GF compared on the displayed total); ties share markings.
Type
Method
Fold 0
Fold 1
Fold 2
F1 ↑
MU ↓
GF(GFf)↓
F1 ↑
MU ↓
GF(GFf)↓
F1 ↑
MU ↓
GF(GFf)↓
Monolithic
MAESTRO
.182
1.0
6091 (.30)
.436
1.0
6091 (.30)
.135
1.0
6091 (.30)
ShaSpec
.645
1.0
6093 (2.17)
.587
1.0
6093 (2.17)
.524
1.0
6093 (2.17)
MBT
.910
1.0
6091 (.56)
.849
1.0
6091 (.56)
.703
1.0
6091 (.56)
DecAlign
.352
1.0
6114 (23.37)
.469
1.0
6114 (23.37)
.265
1.0
6114 (23.37)
Fixed-seq.
MultiModN
.563
1.0
6092 (1.08)
.459
1.0
6092 (1.08)
.379
1.0
6092 (1.08)
Appendix
Table 16: CZU-MHAD , per fold, full modality availability: macro-F1 ( ↑ ), modality usage MU ( ↓ ) and end-to-end GFLOPs ( ↓ ; fusion cost GFf in parentheses). Fold means and sample standard deviations of these values are the cells of Table 1 ; F1 is shown to three decimals. Per column, bold and underlining mark the best and second-best distinct displayed values (F1 highest; MU, GF and GFf lowest, GF compared on the displayed total); ties share markings.
Type
Method
Fold 0
Fold 1
Fold 2
F1 ↑
MU ↓
GF(GFf)↓
F1 ↑
MU ↓
GF(GFf)↓
F1 ↑
MU ↓
GF(GFf)↓
Monolithic
MAESTRO
.516
1.0
6230 ( .15 )
.549
1.0
6230 (.15)
.537
1.0
6230 (.15)
ShaSpec
.638
1.0
6231 (.97)
.550
1.0
6231 (.97)
.615
1.0
6231 (.97)
MBT
.639
1.0
6230 (.26)
.588
1.0
6230 (.26)
.604
1.0
6230 (.26)
DecAlign
.634
1.0
6232 (1.77)
.613
1.0
6232 (1.77)
.599
1.0
6232 (1.77)
Fixed-seq.
MultiModN
.692
1.0
6230 (.49)
.610
1.0
6230 (.49)
.642
1.0
6230 (.49)
Appendix
Table 17: EAV , per fold, full modality availability: macro-F1 ( ↑ ), modality usage MU ( ↓ ) and end-to-end GFLOPs ( ↓ ; fusion cost GFf in parentheses). Fold means and sample standard deviations of these values are the cells of Table 1 ; F1 is shown to three decimals. Per column, bold and underlining mark the best and second-best distinct displayed values (F1 highest; MU, GF and GFf lowest, GF compared on the displayed total); ties share markings.
Type
Method
Fold 0
Fold 1
Fold 2
F1 ↑
MU ↓
GF(GFf)↓
F1 ↑
MU ↓
GF(GFf)↓
F1 ↑
MU ↓
GF(GFf)↓
Monolithic
MAESTRO
.683
1.0
7306 ( .19 )
.721
1.0
7306 (.19)
.689
1.0
7306 (.19)
ShaSpec
.809
1.0
7307 (1.29)
.765
1.0
7307 (1.29)
.566
1.0
7307 (1.29)
MBT
.873
1.0
7306 (.35)
.759
1.0
7306 (.35)
.805
1.0
7306 (.35)
DecAlign
.838
1.0
7310 (4.21)
.709
1.0
7310 (4.21)
.692
1.0
7310 (4.21)
Fixed-seq.
MultiModN
.764
1.0
7306 (.64)
.578
1.0
7306 (.64)
.636
1.0
7306 (.64)
Appendix
Table 18: UTD-MHAD , per fold, full modality availability: macro-F1 ( ↑ ), modality usage MU ( ↓ ) and end-to-end GFLOPs ( ↓ ; fusion cost GFf in parentheses). Fold means and sample standard deviations of these values are the cells of Table 1 ; F1 is shown to three decimals. Per column, bold and underlining mark the best and second-best distinct displayed values (F1 highest; MU, GF and GFf lowest, GF compared on the displayed total); ties share markings.
Figure 11: Class-conditional acquisition on the five datasets outside the main text. Each cell is P(j∈S∣y) , the probability that the deployed policy acquires modality j given true class y , at full availability with the test folds pooled; darker is more often acquired. The red strip right of each panel is that class’s mean acquired-set size ∣S∣ on a 0 – M scale. Panels preserve their aspect ratios and fit within their columns; modality counts are M=3 – 7 and class counts are C=5 – 27 . IEMOCAP , drawn the same way, is Figure 10 .
Method
fp32
fp16
INT8 (QAT)
INT8 (PTQ)
SeMA + ARC (Ours)
.696
.697
.696
.696
AdaMML
.623
.624
.624
.619
DynMM
.663
.661
.663
.663
DyMo
.640
.639
.641
.642
Encoder / format
MB
cos (min tok.)
latency vs fp32
DINOv3-L/16 fp32
1213
1.000 (1.000)
1.00 ×
Appendix
Table 19: Quantization pilot on IEMOCAP. Top: cross-subject test accuracy (three folds) for fp32, fp16, and two INT8 settings. QAT is included only as an experimental comparison; the deployed PTQ setting uses no fine-tuning. fp16 uses bf16 for the ViT-L encoder. Bottom: the shared encoders on Android phone with ONNX Runtime: size, output-feature cosine against fp32 on real inputs (min. per-token cosine in parentheses) and latency relative to fp32. W8A8 = plain dynamic INT8; SQ70 = SmoothQuant( α=0.7 ) then the same quantizer, the deployed format.
Method
IEMOCAP
MM-Fi
CMI
CZU-MHAD
EAV
UTD-MHAD
GPU
MBT
1530.22 / 367.68
4234.19/804.09
15.12/2.15
1324.02/260.59
2032.22/321.31
1036.01/279.43
AdaMML
1595.15/385.60
4196.11/ 789.15
14.61 / 1.57
1312.38 /262.83
2011.12 / 316.99
1018.04 / 275.27
DynMM
1599.79/384.79
4194.79 /796.49
17.98/1.64
1317.80/ 259.45
2017.36/319.48
1020.67/275.42
DyMo
1616.70/388.86
4208.37/799.10
35.50/3.75
1341.98/264.15
2028.04/322.96
1031.99/277.82
SeMARC (Ours)
732.23 / 180.65
179.46 / 32.84
14.39 / 1.60
1315.73 / 260.54
224.42 / 36.74
627.78 / 169.99
Appendix
Table 20: Latency and energy across compute platforms. Each entry reports latency (ms)/energy (J); lower is better. Bold and underlining denote the best and second-best values within each dataset and platform, ranked separately for latency and energy at full precision. Latency is measured end to end for every method and platform. SeMARC uses the deployed policy. Server prediction metrics on GPU and CPU use the complete test folds; phone metrics use the same 50 samples per fold for every method. GPU and CPU energy include idle consumption over the timed interval, whereas phone energy subtracts a calibrated idle baseline.
Method
0%
10%
20%
30%
40%
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
MBT
.64 ±.02
6696 (.41)
.61
6026 (.41)
.60 ±.02
5357 (.41)
.57
4688 (.41)
.55 ±.03
4022 (.41)
MAESTRO
.51 ±.04
6696 (.23)
.45
6026 (.23)
.42 ±.02
5357 (.23)
.38
4688 (.23)
.36 ±.03
4022 (.23)
ShaSpec
.61 ±.03
6697 (1.51)
.58
6027 (1.51)
.56 ±.03
5358 (1.51)
.55
4689 (1.51)
.51 ±.03
4023 (1.51)
DecAlign
.48 ±.01
6707 (11.41)
.46
6037 (11.41)
.45 ±.01
5368 (11.41)
.42
4699 (11.41)
.42 ±.01
4033 (11.41)
MultiModN
.58 ±.06
6696 (.76)
.57
6027 (.76)
.55 ±.07
5357 (.76)
.54
4688 (.76)
.51 ±.06
4023 (.76)
Appendix
Table 21: Macro-F1 ( ↑ ; mean ±std where available) and end-to-end GFLOPs ( ↓ ; fusion cost GFf in parentheses) on IEMOCAP under 0–40% test-time modality missingness. Per column, bold and underlining mark the best and second-best distinct displayed values (F1 highest; GF and GFf lowest); ties share markings; the backbone reference row is not ranked. The 0% primary-method entries use the recomputed three-fold summaries.
Method
0%
10%
20%
30%
40%
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
MBT
.77 ±.19
18098 (.46)
.69
16288 (.46)
.63 ±.17
14480 (.46)
.56
12678 (.46)
.52 ±.14
10896 (.46)
MAESTRO
.18 ±.09
18098 (.28)
.17
16288 (.28)
.14 ±.07
14479 (.28)
.14
12677 (.28)
.14 ±.08
10896 (.28)
ShaSpec
.66 ±.08
18099 (1.68)
.58
16289 (1.68)
.56 ±.04
14481 (1.68)
.47
12679 (1.68)
.44 ±.05
10897 (1.68)
DecAlign
.28 ±.19
18106 (8.45)
.24
16296 (8.45)
.24 ±.12
14488 (8.45)
.21
12685 (8.45)
.19 ±.09
10904 (8.45)
MultiModN
.40 ±.06
18098 (.84)
.37
16289 (.84)
.35 ±.04
14480 (.84)
.32
12678 (.84)
.29 ±.07
10896 (.84)
Appendix
Table 22: Macro-F1 ( ↑ ; mean ±std where available) and end-to-end GFLOPs ( ↓ ; fusion cost GFf in parentheses) on MM-Fi under 0–40% test-time modality missingness. Per column, bold and underlining mark the best and second-best distinct displayed values (F1 highest; GF and GFf lowest); ties share markings; the backbone reference row is not ranked. The 0% primary-method entries use the recomputed three-fold summaries.
Method
0%
10%
20%
30%
40%
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
MBT
.55 ±.01
0.55 (.55)
.53
0.55 (.55)
.52 ±.01
0.55 (.55)
.47
0.55 (.55)
.46 ±.02
0.55 (.55)
MAESTRO
.47 ±.02
0.28 (.28)
.42
0.28 (.28)
.37 ±.01
0.28 (.28)
.32
0.28 (.28)
.27 ±.02
0.28 (.28)
ShaSpec
.53 ±.01
2.14 (2.14)
.52
2.14 (2.14)
.50 ±.02
2.14 (2.14)
.48
2.14 (2.14)
.45 ±.01
2.14 (2.14)
DecAlign
.56 ±.02
23.33 (23.33)
.54
23.33 (23.33)
.52 ±.01
23.33 (23.33)
.49
23.33 (23.33)
.45 ±.01
23.33 (23.33)
MultiModN
.45 ±.01
1.07 (1.07)
.44
1.07 (1.07)
.43 ±.01
1.07 (1.07)
.41
1.07 (1.07)
.40 ±.02
1.07 (1.07)
Appendix
Table 23: Macro-F1 ( ↑ ; mean ±std where available) and end-to-end GFLOPs ( ↓ ; fusion cost GFf in parentheses) on CMI under 0–40% test-time modality missingness. Per column, bold and underlining mark the best and second-best distinct displayed values (F1 highest; GF and GFf lowest); ties share markings; the backbone reference row is not ranked. The 0% primary-method entries use the recomputed three-fold summaries.
Method
0%
10%
20%
30%
40%
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
MBT
.82 ±.11
6091 (.56)
.75
5482 (.56)
.69 ±.10
4873 (.56)
.65
4264 (.56)
.57 ±.10
3656 (.56)
MAESTRO
.25 ±.16
6091 (.30)
.22
5482 (.30)
.20 ±.11
4873 (.30)
.17
4264 (.30)
.14 ±.07
3656 (.30)
ShaSpec
.59 ±.06
6093 (2.17)
.54
5484 (2.17)
.49 ±.04
4874 (2.17)
.47
4266 (2.17)
.42 ±.05
3658 (2.17)
DecAlign
.36 ±.10
6114 (23.37)
.34
5505 (23.37)
.29 ±.07
4896 (23.37)
.30
4287 (23.37)
.24 ±.05
3679 (23.37)
MultiModN
.47 ±.09
6092 (1.08)
.44
5482 (1.08)
.41 ±.09
4874 (1.08)
.38
4265 (1.08)
.34 ±.05
3657 (1.08)
Appendix
Table 24: Macro-F1 ( ↑ ; mean ±std where available) and end-to-end GFLOPs ( ↓ ; fusion cost GFf in parentheses) on CZU-MHAD under 0–40% test-time modality missingness. Per column, bold and underlining mark the best and second-best distinct displayed values (F1 highest; GF and GFf lowest); ties share markings; the backbone reference row is not ranked. The 0% primary-method entries use the recomputed three-fold summaries.
Method
0%
10%
20%
30%
40%
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
MBT
.61 ±.03
6230 (.26)
.58
5609 (.26)
.55 ±.02
5001 (.26)
.53
4417 (.26)
.50 ±.01
3871 (.26)
MAESTRO
.53 ±.02
6230 (.15)
.48
5609 (.15)
.45 ±.01
5001 (.15)
.41
4417 (.15)
.36 ±.01
3871 (.15)
ShaSpec
.60 ±.05
6231 (.97)
.58
5610 (.97)
.56 ±.04
5001 (.97)
.52
4418 (.97)
.50 ±.02
3872 (.97)
DecAlign
.62 ±.02
6232 (1.77)
.58
5611 (1.77)
.55 ±.01
5002 (1.77)
.51
4419 (1.77)
.48 ±.01
3872 (1.77)
MultiModN
.65 ±.04
6230 (.49)
.61
5609 (.49)
.59 ±.04
5001 (.49)
.56
4417 (.49)
.54 ±.02
3871 (.49)
Appendix
Table 25: Macro-F1 ( ↑ ; mean ±std where available) and end-to-end GFLOPs ( ↓ ; fusion cost GFf in parentheses) on EAV under 0–40% test-time modality missingness. Per column, bold and underlining mark the best and second-best distinct displayed values (F1 highest; GF and GFf lowest); ties share markings; the backbone reference row is not ranked. The 0% primary-method entries use the recomputed three-fold summaries.
Method
0%
10%
20%
30%
40%
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
F1 ↑
GF(GFf)↓
MBT
.81 ±.06
7306 (.35)
.79 ±.06
6576 (.34)
.73 ±.05
5848 (.34)
.67 ±.06
5129 (.34)
.63 ±.06
4431 (.34)
MAESTRO
.70 ±.02
7306 (.19)
.65 ±.01
6576 (.19)
.57 ±.02
5848 (.19)
.54 ±.02
5129 (.19)
.49 ±.01
4430 (.19)
ShaSpec
.71 ±.13
7307 (1.29)
.66 ±.11
6577 (1.29)
.62 ±.08
5849 (1.29)
.54 ±.08
5130 (1.29)
.47 ±.06
4432 (1.29)
DecAlign
.75 ±.08
7310 (4.21)
.71 ±.09
6580 (4.21)
.62 ±.08
5852 (4.21)
.59 ±.10
5133 (4.21)
.56 ±.06
4434 (4.21)
MultiModN
.66 ±.10
7306 (.64)
.62 ±.08
6576 (.64)
.55 ±.06
5848 (.64)
.46 ±.08
5130 (.64)
.41 ±.07
4431 (.64)
Appendix
Table 26: Macro-F1 ( ↑ ; mean ±std where available) and end-to-end GFLOPs ( ↓ ; fusion cost GFf in parentheses) on UTD-MHAD under 0–40% test-time modality missingness. Per column, bold and underlining mark the best and second-best distinct displayed values (F1 highest; GF and GFf lowest); ties share markings; the backbone reference row is not ranked. The 0% primary-method entries use the recomputed three-fold summaries.
Configuration
IEMOCAP
MM-Fi
CZU-MHAD
EAV
CMI
UTD-MHAD
F1
MU
Latency
F1
MU
Latency
F1
MU
Latency
F1
MU
Latency
F1
MU
Latency
F1
MU
Latency
SeMARC
.669
.65
745
.950
.21
187
.914
.37
1301
.689
.37
217
.556
.75
12.6
.827
.52
659
MultiModN + ARC
.535
.33
199
.377
.20
4004
.485
.16
1294
.647
.39
319
.386
.28
4.4
.734
.40
653
MultiModN
.580
1.00
1530
.397
1.00
4234
.467
1.00
1324
.647
1.00
2032
.447
1.00
15.5
.660
1.00
1036
SeMA + EAMA
.650
.53
338
.945
.22
282
.899
.23
1296
.687
.41
216
.523
.45
6.8
.757
.49
642
SeMA + Joint-AFA
.621
.34
332
.945
.49
2485
.911
.61
1311
.683
.51
603
.517
.62
10.0
.817
.72
780
Appendix
Table 27: Complete component-replacement results: macro-F1 ( ↑ ), modality utilization MU ( ↓ ) and GPU latency t in ms ( ↓ ) per dataset. Table 2 reports the row means. EAMA and Joint-AFA are the pre-backbone adaptations of Appendix C.1 ; EAMA uses uniform costs and the hybrid optimizer. MultiModN F1 and MU are recomputed from the main-table three-fold source. F1 can be compared across backbones on the same task and test set, but only shared-backbone comparisons isolate controller changes.
Category
Variant
SeMA F1
F1 ↑
Δ
GFf↓
MU ↓
SeMA ablations
Training
Fixed modality order
.653
.644 ±.033
−3.69%
.34
.79
No dropout curriculum
.661
.637 ±.043
−4.76%
.30
.70
Fusion
Bottleneck tokens last
.651
.645 ±.018
−3.54%
.29
.67
Learnable aggregation
.675
.670 ±.040
+0.14%
.27
.62
Loss
LCE only
.656
.652 ±.014
−2.46%
.31
.72
Appendix
Table 28: Complete SeMA and ARC ablations on IEMOCAP under full modality availability (mean over three test folds; subscripts are standard deviations over folds). “ SeMA F1” is the retrained backbone alone on all modalities, averaged over 24 orders; ’F1’ is the backbone and policy both retrained; Δ is the relative change of the controlled macro-F1 with respect to .669 . Bold marks the best value per metric.
video
audio
text
MoCap hand
MoCap head
MoCap rot.
tm (ms)
1432.6
64.7
4.1
2.6
1.1
18.4
Relative cost wm(γ)
γ=0
1
1
1
1
1
1
γ=0.25
5.99
2.76
1.39
1.23
1
2.02
γ=0.5
35.9
7.62
1.93
1.52
1
4.07
γ=0.75
215
21.0
2.68
1.88
1
8.20
Appendix
Table 29: Per-modality extraction time tm (idle L40, ms per sample), relative cost wm(γ) of Equation 10 , and the resulting price of acquiring each modality at λ=0.5 with all modalities available ( λwm/∑m′wm′ ) on IEMOCAP.
Acquisition frequency
γ
F1 ↑
MU
GFf↓
Lat. (ms) ↓
Extr.
vid.
aud.
text
hand
head
rot.
0
.669 ±.022
.65
.28
745 ±132
740
.47
.72
1.00
.56
.56
.59
0.25
.672 ±.007
.79
.34
630 ±519
624
.39
.81
1.00
.94
.94
.67
0.5
.658 ±.019
.72
.30
225 ±141
219
.11
.71
1.00
.87
.88
.74
0.75
.670 ±.011
.75
.31
87 ±26
81
.01
.75
1.00
.97
.95
.80
1
.653 ±.007
.71
.30
81 ±17
76
.00
.81
.90
.80
.86
.88
Appendix
Table 30: Complete results for Figure 10 on IEMOCAP ( λ=0.5 with the prior priced at λ , deployed seeds; γ=0 is the deployed SeMARC policy): macro-F1 and end-to-end GPU latency (mean ±std over three test folds), modality utilization, post-extraction GFLOPs, the extraction part of the latency, and the fraction of test samples on which each modality is acquired.
γ=0
0.25
0.5
0.75
1
1.5
2
IEMOCAP
.669 / 745
.672 / 630
.658 / 225
.670 / 87
.653 / 81
.651 / 82
.658 / 230
MM-Fi
.950 / 187
.946 / 58
.954 / 62
.952 / 69
.949 / 139
.947 / 134
.949 / 97
CMI
.556 / 12.6
.548 / 10.1
.550 / 10.4
.552 / 12.2
.555 / 12.1
.554 / 12.2
.548 / 12.4
CZU-MHAD
.914 / 1301
.914 / 1301
.920 / 1307
.916 / 1314
.920 / 1313
.922 / 1315
.922 / 1315
EAV
.689 / 217
.690 / 29
.697 / 156
.694 / 114
.691 / 56
.691 / 55
.692 / 100
UTD-MHAD
.827 / 659
.817 / 617
.800 / 599
.788 / 537
.791 / 544
.791 / 539
.797 / 566
Appendix
Table 31: Macro-F1 / end-to-end GPU latency (ms) for every γ on the six datasets, each fold at its deployed uniform policy’s seed and λtrain , with the prior priced at λtrain (mean over three test folds). γ=0 is the deployed SeMARC policy.
γ\λ
0.05
0.1
0.2
0.5
1
2
0
.661 / 1053
.675 / 994
.668 / 792
.669 / 745
.624 / 238
.598 / 5.5
0.25
.648 / 1207
.670 / 733
.672 / 877
.672 / 630
.620 / 23
.606 / 23
0.5
.664 / 751
.672 / 569
.678 / 818
.658 / 225
.640 / 44
.622 / 29
0.75
.678 / 1093
.676 / 535
.678 / 483
.670 / 87
.662 / 63
.640 / 42
1
.672 / 1097
.677 / 677
.671 / 93
.653 / 81
.656 / 89
.660 / 65
1.5
.670 / 695
.676 / 1082
.671 / 273
.651 / 82
.658 / 92
.647 / 81
Appendix
Table 32: Macro-F1 / end-to-end GPU latency (ms) on IEMOCAP for every (γ,λ) pair (mean over three test folds, deployed seeds, prior priced at λ ). The shaded cell is the deployed SeMARC policy.
Multimodal fusion architectures typically assume all modalities are available at inference, yet sensor failures, acquisition variability, and cost constraints routinely produce incomplete observations. Existing work treats modality absence as a prediction-accuracy problem, leaving a more basic question unanswered: whether a model's confidence estimates remain calibrated when an entire input stream is removed. We argue that missing-modality robustness and calibrated uncertainty are a single coupled property, and introduce Modality-Conditioned Conformal Fusion (MCCF), an architecture that addresses both at once. MCCF combines a multimodal bottleneck fusion backbone trained with modality dropout, per-modality evidential heads producing modality-decomposed Dirichlet distributions, and a Dempster-Shafer combination rule that fuses the per-modality evidence into a joint predictive distribution; an absent modality contributes vacuous evidence that is structurally ignored, so the fused uncertainty automatically reflects the reduced information without test-time imputation. A Mondrian conformal calibration module keyed on the modality-presence mask then provides finite-sample group-conditional coverage for every non-empty modality subset. MCCF is, to our knowledge, the first method with formal coverage guarantees under arbitrary modality availability through architectural integration rather than post-hoc recalibration, and the evidential decomposition yields per-modality vacuity scores that localise uncertainty to the absent modality responsible. Across a synthetic problem and three real multimodal benchmarks, MCCF holds its target coverage on every modality-presence subset, substantially narrows the coverage gap between full and partial modalities relative to a marginal split-conformal baseline, and imposes no measurable accuracy cost relative to temperature-scaled and evidential baselines.
Alireza Moayedikia
Department of Business Technology and Entrepreneurship Swinburne University of Technology Hawthorn, VIC 3122, Australia
Multimodal motion forecasting is inherently under-supervised: each training scene provides only one realized future, yet multiple plausible futures exist. This sparse supervision often leads to mode collapse (redundant hypotheses and insufficient mode coverage) and unreliable confidence ranking when predicting a small set of trajectories. We propose Mode-as-Sequence, a unified decoding framework that translates an unordered mode set into an ordered mode sequence and explicitly models mode-to-mode dependency. Under this framework, we develop two complementary instantiations. ModeSeq performs recurrent mode decoding, where each mode is generated conditioned on the previously generated modes, encouraging diverse, non-redundant hypotheses with calibrated confidence ordering. To remove the mode-by-mode autoregressive bottleneck, we further propose Parallel ModeSeq, which preserves the same causal dependency using masked mode-to-mode self-attention while decoding all modes in a single forward pass, enabling efficient large-K inference and scalable joint-scene prediction. To learn representative modes and calibrated confidence under sparse labels, we introduce Early-Match-Take-All (EMTA) and its joint-scene extension MA-EMTA, together with a lightweight ranking regularizer that reduces confidence inversions. Extensive experiments on large-scale benchmarks demonstrate consistent improvements in both ranking-oriented metrics and best-of-K accuracy across datasets, horizons, and object types. In the Waymo Open Dataset challenges, ModeSeq achieves 1st place in the 2024 LiDAR-free motion prediction track, and Parallel ModeSeq achieves 1st place in the 2025 Interaction Prediction Challenge, validating the effectiveness of Mode-as-Sequence for both accuracy and efficiency.
Zikang Zhou, Haibo Hu, Xinhong Chen +5
City University of Hong Kong · City University of Hong Kong (Dongguan) · Hon Hai Research Institute +1
Multimodal fusion must simultaneously refine modality-specific signals and model cross-modal interactions; two competing objectives typically entangled within the same operation. We propose \textbf{SeRIn} (\textbf{Se}gregate, \textbf{R}efine, \textbf{In}tegrate), a multimodal LM fusion scheme that enforces this separation as an architectural prior. Modality-specific representations evolve along isolated pathways, each refined against its respective encoder context, while a dedicated cross-modal pathway accumulates their joint evolution without contaminating unimodal streams. Full cross-modal interaction is deferred to a final prediction step - ablations confirm that structured interactions, not added capacity, drive the gains; gate analysis under visual corruption reveals emergent modality reweighting without explicit supervision. SeRIn achieves state-of-the-art results on CH-SIMS and CMU-MOSEI, improving all metrics on both benchmarks.
Alexios Filippakopoulos, Elias Kallioras, Nikolaos Xiros +2
National Technical University of Athens, Greece · Athena Research Center, Greece · University of Bern, Switzerland +2