Organizations: Zhejiang University, Hangzhou, China · China Academy of Space Technology, Beijing, China · Innovation and Management Center of the School of Software (Ningbo), Zhejiang University, Ningbo, China · Binjiang Institute of Zhejiang University, Hangzhou, China · Zhejiang Key Laboratory of Digital-Intelligence Service Technology, Hangzhou, China
Audio decisions often depend on evidence that a transcript does not preserve, while their probability estimates can depend on how answer options are ordered. AudioJev maps a waveform, question and supplied alternatives directly to a candidate distribution through one shared full-parameter model. We define order calibration as preserving an answer's probability under meaning-preserving option permutations. Random-derangement SKL training pairs each question with a reordered view in which every alternative changes position, supervises both answers, and aligns the two distributions before applying a symmetric KL penalty. Inference retains a single candidate-scoring forward, with no calibration head or order ensemble. Across three training seeds, AudioJev reaches 68.88%/55.33% mean accuracy on complete MMAU/MMAR and reduces random-order SKL by 43.8%/60.5% relative to the single-view removal ablation. The same model handles intent, environmental sound, note properties, speech activity and conversational transitions. Paired removal ablations and multi-order evaluation measure predictive accuracy and probability stability together, establishing a direct audio interface whose calibration objective acts on candidate meaning rather than presentation position. Inference code and model weights are available at https://github.com/SihanLv/AudioJev-Inference and https://huggingface.co/shlv/AudioJev.
Figures & tables
Figure 1: AudioJev overview. A waveform, question and caller-supplied alternatives define a decision for one shared audio classifier. Training (top): the original and randomly deranged answer lists receive candidate-only supervision, with identity-aligned symmetric KL encouraging consistent probabilities across presentations. Inference (bottom): one forward pass scores the supplied alternatives and returns their probability distribution, supporting different audio tasks through the same interface.
Model
SLURP@8
MINDS
MMSU
MMAU
MMAR
Clotho-AQA
NSynth
Latency (ms) ↓
Direct audio models
AudioJev
95.14 ± 0.40
88.59 ± 0.23
61.70 ± 0.68
68.88 ± 0.46
55.33 ± 0.84
90.93 ± 0.62
67.90 ± 0.92
62.4 ± 1.2
Frozen Omni 3B
85.42
84.67
60.14
68.64
53.50
90.93
45.70
-
Whisper large-v3 + text classifier
Laya English
76.56
54.91
44.90
38.70
35.60
-
-
1206.9 ± 2.5
Laya multilingual
70.31
41.22
37.73
32.71
36.40
-
-
1212.8 ± 0.6
Table 1: Classification accuracy (%) and waveform-to-probabilities latency. AudioJev reports mean ± sample SD over three seeds at the fixed final step. Text pipelines use the same Whisper large-v3 transcripts; timing includes fresh ASR on the fixed five-benchmark panel. AudioJev/Laya timings are refreshed; other reader timings are same-panel references (Appendix F ). Frozen Omni 3B shares AudioJev’s inference architecture. Dashes denote unreported measurements.
Model
Event@8
Event@36
Scene
VAD F1
Orig. AUC
Fresh AUC
Frozen Omni 3B
95.25
84.50
29.50
0.174
0.584
0.678
Frozen Omni 7B
96.00
88.75
33.25
0.387
0.582
0.651
CLAP
94.75
88.50
41.75
0.395
0.365
0.494
Silero VAD
-
-
-
0.873
-
-
Smart Turn (frozen)
-
-
-
-
0.855
0.804
AudioJev
96.17 ± 0.88
91.17 ± 0.95
58.83 ± 1.26
0.935 ± 0.014
0.895 ± 0.002
0.900 ± 0.014
Table 2: Acoustic and conversational tasks. Event and scene columns are accuracy (%); VAD is macro F1 and transition columns are AUROC. AudioJev: mean ± sample SD over three seeds. Fixed references use their established task definitions; Smart Turn transfers completion prediction to observed next-speaker change. Dashes indicate inapplicable tasks.
Task
Random SKL ↓
Random TV ↓
Flip (%) ↓
SLURP (all)
0.0387 ± 0.0064
0.0169 ± 0.0025
1.59 ± 0.37
MINDS
0.0623 ± 0.0113
0.0487 ± 0.0066
4.32 ± 1.69
MMSU
0.0805 ± 0.0123
0.0924 ± 0.0101
18.10 ± 2.29
Scene
0.0382 ± 0.0131
0.0602 ± 0.0086
6.33 ± 0.76
Event/activity
0.0157 ± 0.0046
0.0080 ± 0.0009
0.70 ± 0.19
Transition
0.0039 ± 0.0006
0.0114 ± 0.0055
0.42 ± 0.38
Table 3: Random-order calibration: mean ± sample SD over three seeds. Distributions are compared with the original ordering after semantic alignment. SKL and TV measure distribution differences; flip is the percentage of changed semantic answers. Random permutations are drawn independently of training partners. Lower is better.
Test
Acc. (%)
Macro F1
BA
AUROC
Transition
86.00 ± 2.95
0.741 ± 0.033
0.842 ± 0.007
0.895 ± 0.002
New sessions
85.67 ± 2.90
0.731 ± 0.031
0.823 ± 0.022
0.900 ± 0.014
Table 4: Causal next-speaker prediction (mean ± sample SD over three seeds). BA is balanced accuracy. Macro F1 and BA describe the native argmax decision; AUROC measures the ranking of switch probabilities.
MMAU full
MMSU
MMAR
Configuration
Acc. (%)
SKL ↓
Flip (%) ↓
Acc. (%)
SKL ↓
Flip (%) ↓
Acc. (%)
SKL ↓
Flip (%) ↓
w/o paired training
69.47 ± 0.32
0.1239 ± 0.0312
14.10 ± 1.29
62.55 ± 0.52
0.1577 ± 0.0402
19.91 ± 1.72
55.93 ± 0.42
0.1703 ± 0.0707
18.20 ± 2.94
AudioJev (RD-SKL)
68.88 ± 0.46
0.0697 ± 0.0050
13.09 ± 1.63
61.70 ± 0.68
0.0805 ± 0.0123
18.10 ± 2.29
55.33 ± 0.84
0.0674 ± 0.0108
15.44 ± 0.56
Table 5: Paired derangement training versus its removal ablation (mean ± sample SD over three seeds). Original examples, seeds, updates and final steps match. Accuracy is native-order percent; SKL and flip use independent non-cyclic reordering. The ablation removes the second view and SKL together.
MMAU full
MMSU
MMAR
λ ( n )
Acc. (%)
SKL ↓
Flip (%) ↓
Acc. (%)
SKL ↓
Flip (%) ↓
Acc. (%)
SKL ↓
Flip (%) ↓
0.25 (1)
68.97
0.0721
12.17
62.48
0.0871
16.87
55.50
0.0771
15.74
0.5 (3)
68.88 ± 0.46
0.0697 ± 0.0050
13.09 ± 1.63
61.70 ± 0.68
0.0805 ± 0.0123
18.10 ± 2.29
55.33 ± 0.84
0.0674 ± 0.0108
15.44 ± 0.56
0.75 (1)
68.95
0.0734
11.09
61.89
0.0911
15.47
54.70
0.0858
13.42
1 (3)
68.68 ± 1.12
0.0743 ± 0.0065
13.33 ± 0.09
61.99 ± 0.79
0.0928 ± 0.0082
17.89 ± 0.45
55.50 ± 0.66
0.0789 ± 0.0129
15.78 ± 1.75
1.5 (1)
68.04
0.0525
12.23
60.98
0.0571
17.48
54.00
0.0527
15.84
Table 6: SKL-weight sensitivity on MMAU/MMSU/MMAR. Rows with n=3 report mean ± sample SD; n=1 rows report one seed. Accuracy uses native order; SKL and flip use independent random reordering of 9,968/3,931/991 eligible questions. All runs use the fixed final checkpoint.
Evaluation order
Model
MMAU SKL ↓
MMSU SKL ↓
MMAR SKL ↓
One-step rotation
w/o paired training
0.1417 ± 0.0375
0.1826 ± 0.0549
0.2016 ± 0.1030
AudioJev (RD-SKL)
0.0750 ± 0.0057
0.0835 ± 0.0137
0.0742 ± 0.0107
Half-cycle rotation
w/o paired training
0.1509 ± 0.0359
0.1975 ± 0.0470
0.2013 ± 0.0738
AudioJev (RD-SKL)
0.0927 ± 0.0088
0.1143 ± 0.0181
0.0935 ± 0.0110
Random non-cyclic
w/o paired training
0.1239 ± 0.0312
0.1577 ± 0.0402
0.1703 ± 0.0707
AudioJev (RD-SKL)
0.0697 ± 0.0050
0.0805 ± 0.0123
0.0674 ± 0.0108
Table 7: Transfer across evaluation orders: mean ± sample SD over three paired training seeds. Each reordered distribution is compared with native order after semantic alignment, using the same 9,968/3,931/991 eligible MMAU/MMSU/MMAR questions for both models. All comparisons use fixed final checkpoints. Lower SKL is better; bold marks the lower mean within each order.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Capability
Dataset
Training questions
Event
ESC-50
1,024
Scene
TUT2018
1,024
Activity
Oration
512
Transition
SpokenWOZ
1,024
Intent
SLURP
512
Sound questions
Clotho-AQA
2,048
Appendix
Table 8: Original questions in the final general-audio stage; each has two training views. Repeated questions about a recording are not independent audio sources.
Seed
General step
MMAU
MMSU
MMAR
01
2048
69.41
62.45
56.30
02
2048
68.66
61.13
54.80
03
2048
68.56
61.51
54.90
Appendix
Table 9: Final RD-SKL models: fixed final general step and native-order accuracy (%). Seed labels abbreviate 20261001–20261003.
Seed
Public (1k)
Hidden (9k)
Full (10k)
01
71.10
69.22
69.41
02
71.10
68.39
68.66
03
70.50
68.34
68.56
Appendix
Table 10: MMAU partition accuracy (%) for the final method. Full accuracy adds public and hidden correct counts before division by 10,000.
Dataset
Native
Canonical
SLURP (all)
94.62 ± 0.29
94.64 ± 0.49
MINDS
88.59 ± 0.23
88.64 ± 0.48
MMSU
61.70 ± 0.68
62.27 ± 1.04
Scene
58.83 ± 1.26
59.17 ± 3.11
Scene (permuted)
59.25 ± 2.78
59.17 ± 3.11
Event/activity
95.96 ± 0.56
95.90 ± 0.55
Appendix
Table 11: Final-model accuracy (%): mean ± sample SD over three seeds. Native order is the main interface. Canonical order is a separate deterministic-wrapper diagnostic.
Seed
Clotho
NSynth
01
90.36
68.95
02
91.59
67.19
03
90.83
67.58
Appendix
Table 12: Held-out sound questions and note properties for each final RD-SKL seed: accuracy (%).
Seed01 configuration
MMAU acc.
SKL ↓
MMSU acc.
SKL ↓
MMAR acc.
SKL ↓
Ablation: no paired training
69.55
0.0879
63.14
0.1145
56.40
0.1055
Cyclic pairs, CE
69.87
0.1182
61.51
0.1534
56.10
0.1167
Cyclic pairs, CE+SKL ( λ=1 )
69.48
0.0547
61.56
0.0631
54.00
0.0509
Deranged pairs, CE+SKL ( λ=1 )
69.94
0.0667
61.82
0.0869
56.20
0.0696
AudioJev: deranged pairs ( λ=0.5 )
69.41
0.0644
62.45
0.0742
56.30
0.0548
Appendix
Table 13: Seed01 pairing controls with explicit SKL weights. Accuracy is native-order percent; SKL measures an independent non-cyclic random reordering. All two-view arms share original examples and update budgets. The single-view row is the removal ablation. The cyclic CE row is not a matched random-derangement CE control, so differences against it are recipe effects.
Pipeline
ASR
Reader
Total
p50
p95
AudioJev (RD-SKL)
0.0
62.4
62.4 ± 1.2
59.6
84.7
ASR + Laya English
1177.8
29.1
1206.9 ± 2.5
557.3
7248.0
ASR + Laya multilingual
1188.8
24.1
1212.8 ± 0.6
564.1
7304.1
Appendix
Table 14: Fresh final-model and ASR–Laya timing on the fixed panel, in milliseconds. Total means include sample SD across three repetition means.
Task/order
SKL ↓
TV ↓
Flip (%) ↓
MMAU full: one-step
0.0750 ± 0.0057
0.0890 ± 0.0059
14.04 ± 2.14
MMAU full: half-cycle
0.0927 ± 0.0088
0.1025 ± 0.0090
16.31 ± 2.99
MMAU full: random
0.0697 ± 0.0050
0.0829 ± 0.0052
13.09 ± 1.63
MMAU full: reverse
0.0960 ± 0.0101
0.1036 ± 0.0077
16.64 ± 2.88
MMSU: one-step
0.0835 ± 0.0137
0.0977 ± 0.0121
19.68 ± 2.52
MMSU: half-cycle
0.1143 ± 0.0181
0.1189 ± 0.0161
23.63 ± 3.17
Appendix
Table 15: Multi-order diagnostics for the final method: mean ± sample SD over three seeds. Probability differences are computed against native order after semantic alignment; flip rates are percentages.