Audio Token Attention Is Predictable Before the Language Model Runs
Organizations: The University of Texas at Austin
Abstract
A large audio language model (LALM) turns a minute of speech into 750-1,500 tokens and prefills every one. Image-token pruning often cuts after the language model's first layers, where image tokens draw little attention. Audio tokens draw much more attention there, and their ranking is still far from final, so audio needs a ranking before the language model runs. Surprisingly, the attention an audio token will receive across the language model is already linearly predictable from its encoder output, before the language model runs. A linear map, fitted in closed form without labels, predicts this all-layer attention ranking at on eleven of thirteen LALMs. Our method, Triage, cuts audio tokens by this prediction and, on multiple choice, cuts again at layer 2, correcting the prediction with the attention observed there. Triage sets its compression without labels, under two budgets that limit how far its output may differ from the model's own full-audio output. At the conservative budget, its word error rate and accuracy stay within .04 of full audio. At the aggressive budget, Triage beats every baseline in all twelve transcription cases. On multiple choice, at 2.2-5x compression, it outperforms DART, the strongest baseline on average, by .043 in mean accuracy. Because it cuts before the language model, it raises the audio that fits in Qwen2.5-Omni-3B's context window from 21.8 to about 62 minutes. At its most compressive point, Triage lets one GPU serve 4x as many concurrent 5-minute streams of that model. Project page: https://audio-triage.github.io
Figures & tables
| against | the prior | enc. self-attn | acoustic energy | state norm | self-information |
| Ultravox, B | |||||
| Phi-4, B | |||||
| Voxtral, B | |||||
| Qwen2.5-Omni, B | |||||
| Qwen2.5-Omni, B | |||||
| Qwen3-Omni, B |
| Qwen2.5-Omni-3B | Qwen3-Omni-30B | Voxtral | Phi-4 | ||||||||||||
| (a) WER | Libri | FLEURS | TED | Libri | FLEURS | TED | Libri | FLEURS | TED | Libri | FLEURS | TED | |||
| full audio | .083 | .153 | .215 | .011 | .037 | .020 | .020 | .043 | .032 | .017 | .045 | .085 | |||
| [2pt/1.5pt] conservative ( ) | 1.54 | 1.11 | 1.25 | 1.54 | 1.54 | 1.67 | 1.54 | 1.54 | 2.00 | 1.54 | 1.54 | 1.54 | |||
| ours | .099 | .177 | .197 | .025 | .050 | .049 | .025 | .043 | .071 | .045 | .064 | .102 | |||
| [2pt/1.5pt] aggressive ( ) | 2.00 | 1.18 | 1.67 | 2.00 | 2.00 | 2.00 | 2.86 | 2.86 | 2.38 | 2.86 | 2.86 | 2.86 | |||
| uniform pooling | .348 | .325 | .403 | .160 | .140 | .123 | .398 | .385 | .241 | .383 | .368 | .488 | |||
Appendix figures & tables34 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | the passage | tokens | mass at | mass / token | answered right |
| Qwen3-Omni-30B | heard | ||||
| ( ) | read, as a page | ||||
| read, as text | |||||
| Qwen2.5-Omni-3B | heard | ||||
| ( ) | read, as a page | ||||
| read, as text |
| audio tokens kept | |||
| linear prior | |||
| MLP prior | |||
| best baseline | ( vad ) | ( dart ) | ( dart ) |
| true attention (oracle) |
| Model (tok/s) | audio | full (ms) | ||||
| Qwen2.5-Omni-3B ( ) | min | |||||
| min | ||||||
| Voxtral-Mini-3B ( ) | min | |||||
| min | ||||||
| Phi-4-multimodal ( ) | min | |||||
| min |
| before fitting | after fitting | ||
| model | audio front end | ||
| above the threshold: the prior works | |||
| Qwen3-Omni-30B (MoE) | trained together | ||
| SeaLLMs-Audio-7B [ 63 ] | Qwen2-Audio Qwen2.5 | ||
| Aero-1-Audio [ 65 ] | Qwen2-Audio Qwen2.5 | ||
| Ultravox-v0.5-Llama-3.2-1B [ 53 ] | Whisper Llama-3.2 | ||
| Model | benchmark | prior | obs. |
| Qwen 3B | MMSU | .678 | .516 |
| DREAM | .644 | .467 | |
| RACE | .564 | .497 | |
| Qwen 30B | MMSU | .574 | .414 |
| DREAM | .577 | .401 | |
| RACE | .548 | .391 |
| Qwen2.5-Omni-3B | Qwen3-Omni-30B (MoE) | ||||||
| Benchmark | in-dom. | univ. | 0-shot | in-dom. | univ. | 0-shot | |
| speech QA | |||||||
| MMSU | .698 .006 | .759 .003 ∗∗ | .756 .004 | .653 .007 ∗∗ | .629 .007 | .613 .007 | |
| AudioM.-RACE | .666 .002 | .713 .002 ∗∗ | .690 .002 | .561 .003 | .559 .003 | .549 .003 | |
| DREAM | .675 .004 | .745 .003 ∗∗ | .740 .003 | .646 .005 | .645 .005 | .635 .005 | |
| transcription | |||||||
| Condition | fwd | tokens | 20 | 25 | 30 | 35 | 40 | 45 min | mean |
| oracle span (ceiling) | 1 | — | .847 | .857 | .873 | .803 | .800 | .843 | .837 |
| chunked, full audio | 1 | .730 | fails: every question over context | — | |||||
| single call (today’s default) | 1 | .513 | .500 | .497 | .410 | .470 | .500 | .482 | |
| per-minute window voting | .700 | .717 | .740 | .667 | .677 | .733 | .706 | ||
| Triage, conserv. (.50/.35) | 1 | .787 | .773 | .780 | .630 | .700 | .713 | .731 | |
| Triage, aggr. (.35/.20) | 1 | .733 | .763 | .787 | .607 | .697 | .743 | .722 | |
| min, all needles required | tokens | ||||
| full audio (uncompressed reference) | |||||
| Triage, conserv. (.50/.35) | |||||
| Triage, aggr. (.35/.20) | |||||
| Stage 1 alone | |||||
| universal prior | |||||
| universal prior | |||||
| conservative | aggressive | |||
| Method | acc | [95% CI] | acc | [95% CI] |
| (a) Qwen3-Omni-30B, of and , reference full audio | ||||
| full audio (reference) | ||||
| no audio | [ ] | |||
| one-minute window vote | [ ] | |||
| pool | [ ] | [ ] | ||
| Qwen 30B, conservative | Qwen 30B, aggressive | Qwen 3B, aggressive | ||||
| Method | acc | [95% CI] | acc | [95% CI] | acc | [95% CI] |
| reference | ||||||
| no audio | [ ] | [ ] | ||||
| one-minute window vote | [ ] | [ ] | ||||
| pool | [ ] | [ ] | [ ] | |||
| vad | [ ] | [ ] | [ ] | |||
| Model | Method | encoder | language-model prefill | end to end |
| Qwen 30B | no audio | — | ||
| full audio (reference) | ||||
| one-minute window vote | ||||
| Triage, Stage 1 only: top- , conservative | ||||
| Triage, both stages, conservative | ||||
| Triage, Stage 1 only: top- , aggressive |
| Signal | MMSU | DREAM | RACE | MMSU, 2nd sample |
| Acoustic energy | ||||
| Self-information ( ) | ||||
| State norm |
| Signal | Representation | vs. |
| The prior (ours) | encoder | |
| Self-information ( ) | encoder | |
| Self-information ( ) | raw waveform | |
| Self-information ( ) | raw waveform | |
| Spectral entropy | raw waveform | |
| Acoustic energy | raw waveform |
| Bench | Operating point | full | top- (ours) | self-info | inverse | VAD | random |
| MMSU | conserv. | .603 | .603 | .583 | .585 | .595 | .608 |
| aggressive | .583 | .515 | .520 | .552 | .505 | ||
| DREAM | conserv. | .894 | .861 | .582 | .801 | .831 | .747 |
| aggressive | .806 | .489 | .584 | .622 | .554 | ||
| RACE | conserv. | .818 | .835 | .738 | .777 | .820 | .782 |
| aggressive | .820 | .615 | .688 | .718 | .677 |
| LibriSpeech | FLEURS | |||
| Selector | aggressive | conservative | aggressive | conservative |
| Qwen2.5-Omni-3B | ||||
| kept | ||||
| full audio (ceiling) | ||||
| bin coverage (ours) | ||||
| stride (decimation) | ||||
| (a) Retained set, kept | |||
| Metric | Ours | attn top- | Random |
| Silence fraction (TEDLIUM) | .018 | .088 | .154 |
| Function-word share (LibriSpeech) | .144 | .190 | .291 |
| Mean RMS energy | .091 | .074 | .060 |
| (b) What each stage drops ( -scored) | |||
| Removed at | prior | observed | all-layer |
| between the prior and | on LibriSpeech | on FLEURS |
| no controls | ||
| controlling for waveform descriptors | ||
| (position, RMS, spectral centroid, flatness, rolloff, zero-crossing rate, onset strength) | ||
| controlling for lexical descriptors | ||
| (in-word, content-word, distance to a word boundary, word duration) | ||
| controlling for both |
| Voxtral-Mini-3B ( layers) | Qwen2.5-Omni-3B ( layers) | |||
| summed over | share of | share of | ||
| every layer | ||||
| layer and deeper | ||||
| layer and deeper | ||||
| the deepest quarter | ||||
| layer alone | ||||
| Control (same #tokens removed) | delete | delete | delete |
| Voxtral-Mini-3B, WER | |||
| at random | [ ] | [ ] | [ ] |
| the prior’s own bottom | [ ] | [ ] | [ ] |
| the highest-energy frames | [ ] | [ ] | [ ] |
| matched to its position | [ ] | [ ] | [ ] |
| self-information | [ ] | [ ] | [ ] |
| Bench | kept | full | obs tail | obs | prior | precision fusion |
| MMSU | ||||||
| DREAM | ||||||
| prior shortlist | random shortlist | effect | ||||||
| Bench | Operating point | full | observed- only | precision fusion | observed- only | precision fusion | ||
| MMSU | conserv. (.85/.65) | .603 | .600 | .600 | .595 | .593 | ||
| aggr. (.35/.25) | .580 | .585 | .548 | .538 | ||||
| DREAM | conserv. (.65/.45) | .892 | .878 | .865 | .788 | .805 | ||
| aggr. (.50/.20) | .815 | .860 | .675 | .713 | ||||
| RACE | conserv. (.85/.65) | .818 | .815 | .823 | .808 | .813 | ||
| Design choice | acc |
| Stage 2, fusion rule on a fixed shortlist (prior only: ) | |
| observed-only ( ) | .762 |
| product | .765 |
| uniform blend ( ) | .766 |
| max-rank | .767 |
| cascade (multi-layer) | .767 |
| direct cut | achieved vs direct cut | |||||
| Benchmark | kept | SD / SE | SD ratio | exact | ||
| DREAM | ||||||
| DREAM | ||||||
| DREAM | ||||||
| RACE | ||||||
| RACE | ||||||
| Benchmark | What it tests | Metric | Scored in |
| Multiple-choice QA | |||
| MMSU [ 35 ] | targeted listening, s | gold accuracy | Tab. H.3 |
| DREAM [ 36 ; 49 ] | dialogue comprehension, s | gold accuracy | Tab. H.3 |
| AudioMarathon-RACE [ 37 ; 55 ] | reading comprehension, min | gold accuracy | Tab. H.3 |
| Transcription | |||
| LibriSpeech [ 32 ] | read-aloud speech, test-clean | corpus WER | Tab. 2 |
| Model | Bench | Triage | DART | layer 2 | shortlist |
| Qwen 3B | MMSU | .585 | .547 | .550 | .565 |
| DREAM | .863 | .757 | .738 | .778 | |
| RACE | .820 | .805 | .812 | .818 | |
| Qwen 30B | MMSU | .682 | .620 | .652 | .618 |
| DREAM | .943 | .907 | .853 | .843 | |
| Voxtral | MMSU | .525 | .532 | .537 | .448 |
| Selector | MMSU | DREAM | RACE |
| Qwen2.5-Omni-3B , two disjoint samples, A / B ( ; RACE each) | |||
| precision fusion (ours) | / | / | / |
| top- (prior only, ours) | / | / | / |
| DART | / | / | / |
| HeadRouter N | / | / | / |
| FastV N | / | / | / |
| model | bench | ours | dart | Holm | |||
| Qwen2.5-Omni-3B | MMSU | ||||||
| DREAM | ✓ | ||||||
| RACE | |||||||
| Qwen3-Omni-30B | MMSU | ✓ | |||||
| DREAM | ✓ | ||||||
| RACE |
| keep | keep | ||||
| Target : query rows | direct cut | achieved | direct cut | achieved | |
| trailing rows | .658 | .817 | .782 | .780 | .695 |
| all text rows | .677 | .825 | .825 | .818 | .770 |
| text rows, template tail dropped | .675 | .827 | .823 | .810 | .772 |
| audio rows (ours) | .707 | .815 | .825 | .798 | .815 |
| all rows | .714 | .815 | .825 | .802 | .812 |
| Qwen2.5-Omni-3B ( min) | Qwen3-Omni-30B (MoE, min) | |||||||||
| Task | Budget | prefill / KV% | streams | tok/s | prefill / KV% | streams | tok/s | |||
| MCQ | conserv. | .65/.45 | 2.3 / | 2 | 2.1 | .85/.55 | 1.4 / | 2 | 1.9 | |
| MCQ | aggr. | .50/.20 | 3.7 / | 4 | 4.5 | .85/.35 | 1.6 / | 3 | 2.8 | |
| ASR | conserv. | .85/.65 | 1.6 / | 1.5 | 1.5 | .90/.65 | 1.2 / | 1.5 | 1.5 | |
| ASR | aggr. | .80/.50 | 2.0 / | 2 | 1.9 | .70/.50 | 1.5 / | 2 | 2.0 | |
| Config | prefill | KV% | streams | tok/s |
| Qwen2.5-Omni-3B (5-min), aggressive (MCQ) | ||||
| full ( ) | ms | 256 | 666 | |
| Stage-1 only ( ) | ms ( ) | 512 | 1326 ( ) | |
| + Stage-2 ( ) | ms ( ) | 1024 ∗ | 2993 ( ) | |
| Qwen3-Omni-30B (40-min), aggressive (MCQ) | ||||
| full ( ) | s | 8 | 28 | |
| Benchmark | Model | audio full | ours (conserv.) | ours (aggr.) | Whisper text |
| MMAU-mini sound | Qwen 3B | .62 | .62 | .59 | .44 |
| Qwen 30B | .66 | .63 | .66 | .51 | |
| MMAR sound | Qwen 3B | .54 | .54 | .48 | .39 |
| Qwen 30B | .66 | .66 | .62 | .42 | |
| IEMOCAP emotion | Qwen 3B | .62 | .62 | .54 | .57 |
| Qwen 30B | .65 | .58 | .53 | .56 |
| Model | Audio | ours (compressed) | cascade (Whisper text) | speedup |
| Qwen 3B | 30 s | ms | s | |
| 1 min | ms | s | ||
| 5 min | ms | s | ||
| Qwen 30B | 30 s | s | s | |
| 5 min | s | s | ||
| 10 min | s | s |