ASD: Accelerating On-Demand Audio Understanding on Mobile Devices
Organizations: The University of Texas at Austin · Shanghai Jiao Tong University
Abstract
Speculative decoding accelerates autoregressive generation by using a smaller drafter to propose tokens for batched verification by a larger target. However, conventional speculative decoding couples drafting to the target's evolving verified prefix, serializing drafting and verification. We ask whether this dependency is necessary for source-conditioned generation. Our key observation is that, for audio language models, the input audio and user request can provide useful speculative candidates without following the target's evolving text prefix. We propose ASD (Audio Speculative Speculative Decoding), which enables target-decoupled drafting: an audio-conditioned drafter follows its own generation history while the target independently verifies and corrects ready candidates. Without usable candidates, the target advances alone. Thus, target feedback determines which candidates are committed but no longer determines when the drafter can make progress, enabling drafting and verification to proceed concurrently while retaining target-side verification and correction. We implement ASD in MNN for Android and evaluate two target models across four phones, seven datasets, and three tasks covering 12.2 hours of audio. Across four phones, ASD improves pooled ASR throughput by 42-76% over target-only decoding, while only 5.7% of evaluation windows are slower than target-only, compared with 58.1-63.0% for speculative baselines. For ASR, ASD reaches 97.33-98.20% of a hindsight per-window oracle's pooled throughput over the evaluated drafter/budget catalog. Native on-demand execution with a 7B target achieves up to 78% higher throughput than target-only. These results show that source-conditioned audio generation can relax the conventional dependence of speculative drafting on the target's evolving output prefix, exposing substantial parallelism for efficient inference.
Figures & tables
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Redmi K70 Pro | Xiaomi 11 Pro | Xiaomi 14 | Redmi Note 9 Pro | |
| Model ID | 23117RK66C | M2102K1AC | 23127PN0CC | M2007J17C |
| Snapdragon SoC | 8 Gen 3 | 888 | 8 Gen 3 | 750G |
| Process (nm) | 4 | 5 | 4 | 8 |
| CPU family | Kryo (X4 prime) | Kryo 680 | Kryo (X4 prime) | Kryo 570 |
| CPU cores | 8 | 8 | 8 | 8 |
| Peak CPU (GHz) | 3.3 | 2.84 | 3.3 | 2.2 |
| Check | Windows | Mean absolute error (%) | Pooled error (%) |
|---|---|---|---|
| Initial validation | 9 | 5.81 | |
| Fresh repeat | 5 | 6.54 |
| Redmi K70 Pro | Xiaomi 11 Pro | |||||||||||
| Method | Libri | AMI | Flrs | KeS | M4S | All | Libri | AMI | Flrs | KeS | M4S | All |
| Target-only | 16.92 | 16.92 | 16.94 | 17.00 | 17.04 | 16.96 | 8.33 | 8.33 | 8.34 | 8.37 | 8.39 | 8.35 |
| Standard SD | 21.52 | 15.46 | 19.91 | 17.63 | 18.26 | 18.97 | 11.77 | 8.39 | 10.87 | 9.59 | 9.94 | 10.34 |
| SpecASR | 21.44 | 11.19 | 18.35 | 15.44 | 15.61 | 16.81 | 12.46 | 7.32 | 10.97 | 9.37 | 9.64 | 10.21 |
| As 2 d | 35.37 | 18.60 | 28.22 | 27.57 | 25.59 | 27.92 | 18.94 | 9.82 | 14.98 | 14.24 | 13.22 | 14.68 |
| Gain (%) | +64.4 | +9.9 | +41.7 | +56.4 | +40.2 | +47.2 | +52.0 | +17.0 | +36.6 | +48.4 | +33.0 | +41.9 |
| Setting | LibriSpeech | AMI | FLEURS | KeSpeech | M4Singer | All five |
|---|---|---|---|---|---|---|
| Redmi K70 Pro | ||||||
| 19.77 | 15.56 | 18.67 | 17.15 | 17.62 | 18.09 | |
| 21.41 | 15.82 | 19.87 | 17.90 | 18.43 | 19.10 | |
| 21.97 | 15.28 | 20.19 | 17.65 | 18.35 | 19.13 | |
| 22.93 | 15.19 | 20.90 | 17.81 | 18.62 | 19.55 | |
| Mean | 21.52 | 15.46 | 19.91 | 17.63 | 18.26 | 18.97 |
| Fixed SD | LibriSpeech | AMI | FLEURS | KeSpeech | M4Singer | All five |
|---|---|---|---|---|---|---|
| Standard SD (K8) | 22.93 | 15.19 | 20.90 | 17.81 | 18.62 | 19.55 |
| Standard SD (K12) | 22.79 | 12.45 | 19.53 | 15.55 | 16.52 | 17.76 |
| Case group | Standard SD | SpecASR | As 2 d |
|---|---|---|---|
| A | +28.8 [+24.3, +32.3] | +31.2 [+15.6, +42.4] | +124.4 [+104.1, +134.6] |
| B | -16.7 [-19.4, -14.5] | -14.7 [-20.1, -10.8] | +64.4 [+61.9, +66.6] |
| C | -22.3 [-27.6, -18.2] | -24.3 [-32.1, -18.4] | +61.2 [+54.5, +64.6] |
| D | -29.2 [-36.2, -24.2] | -31.9 [-40.1, -25.2] | +54.6 [+45.7, +60.1] |
| E | -27.9 [-33.7, -21.8] | -30.7 [-39.9, -22.5] | +50.8 [+38.3, +58.6] |
| F | -39.6 [-44.0, -35.1] | -48.0 [-53.1, -43.0] | +21.4 [+3.3, +33.0] |
| Phone | Method | Faster | Slower | Median loss | Max. loss |
|---|---|---|---|---|---|
| Redmi K70 Pro | Standard SD | 541 | 147 | 9.8% | 62.1% |
| SpecASR | 375 | 313 | 19.2% | 74.3% | |
| As 2 d | 641 | 47 | 14.2% | 21.9% | |
| Xiaomi 11 Pro | Standard SD | 611 | 77 | 8.3% | 58.5% |
| SpecASR | 553 | 135 | 10.7% | 67.8% | |
| As 2 d | 657 | 31 | 7.4% | 20.3% |
| Case | Device/data | Comparison | TPS | (%) | |
|---|---|---|---|---|---|
| Supply | Mi 14/Libri | Qwen Parakeet | 28.36 46.33 | 74.0 69.1 | 5 |
| Verification | Mi 14/AMI | Qwen Parakeet | 22.11 22.74 | 41.1 23.9 | 5 |
| TD, high acc. | K70/Libri | Serial TD | 15.41 29.81 | 77.4 81.7 | 5 |
| TD, low acc. | K70/KeSpeech | Serial TD | 9.62 15.80 | 44.9 24.3 | 1 |
| TD, zero acc. | K70/M4Singer | Serial TD | 11.85 12.52 | 80.4 0.0 | 1 |
| Prefix lookahead | K70/AMI | Serial lookahead | 9.25 9.41 | 37.0 37.0 | 2 |
| Drafter | Max. | TPS | Gain (%) | Extra tokens | Target calls |
|---|---|---|---|---|---|
| Target-only | scalar | 14.58 | 0.0 | 0 | 211 |
| Parakeet CTC | 3 | 15.44 | 5.9 | 33 | 178 |
| Parakeet CTC | 4 | 16.25 | 11.5 | 39 | 172 |
| Parakeet CTC | 8 | 15.66 | 7.4 | 43 | 168 |
| SenseVoice | 3 | 16.03 | 10.0 | 24 | 187 |
| SenseVoice | 4 | 16.03 | 9.9 | 26 | 185 |
| Method | Energy (J) | Mean power (W) | Response (s) | Body tokens | J/token |
|---|---|---|---|---|---|
| Target-only | 341.6 | 10.47 | 32.64 | 256 | 1.334 |
| As 2 d | 239.4 | 11.33 | 21.13 | 248 | 0.965 |
| Metric | With cooler | Without cooler |
|---|---|---|
| Mean encoder / prefill / decode (s) | 0.70 / 2.24 / 0.68 | 0.94 / 2.95 / 0.78 |
| Mean complete block processing (s) | 3.63 | 4.68 |
| Active / resource RTF | 0.907 / 1.007 | 1.171 / 1.201 |
| Decode throughput (tokens/s) | 13.87 | 12.23 |
| p95 completion lag (s) | 4.68 | 103.39 |
| Final completion lag (s) | 4.09 | 120.50 |