Speculative decoding accelerates autoregressive generation by using a smaller drafter to propose tokens for batched verification by a larger target. However, conventional speculative decoding couples drafting to the target's evolving verified prefix, serializing drafting and verification. We ask whether this dependency is necessary for source-conditioned generation. Our key observation is that, for audio language models, the input audio and user request can provide useful speculative candidates without following the target's evolving text prefix. We propose AS2D (Audio Speculative Speculative Decoding), which enables target-decoupled drafting: an audio-conditioned drafter follows its own generation history while the target independently verifies and corrects ready candidates. Without usable candidates, the target advances alone. Thus, target feedback determines which candidates are committed but no longer determines when the drafter can make progress, enabling drafting and verification to proceed concurrently while retaining target-side verification and correction. We implement AS2D in MNN for Android and evaluate two target models across four phones, seven datasets, and three tasks covering 12.2 hours of audio. Across four phones, AS2D improves pooled ASR throughput by 42-76% over target-only decoding, while only 5.7% of evaluation windows are slower than target-only, compared with 58.1-63.0% for speculative baselines. For ASR, AS2D reaches 97.33-98.20% of a hindsight per-window oracle's pooled throughput over the evaluated drafter/budget catalog. Native on-demand execution with a 7B target achieves up to 78% higher throughput than target-only. These results show that source-conditioned audio generation can relax the conventional dependence of speculative drafting on the target's evolving output prefix, exposing substantial parallelism for efficient inference.
Figures & tables
Figure 1: As 2 d overlaps drafting and verification for on-demand audio understanding. (a) Background audio prefill prepares LM KV before the request. Response latency includes instruction prefill, decoding, and any residual audio preparation. (b) Longer replies require more sequential target steps. (c) Batched verification saves target work, but serial drafting can consume those savings on a phone. (d) As 2 d keeps the target advancing: it verifies ready candidates ( Vi ) or takes a target-only step ( T ) without waiting for drafting. Audio-based drafting ( Di ) runs alongside, independent of the latest verified prefix. Dashed arrows supply ready candidates at target boundaries. The target determines committed tokens. Decoding timelines are schematic and start after the request.
Figure 2: As 2 d decouples candidate generation from target verification. Request zu=(xu,pu) supplies fixed audio and an instruction to target-decoupled drafting. Verification and alignment processes pu using retained target KV Cu . The target verifies ready candidates or decodes alone while the producer continues independently. Alignment updates the buffer. Device-profiled configuration combines measured phone costs with development traces to select a fixed budget for the given drafter.
Figure 3: Drafting is costly, yet need not wait for updated text. (a) Native serial SD pairs a CPU 0.6B drafter with an OpenCL 1.7B target. Time shares pool five runs per phone on one 59.5-second audio, with seven candidates per round. Drafting includes prefix catch-up. Loading and initial prefill are excluded. (b) We compare acceptance with and without preceding transcript context. The comparison uses 120 four-second blocks per corpus, a 32-token proposal cap, and identical verification context. Both panels use Qwen3-ASR. Appendix B.3 gives both protocols.
Figure 4: As 2 d publishes, verifies, and reuses partial drafts. (a) Only published candidates enter verification. (b) The target accepts AB , corrects C to X , and commits ABX . (c) The target’s next token D matches the buffer and supplies the root for verifying ready EF . The producer continues unchanged. The target verifies each reused candidate. Tokens and batch lengths are illustrative.
Figure 5: As 2 d improves decoding throughput across four phones. Each panel contains the same 688 complete-audio test windows from five datasets, with separately measured phone-cost replay. Curves are unsmoothed. Farther right indicates higher TPS. SpecASR uses the serial ASP + recycling port. SD averages K=5,6,7,8 TPS within each window. Horizontal scales differ by phone.
Table 6Table 7Table 8
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Batching savings depend on phone and width. We report per-token verification cost reductions for Qwen3-ASR 1.7B on OpenCL relative to the K=1 verifier. Costs average 15 calls. Bars scale one call-time standard deviation by the same baseline. These are verifier savings, not full-request gains.
Redmi K70 Pro
Xiaomi 11 Pro
Xiaomi 14
Redmi Note 9 Pro
Model ID
23117RK66C
M2102K1AC
23127PN0CC
M2007J17C
Snapdragon SoC
8 Gen 3
888
8 Gen 3
750G
Process (nm)
4
5
4
8
CPU family
Kryo (X4 prime)
Kryo 680
Kryo (X4 prime)
Kryo 570
CPU cores
8
8
8
8
Peak CPU (GHz)
3.3
2.84
3.3
2.2
Appendix
Table 6: The four phones span different hardware and software configurations. RAM denotes OS-visible total memory, and CPU clocks are specified maxima. The lower block lists the fixed models, placement, and proposal budgets used for the ASR comparison.
Check
Windows
Mean absolute error (%)
Pooled error (%)
Initial validation
9
5.81
−4.58
Fresh repeat
5
6.54
−4.95
Appendix
Table 7: The table summarizes replay TPS errors in the initial check and fresh repeat.
Redmi K70 Pro
Xiaomi 11 Pro
Method
Libri
AMI
Flrs
KeS
M4S
All
Libri
AMI
Flrs
KeS
M4S
All
Target-only
16.92
16.92
16.94
17.00
17.04
16.96
8.33
8.33
8.34
8.37
8.39
8.35
Standard SD
21.52
15.46
19.91
17.63
18.26
18.97
11.77
8.39
10.87
9.59
9.94
10.34
SpecASR
21.44
11.19
18.35
15.44
15.61
16.81
12.46
7.32
10.97
9.37
9.64
10.21
As 2 d
35.37
18.60
28.22
27.57
25.59
27.92
18.94
9.82
14.98
14.24
13.22
14.68
Gain (%)
+64.4
+9.9
+41.7
+56.4
+40.2
+47.2
+52.0
+17.0
+36.6
+48.4
+33.0
+41.9
Appendix
Table 8: As 2 d improves pooled throughput over the baselines on all four phones. Values report phone-profile replay TPS on the same 688 windows. Standard SD pools tokens and time for each fixed K=5,6,7,8, then averages the four TPS values. All includes all five datasets. Bold identifies As 2 d , and underlines mark the best baseline per phone and dataset. Gain is the relative TPS increase over that baseline, excluding ablations. Xiaomi 14 retains Qwen/budget 7, while Note 9 Pro uses validation-selected fixed Qwen/budget 3. Matched without-TD traces cover all 688 windows. Section 5.3 specifies the settings, and Appendix D.2 describes calibration.
Setting
LibriSpeech
AMI
FLEURS
KeSpeech
M4Singer
All five
Redmi K70 Pro
K=5
19.77
15.56
18.67
17.15
17.62
18.09
K=6
21.41
15.82
19.87
17.90
18.43
19.10
K=7
21.97
15.28
20.19
17.65
18.35
19.13
K=8
22.93
15.19
20.90
17.81
18.62
19.55
Mean
21.52
15.46
19.91
17.63
18.26
18.97
Appendix
Table 9: Standard SD throughput varies with fixed verification width on the same 688 windows. Physical K includes one known root. For each fixed configuration, TPS pools tokens and decode time. Mean denotes the equal-weight arithmetic mean of TPS at K=5,6,7,8. Range shows variation across these configurations, not run-to-run uncertainty. Both phones use the same averaging rule.
Fixed SD
LibriSpeech
AMI
FLEURS
KeSpeech
M4Singer
All five
Standard SD (K8)
22.93
15.19
20.90
17.81
18.62
19.55
Standard SD (K12)
22.79
12.45
19.53
15.55
16.52
17.76
Appendix
Table 10: Fixed-width SD at K8 outperforms K12 on all five datasets in the K70 replay cohort. K includes one known root. These controls are separate from the main-table K5–8 mean. Neither configuration is claimed to be a validation-selected optimum.
Case group
Standard SD
SpecASR
As 2 d
A
+28.8 [+24.3, +32.3]
+31.2 [+15.6, +42.4]
+124.4 [+104.1, +134.6]
B
-16.7 [-19.4, -14.5]
-14.7 [-20.1, -10.8]
+64.4 [+61.9, +66.6]
C
-22.3 [-27.6, -18.2]
-24.3 [-32.1, -18.4]
+61.2 [+54.5, +64.6]
D
-29.2 [-36.2, -24.2]
-31.9 [-40.1, -25.2]
+54.6 [+45.7, +60.1]
E
-27.9 [-33.7, -21.8]
-30.7 [-39.9, -22.5]
+50.8 [+38.3, +58.6]
F
-39.6 [-44.0, -35.1]
-48.0 [-53.1, -43.0]
+21.4 [+3.3, +33.0]
Appendix
Table 11: Medians and interquartile ranges summarize each illustrative regime. All windows are retained. Values are per-window TPS changes (%) relative to target-only, rather than changes in pooled TPS. The intervals describe variation across windows, not confidence intervals. Negative values indicate slowdowns.
Phone
Method
Faster
Slower
Median loss
Max. loss
Redmi K70 Pro
Standard SD
541
147
9.8%
62.1%
SpecASR
375
313
19.2%
74.3%
As 2 d
641
47
14.2%
21.9%
Xiaomi 11 Pro
Standard SD
611
77
8.3%
58.5%
SpecASR
553
135
10.7%
67.8%
As 2 d
657
31
7.4%
20.3%
Appendix
Table 12: The counts compare complete-window throughput with target-only. All 688 windows per phone are retained. Losses measure TPS reductions among slower windows.
Case
Device/data
Comparison
TPS
α (%)
n
Supply
Mi 14/Libri
Qwen → Parakeet
28.36 → 46.33
74.0 → 69.1
5
Verification
Mi 14/AMI
Qwen → Parakeet
22.11 → 22.74
41.1 → 23.9
5
TD, high acc.
K70/Libri
Serial → TD
15.41 → 29.81
77.4 → 81.7
5
TD, low acc.
K70/KeSpeech
Serial → TD
9.62 → 15.80
44.9 → 24.3
1
TD, zero acc.
K70/M4Singer
Serial → TD
11.85 → 12.52
80.4 → 0.0
1
Prefix lookahead
K70/AMI
Serial → lookahead
9.25 → 9.41
37.0 → 37.0
2
Appendix
Table 13: Native cases compare source and scheduling choices. Arrows follow the comparison column, and n gives measured repetitions per arm. Acceptance α divides pooled accepted draft tokens by pooled proposed draft tokens. Rows retain their own runtime boundaries. Gains are not pooled across rows.
Drafter
Max. K
TPS
Gain (%)
Extra tokens
Target calls
Target-only
scalar
14.58
0.0
0
211
Parakeet CTC
3
15.44
5.9
33
178
Parakeet CTC
4
16.25
11.5
39
172
Parakeet CTC
8
15.66
7.4
43
168
SenseVoice
3
16.03
10.0
24
187
SenseVoice
4
16.03
9.9
26
185
Appendix
Table 14: Native decoding compares drafters on one 60-second AMI development window on Redmi K70 Pro. TPS pools five measured runs per configuration. Target-only uses scalar decoding. Gain is relative to the interleaved target-only control. Extra tokens and target calls are per-run medians. Extra tokens exclude the already-known first token of each admitted chain. All 35 runs agree through EOS (212 tokens).
Method
Energy (J)
Mean power (W)
Response (s)
Body tokens
J/token
Target-only
341.6
10.47
32.64
256
1.334
As 2 d
239.4
11.33
21.13
248
0.965
Appendix
Table 15: As 2 d uses less estimated phone energy on one 60s Xiaomi 14 window. Each row reports one formal execution after warmup.
Figure 9: The cooled configuration keeps pace over 600s, while the uncooled case develops backlog. Decode throughput, prefill time, and processing RTF use trailing windows of at most 15 blocks. Completion lag and queue wait show every block. Panels (a)–(e) use input-arrival time. Temperature uses elapsed wall time and includes final drain. The dotted line marks RTF 1. Both cases use 75% token retention, but source deadlines and initial thermal states differ as described above.
Metric
With cooler
Without cooler
Mean encoder / prefill / decode (s)
0.70 / 2.24 / 0.68
0.94 / 2.95 / 0.78
Mean complete block processing (s)
3.63
4.68
Active / resource RTF
0.907 / 1.007
1.171 / 1.201
Decode throughput (tokens/s)
13.87
12.23
p95 completion lag (s)
4.68
103.39
Final completion lag (s)
4.09
120.50
Appendix
Table 16: The table summarizes both complete native deployment runs. Temperature comes from the battery sensor. Source-lifetime and initial-state differences preclude cooling-only attribution.
Speculative decoding speeds up generation by letting a cheap draft propose several tokens that a target model checks in one pass. In the single-model form, the draft is a lightweight module attached to the target rather than a separate model. Applying this design to Automatic Speech Recognition (ASR) introduces an extra problem. The draft can read the whole audio at every step, yet its proposals get worse as it runs on its own. Access is not localization. The accepted text keeps the transcript position explicit, but the draft must also track the changing audio position. In the primary matched comparison, per-step audio access changes the first proposal modestly but roughly doubles later-proposal acceptance. Fixed-width windows show that the audio position explains part of this gap. A correctly placed window recovers continuation, while an equally narrow window at the wrong position reduces it. Late-draft median error reaches 21 frames in the hardest reported condition, while target attention during verification stays within a 2-frame median. We test two ways to reduce this drift. The first reads the audio position from verification attention and uses it to guide the next draft round. It saves time only when the extra accepted tokens offset the readout cost. The second is AnchorDraft, which teaches the draft to track the audio position during training without changing the inference graph. The trained draft improves end-to-end speed at both tested target scales. These results show that ASR self-speculation depends on token prediction, audio-position tracking, and draft cost.
While Large Multimodal Models excel in comprehension, high-throughput inference engines lack native support for multimodal generation. This is severe in Speech Language Models, where generating multi-layered audio tokens via decoupled AR+NAR or synchronous Multi-Token Prediction (MTP) with delay-pattern interleaving conflicts with standard single-stream loops. We present a vLLM-based inference pipeline for unified speech understanding and generation. We extend autoregressive decoding to natively execute delay-pattern de-interleaving and coordinated multi-stream sampling, integrating an on-GPU acoustic decoder for end-to-end waveform synthesis. Crucially, we overcome the shared intuition that Classifier-Free Guidance (CFG) halves throughput. By co-scheduling paired conditional and unconditional requests within a continuous batch, our CFG implementation sustains 80% of non-CFG throughput, absorbing dual-request and logit merging overheads. We open-source our framework.
Speculative decoding accelerates autoregressive generation by verifying a draft block with a target model in parallel. Under standard greedy verification, decoding stops at the first draft token that differs from the target argmax, discarding the remaining target-scored suffix. Although accepting such a mismatch changes the decoding trajectory, it can make a contiguous suffix reusable when its tokens remain target-greedy under the realized prefix. In this paper, we introduce \textbf{Approximate Speculative Decoding (ASD)}, a training-free verifier that replaces binary first-mismatch truncation with budgeted longest-prefix selection. ASD accepts selected mismatches subject to a local target-logit regret gate, a per-block exception cap, and a persistent request-level regret budget, then reuses the contiguous target-greedy suffix without additional approximate decisions or target-model forward passes. ASD requires neither a new draft model nor fine-tuning, and exactly reduces to standard greedy verification when the budget is zero. Experiments show that ASD improves fixed-workload throughput by 3.05%--15.26% over matched strict verification and averages a 7.78% gain across seven Qwen3-14B + DSpark-14B tasks. On DeepSeek-V4-Flash (284B) with DSpark it also raises verifier-side acceptance by roughly 10%--16% on GSM8K and MATH-500 in an FP4-to-FP8 compatibility setting. The source code is publicly available at: https://github.com/Kissmetothemoon/ASD
Yuannuo Feng, Zegang Peng, Yuxin Xie +5
School of Integrated Circuit Science and Engineering, Beihang University, Beijing, China · Department of Precision Instrument, Tsinghua University, Beijing, China · Faculty of Engineering, The University of Hong Kong, Hong Kong SAR, China +2