Retrieval-based speculative decoding (SD) drafts tokens by copying continuations from existing text, which suits coding agents that repeatedly reproduce code, logs, and earlier attempts. Yet existing methods fall short in agent pipelines: much of the reusable text is missing from their corpora or stored in a form that differs from what the agent emits, and their draft lengths ignore that accept length varies across agents and drifts over turns. We present AgSpec, a framework that supplies the corpus and draft-length policies that existing retrieval engines lack in coding-agent pipelines. AgSpec retrieves from session, workspace, and global corpora, retaining the ongoing session trajectory and indexing opened files in the agent's emission format. It bounds each agent's draft length with an offline-profiled cap and adapts the length online from verification feedback. On two repository-level multi-agent coding benchmarks, AgSpec outperforms five retrieval-based drafters and EAGLE-3 in most evaluated settings, raising generation throughput over autoregressive decoding up to 4.37× at batch size 1 and 4.76× at batch size 16. AgSpec also remains effective on benchmarks without a repository or a multi-agent pipeline, showing that its gains generalize to coding agents broadly.
Figures & tables
Figure 1: (a) Accept length when restoring availability, then matchability. (b) Achievable accept length for Gemma3-27B, taking the best available continuation at each step. Markers are medians and bars the interquartile range. (c) Decoding throughput of SAM-Decoding and AgSpec -SAM relative to no speculation for Gemma3-27B. All data is from SWE-bench Verified.
Figure 2: The same code in two forms: as the workspace holds it, and as a unified-diff patch emits it. Red marks the prefix each copied line gains.
Figure 3: Overview of AgSpec . Retrieval consults the session, workspace, and global corpora. Verification extends the output, updates the session corpus, and provides feedback to the length controller. The workspace corpus grows as artifacts are opened and is discarded when the session ends.
Figure 4: Main comparison across two repository benchmarks and three models. Rows show generation throughput and AL at batch sizes 1 and 16 ; higher is better. The number above each bar is its ratio to the no-speculation bar in that panel. no spec. denotes AR generation.
Figure 5: Non-repository and single-agent evaluation at batch size 1 across three models. LiveCodeBench evaluates AgSpec without repository context, using the global reference and session corpora with an empty workspace corpus. Terminal-Bench evaluates AgSpec when agent functions are not represented by separate explicit roles.
Figure 6: Corpus ablation on SWE-bench Verified at batch size 1 . (a) AL as corpora are added; the number above each bar is its ratio to G alone, and Ccall is a per-call corpus of the current prompt and output. (b) Coder AL gain per turn from indexing workspace artifacts in emission form, relative to raw file text only.
Figure 7: Draft-length control on SWE-bench Verified with Devstral. (a) Generation throughput (top) and AL (bottom) of fixed draft lengths K (every agent capped at K , no online scale) and AgSpec , normalized to K=8 within each group. (b) Mean draft tokens per drafting step by turn at batch size 1 , split into accepted (solid) and rejected (hatched) tokens. (c) Fixed K=32 without and with the online scale enabled, at batch size 16 : throughput (left) and case study of coder call per decoding step (right), with accepted tokens solid and rejected tokens lightly shaded.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Hardware
One NVIDIA RTX PRO 6000 Blackwell Server Edition (96 GB) per run.
Serving
vLLM 0.17.1 (V1 engine), TP1, eager execution, prefix caching off, chunked prefill on, and maximum model length 16,000 .
Models
Devstral-Small-2507 AWQ 4-bit, Gemma3-27B-it W4A16, and Qwen3.6-27B AWQ INT4, with greedy decoding.
Global corpus G
One fixed corpus per model and benchmark, constructed before serving.
Session corpus Ss
Trailing 2,048 prompt tokens and generated tokens, shared within a session and reset afterward.
Workspace corpus Wr
Current-session repository artifacts in raw and emission forms, reset after the session.
Retrieval
Minimum match length 1 and margin δ=5 against G . G and Ss use frequency; Wr uses the first occurrence.
Appendix
Table 1: Configuration of the reported AgSpec experiments.
Method
K
Settings
PLD
10
n -gram length 3 – 5 over the prompt
REST
10
Match threshold 2 ; per-model datastore
FastCoder
10
Match threshold 2 ; same datastore as REST
SAM-Decoding
15
Match threshold 2 ; 1+15 verified tokens per step
SuffixDecoding
24
Tree depth 24 , speculation factor 1.0 , minimum token probability 0.1
EAGLE-3
3
Chain depth 3 ; per-model drafter checkpoint
Appendix
Table 2: Baseline draft lengths and engine settings.
Benchmark
Model
λ
1−λ
τ used
Kmax used
TeamBench
Devstral
0.022
0.979
0.944
4 / 14 / 19
Gemma3
0.133
0.867
0.944
4 / 24 / 11
Qwen3.6
0.051
0.949
0.896
15 / 3 / 7
SWE-bench Verified
Devstral
0.299
0.701
0.950
24 / 24 / 24 / 24 / 24
Gemma3
0.299
0.701
0.950
24 / 24 / 24 / 24 / 24
Qwen3.6
0.299
0.701
0.950
24 / 24 / 24 / 24 / 24
Appendix
Table 3: Break-even threshold re-measured on the final batch-size-16 runs (steps with b≥12 ), with the τ and per-agent caps Kmax used in the reported runs. Caps are listed in agent order: planner / executor / verifier for TeamBench, and localizer / coder / executor / feedback / summarizer for SWE-bench Verified; on SWE-bench Verified only the first three are set explicitly and the other two default to the engine budget. On SWE-bench Verified, all three models use the λ measured with Devstral.
Control
Devstral
Gemma3
Qwen3.6
Per-call session corpus
−51.4%∗∗∗
−44.2%∗∗∗
−39.3%∗∗∗
Per-agent session corpus
−2.3%∗∗∗
−1.7%∗∗∗
−1.4%∗∗∗
Cross-session workspace growth
−0.2%
−0.5%
+0.3%
Workspace pooled into G
+0.2%
0.0%
+0.5%
Appendix
Table 4: Additional corpus controls on SWE-bench Verified at batch size 1 . Values are AL changes relative to the reported configuration. Asterisks mark a two-sided Wilcoxon signed-rank test over paired sessions ( ∗p<0.05 , ∗∗p<0.01 , ∗∗∗p<0.001 ; no multiple-comparison correction).
δ
Devstral
Gemma3
Qwen3.6
0
93.8∗∗
93.3∗∗∗
90.8∗∗∗
2
95.0∗∗
94.7∗∗∗
93.6∗∗
5†
94.6
94.1
93.3
8
94.3∗∗∗
93.8∗∗∗
93.2∗∗∗
12
94.2∗∗∗
93.7∗∗∗
93.2∗∗∗
∞
94.1∗∗∗
93.6∗∗∗
93.2∗∗∗
Appendix
Table 5: Preference-margin sensitivity on SWE-bench Verified ( b=1 ). Values are oracle-normalized continuation agreement (%).
Rule
Global G
Session Ss
Workspace Wr
first occurrence
−18.1∗∗∗
−4.2∗∗∗
0†
Most recent
−7.9∗∗∗
+1.1∗∗∗
−2.6∗∗
Most frequent
0†
0†
−1.0
REST-style trie
+6.5∗∗∗
+0.1
+0.3
Recency–frequency
−0.8∗∗∗
+0.6∗∗∗
−1.1
Appendix
Table 6: Continuation-selection replay on SWE-bench Verified. Values are changes in reference-matching continuation length (%) relative to the reported rule, averaged across three models.
Figure 8: Draft-token composition at batch size 16 over decoding steps with at least 12 active requests. Top: mean accepted (solid) and rejected (hatched) draft tokens per drafting step; labels report rejection rates. Bottom: the ratio of accepted to rejected draft tokens, on a log scale, with the dotted line at equal counts. No-speculation is omitted.
Figure 9: Speed-up over no speculation by turn at batch size 1 on SWE-bench Verified and TeamBench.
Figure 10: Speed-up over no speculation by agent and turn at batch size 1 . TeamBench results are geometric means across available models.
Benchmark
Model
Resident
On disk
Load
Tokens
SWE-bench Verified
Devstral
3,603 MB
214 MB
12.4 s
–
Gemma3
4,121 MB
253 MB
14.1 s
–
Qwen3.6
4,295 MB
258 MB
14.7 s
–
TeamBench
Devstral
2,484 MB
144 MB
9.2 s
1.94 M
Gemma3
3,180 MB
189 MB
11.5 s
2.51 M
Qwen3.6
3,640 MB
214 MB
14.0 s
2.94 M
Appendix
Table 7: Global-corpus memory footprint and load time. Token counts are reported for TeamBench.
Benchmark
Model
Engine
AgSpec
Rest
AgSpec share
SWE-bench Verified
Devstral
SAM
5.05 ms
98.1 ms
4.9%
Suffix
2.80 ms
78.7 ms
3.4%
Gemma3
SAM
5.81 ms
132.8 ms
4.2%
Suffix
2.57 ms
142.6 ms
1.8%
Qwen3.6
SAM
2.73 ms
162.6 ms
1.7%
Suffix
2.54 ms
179.4 ms
1.4%
Appendix
Table 8: Per-step split at batch size 16 , over steps with at least 12 active requests. AgSpec is the proposer: corpus queries, draft selection, and appending the step’s tokens. Rest is the target model and the engine.