Tool-augmented speech assistants typically serialize automatic speech recognition, large language model inference, and external tool execution. As a result, tool latency is incurred only after the user has finished speaking and the LLM has identified the required tool calls. We present speculative tool execution for on-device cascaded voice agents, which predicts tool requests from partial ASR hypotheses and initiates tool execution while speech is still being received, thereby reducing end-to-end response latency. Our approach introduces a Predictor module that anticipates tool calls during speech recognition, executes them speculatively, and caches the results. The cached outputs are then injected into the LLM prompt, enabling faster responses. Additionally, to mitigate errors caused by user self-corrections during speech, we employ a rule-based validation mechanism that selectively injects only valid cached results. As a final safeguard, the LLM retains the ability to issue tool calls directly, ensuring that the latency of our framework is upper-bounded by the baseline serial execution pipeline in the worst case. We evaluate our method using live measurements from a fully implemented Android voice assistant. Our approach reduces the median time-to-first-audio from 5.79,s to 4.60,s and decreases the standard deviation from 3.49,s to 2.81,s, resulting in more predictable response latency.
Figures & tables
Figure 1: Timing block diagram of speculative tool calling. Unlike the cascaded baseline, our method hides part or all of the tool-calling latency during ASR. The figure illustrates a case where one of two required tool calls results in a cache hit.
Figure 2: Overview of our methodology. During streaming ASR processing, the Predictor anticipates potential tool calls and retrieves relevant information via a Search API. The retrieved results are stored in the Speculative Results Cache and directly injected into the LLM through path (b). When the LLM later issues the corresponding tool call, the system first checks the cached results through path (c) for verification. As a fallback, the standard tool-calling workflow is executed through path (a), corresponding to the conventional LLM tool invocation process. In our setup, the ASR model runs on the CPU, while the W4A16 quantized LLM runs on the NPU.
System
p50
p99
FT p99
Mean ± Std
F1
(s)
(s)
(s)
(s)
(%)
Cascaded
5.79
19.76
22.44
6.69 ± 3.49
44.8
Speculative (Ours)
4.60
15.10
11.71
5.99 ± 2.81
60.6
Table 1: Main results. Cascaded denotes the non-speculative baseline. TTFA is measured from the end of speech input to the first TTS audio output. FT denotes the time from ASR completion to the F irst T ool call.
Variant
p50
p95
Mean ± Std
F1
Ready
Wasted
(s)
(s)
(s)
(%)
(%)
calls
cascaded
5.79
11.92
6.69 ± 3.49
44.8
0.0
0
speculative prefetch only
6.03
11.19
6.42 ± 2.69
44.8
0.0
47
verified fallback
6.05
11.35
6.53 ± 2.97
43.5
10.2
38
direct cache injection
5.72
11.05
6.17 ± 2.68
63.0
46.6
6
direct injection + verified fallback
4.60
11.39
5.99 ± 2.81
60.6
47.7
5
Table 2: Ablation study. Components are added incrementally to the cascaded baseline. Ready measures cache availability when a search result is needed, and Wasted counts speculative tool calls whose outputs are never consumed.
Set
n
p50
F1
False read
Spec calls
Wasted
(s)
(%)
(%)
calls
Hard negatives
22
6.24
N/A
9.1
2
2
Adversarial
24
8.40
90.2
12.3
21
21
Table 3: Robustness diagnostics, with key outcomes in bold. N/A denotes an inapplicable metric.
Router input
F1 ↑
Exact count ↑
False read ↓
Fire position ↓
(%)
(%)
(%)
(%)
Rule word prefix
37.5
49.0
22.2
50.6
Rule clause prefix
36.5
49.0
22.2
70.6
Rule full utterance
39.4
50.0
22.2
100.0
LLM prefix 50%
40.0
67.0
8.9
49.8
LLM prefix 75%
51.9
68.0
22.2
74.8
Table 4: Offline router accuracy reference. Fire position is the mean fraction of transcript words observed before the first prediction. Bold is the best value per column.
While low-latency interaction is critical for spoken dialogue, cascaded architectures are often bottlenecked by reactive turn-completion detection. We propose Endpoint Anticipation, shifting from reactive detection to proactive forecasting of end-of-turn signals. Our speech-based model anticipates endpoints upto 2.56 seconds in advance, enabling speculative execution of LLM and TTS pipelines on partial context. We introduce metrics to quantify the trade-off between realized latency reduction and computational redundancy. Evaluation across conversational and task-oriented datasets shows our model consistently outperforms competitive VAP-based baselines. Integration with the Unmute framework demonstrates a 505 ms average latency reduction with a 28.4% increase in speculative computation, effectively masking sequential bottlenecks to enable complex reasoning in real-time speech-to-speech interaction.
Sathvik Udupa, Shinji Watanabe, Petr Schwarz +1
Brno University of Technology, Czechia · Carnegie Mellon University, United States
Large language model agents increasingly act through stateful tools, yet model generation and environment execution remain serialized at every step. As decoding accelerates, tool execution becomes a growing bottleneck. Existing action- or observation-only speculation leaves much of this latency exposed: value is concentrated in a few slow calls, some outcomes emerge only through execution, and longer lookahead typically requires an increasingly unlikely chain of action predictions. We present AOSpec, a lossless framework that co-speculates actions and observations across the full agent-environment loop. Expected Value Decoding (EVD) directs observation speculation toward outcomes with the greatest expected latency benefit, optimizing expected time hidden rather than hit rate. For outcomes only execution can reveal, AOSpec launches latency-critical target actions in isolated forks that contain their effects, while Joint Action-State Verification (JASV) verifies both the action and its origin state against committed execution before reuse. JASV recasts long-horizon action dependency from full-chain prediction into target action-state verification, breaking the lookahead--accuracy tradeoff and unlocking long-range overlap without sacrificing serial semantics. Across Terminal-Bench serving settings spanning four harnesses, five actor models, and five serving speeds, AOSpec outperforms every practical baseline, reducing mean end-to-end latency by 11.8-32.5% and p99 latency by up to 42.8%. Its gains increase as decoding accelerates, and its observation model transfers from Terminal-Bench to SWE-bench Verified without retraining.
There is a growing demand for agentic AI technologies for a range of downstream applications like customer service and personal assistants. For applications where the agent needs to interact with a person, real-time low-latency responsiveness is required; for example, with voice-controlled applications, under 1 second of latency is typically required for the interaction to feel seamless. However, if we want the LLM to reason and execute an agentic workflow with tool calling, this can add several seconds or more of latency, which is prohibitive for real-time latency-sensitive applications. In our work, we propose Speculative Interaction Agents to enable real-time interaction even for agents with complex multi-turn tool calling. We propose Asynchronous I/O, which decouples the core agent reason-and-act thread from waiting for additional information from either the user or environment, thereby allowing for overlapping agentic processing while waiting on external delays. We also propose Speculative Tool Calling as a method to manage task execution when the agent is still unsure if it has received the full information or if additional user information may later be provided. For strong cloud models, our method can be applied out-of-the-box to existing real-time cloud APIs, providing 1.3-1.7× speedups with minor accuracy loss. To enable real-time interaction with small edge-scale models, we also present a clock-based training methodology that adapts the model to handle streaming inputs and asynchronous responses, and demonstrate a synthetic data generation strategy for SFT. Altogether, this approach provides 1.6-2.2× speedups with the Qwen2.5-3B-Instruct and Llama-3.2-3B-Instruct models across multiple tool calling benchmarks.