cs.SDOct 6, 2026

Hiding Tool Latency in On-Device Cascaded Voice Agent through Speculative Execution

Authors: Kyudan Jung, Hyunsin Park, Yoonhyung Lee, Jinhwan Park, Jinhyeok Yang, KiHyun Nam, Jaegul Choo, Jinkyu Lee

Organizations: Qualcomm AI Research · KAIST AI

Abstract

Tool-augmented speech assistants typically serialize automatic speech recognition, large language model inference, and external tool execution. As a result, tool latency is incurred only after the user has finished speaking and the LLM has identified the required tool calls. We present speculative tool execution for on-device cascaded voice agents, which predicts tool requests from partial ASR hypotheses and initiates tool execution while speech is still being received, thereby reducing end-to-end response latency. Our approach introduces a Predictor module that anticipates tool calls during speech recognition, executes them speculatively, and caches the results. The cached outputs are then injected into the LLM prompt, enabling faster responses. Additionally, to mitigate errors caused by user self-corrections during speech, we employ a rule-based validation mechanism that selectively injects only valid cached results. As a final safeguard, the LLM retains the ability to issue tool calls directly, ensuring that the latency of our framework is upper-bounded by the baseline serial execution pipeline in the worst case. We evaluate our method using live measurements from a fully implemented Android voice assistant. Our approach reduces the median time-to-first-audio from 5.79,s to 4.60,s and decreases the standard deviation from 3.49,s to 2.81,s, resulting in more predictable response latency.

Figures & tables

Explore similar work

CardsList
  1. Endpoint Anticipation for Low-Latency Spoken Dialogue

    Jun 11, 2026Sathvik Udupa, Shinji Watanabe, Petr Schwarz +1Turn-TakingSpeech Language Models

  2. AOSpec: Action and Observation Co-Speculation for Low-Latency Agent Serving

    Aug 1, 2026Hao Mark Chen, Jinnan Guo, Wayne Luk +1Self-SpeculativeSmooth Execution

  3. Speculative Interaction Agents: Building Real-Time Agents with Asynchronous I/O and Speculative Tool Calling

    May 13, 2026Coleman Hooper, Minwoo Kang, Suhong Moon +7Agentic Reasoning