Skip to content
AI.info

Research

Hiding Tool Latency in On-Device Cascaded Voice Agent through Speculative Execution

Hiding Tool Latency in On-Device Cascaded Voice Agent through Speculative Execution Overview Research area: On-device speech systems — specifically cascaded automatic speech recognition (ASR), large l

Hiding Tool Latency in On-Device Cascaded Voice Agent through Speculative Execution
arXiv
2610.07641
Published
2026-10-06
Authors
Kyudan Jung, Hyunsin Park, Yoonhyung Lee, Jinhwan Park, Jinhyeok Yang, KiHyun Nam, Jaegul Choo, Jinkyu Lee

AI summary

Hiding Tool Latency in On-Device Cascaded Voice Agent through Speculative Execution

Overview

Research area: On-device speech systems — specifically cascaded automatic speech recognition (ASR), large language model (LLM), and text-to-speech (TTS) voice assistants that call external tools (e.g., remote web search).

Technical level: Intermediate. The paper assumes familiarity with ASR/LLM/TTS pipelines and tool calling, but its core idea is explained with concrete timing diagrams and consumer-device measurements.

Scope: The paper proposes and evaluates a speculative tool-execution scheme that predicts and launches tool calls from partial ASR hypotheses, hiding tool latency inside the user's speaking time on a commercial Android device.

What This Paper Is About

Conventional cascaded voice assistants run ASR, LLM inference, and external tool execution one after another, so any remote tool latency (such as a web search) is only incurred after the user stops speaking and the LLM decides which tool to call. The authors' goal is to overlap that tool latency with the remaining speech, so the user waits less and the wait time becomes more consistent. They build a full Android voice assistant and measure how much of the tool delay can be hidden without hurting correctness.

Key Contributions

  1. A speculative tool-execution framework for on-device cascaded voice agents. A Predictor module watches partial ASR hypotheses, anticipates likely tool calls, executes them speculatively, and stores the outputs in a speculative results cache.
  2. Direct injection of cached results into the LLM prompt. Rather than waiting for the LLM to issue a tool call, cached speculative results are inserted into the prompt before decoding, so the model can use information gathered while the user was still speaking.
  3. A rule-based validation and verified-fallback mechanism. Clause-level segmentation plus selective injection handles user self-corrections, and if the LLM still issues a tool call, the system first checks the cache (Figure 2, path c) before falling back to normal tool execution (path a). This keeps worst-case latency upper-bounded by the baseline serial pipeline.
  4. Live, full-system evaluation on a commercial Android device, including an ablation study, hard-negative and adversarial robustness sets, and an offline comparison against an LLM-based predictor.

Main Findings

  • Median latency reduced: Median time-to-first-audio (TTFA), measured from the end of speech input to the first TTS audio output, fell from 5.79 s in the cascaded baseline to 4.60 s with the speculative system.
  • More predictable responses: Standard deviation dropped from 3.49 s to 2.81 s, and p99 TTFA fell from 19.76 s to 15.10 s.
  • Tool-call timing improved: FT p99 (time from ASR completion to the first tool call) fell from 22.44 s to 11.71 s.
  • Higher tool-selection accuracy: Tool-calling F1 rose from 44.8% to 60.6%; the authors hypothesize the small 3B model benefits from additional task-relevant context in the prompt.
  • Illustrative case: For a 15-second query requiring approximately 10 seconds of search-tool execution, TTFA was reduced from 17 seconds to 8 seconds.
  • Ablation — direct cache injection is the main driver: It reduced mean latency from 6.69 s to 6.17 s and raised F1 from 44.8% to 63.0%. Verified fallback alone left latency and F1 close to baseline (6.53 s mean, 43.5% F1) with only a 10.2% cache-ready rate.
  • Ablation — combining both gives the best overall balance: Direct injection plus verified fallback achieved the lowest median (4.60 s) and mean (5.99 s) latency, the highest cache-ready rate (47.7%), and the fewest wasted speculative calls (5), though its p95 (11.39 s) and F1 (60.6%) were slightly worse than direct injection alone.
  • Robust to hard negatives: On 22 hard negatives containing search-related expressions that require no search, the false read rate was 9.1%, with 2 speculative calls and 2 wasted calls. F1 is not reported for this set because no ground-truth search requests exist.
  • Robust to adversarial corrections: On 24 adversarial examples with corrections, retractions, late intents, or topic changes, F1 was 90.2% with a 12.3% false read rate, and 21 speculative calls of which 21 were wasted — errors that do not modify alarms, calendars, or other user data because speculation is restricted to read-only tools.
  • The LLM based router is more accurate offline but not practical: In an offline comparison, the rule-based predictors scored 37.5% F1 (word prefix), 36.5% (clause prefix), and 39.4% (full utterance), while the LLM predictor reached 40.0% (prefix 50%), 51.9% (prefix 75%), and 59.9% (full utterance). The authors note these offline numbers exclude inference latency, and that on-device LLM prediction is unlikely to finish before ASR finalization, so they favor the lightweight rule-based predictor.
  • Design parameters: The predictor fires after a minimum of 6 words, after every 5 additional words, or when a pause longer than 300 ms is detected; clause topics are drawn from a fixed set of 12 coarse-grained information-seeking categories.
  • Residual tool calling: In the evaluation set the LLM still issued tool calls in 13% of cases, and 7% of cases redundantly requested tools for information already provided.

Methodology in Plain English

The researchers start from a timing model of the standard pipeline: the user starts speaking at t0, ASR finalizes the transcript, the LLM decides whether and which tool to call, the tool runs, a second LLM pass writes the reply, and TTS speaks it. The time they want to eliminate is the tool execution minus the ASR completion — dead time that currently sits after the user has stopped talking.

Their fix is to guess early. While ASR is still streaming, a lightweight rule-based Predictor checks the partial transcript at fixed trigger points. It splits the partial hypothesis into clauses using punctuation and coordinating words like "and," "but," "also," "then," and "while," and each clause is independently mapped to a candidate tool call, with a coarse topic label. This clause-level view lets multiple independent information-seeking intents in one utterance be dispatched in parallel. Predicted calls are executed immediately and their results go into a speculative results cache.

Once ASR finalizes, cached speculative results are injected straight into the LLM prompt so the model can answer using information that was fetched while the user was still speaking. If the LLM nevertheless asks for a tool, the system first checks whether a matching result is already cached and reuses it; otherwise it runs the tool through the normal path. This fallback is what bounds worst-case latency at the serial baseline.

The evaluation is fully live rather than simulated. The dataset has 100 manually written English utterances covering single- and multi-search requests, calendar and alarm requests, and requests needing no tool, plus 22 hard negatives and 24 adversarial examples, for 146 samples total. Each utterance is synthesized with Gemini-2.5-Flash-Preview-TTS and streamed into the device in real time; the dataset is approximately 1.5 hours of audio and a full run takes approximately 3 hours including TTS generation and live network calls. The device is a Samsung Galaxy S26 Ultra (SM-S948N, Snapdragon Elite Gen 5), with a streaming Conformer ASR model and TTS on the CPU, a quantized Llama-3.2-3B QAIRT bundle (W4A16) on the NPU via the Qualcomm QNN stack, and CPU-side audio I/O, predictor, cache management, and HTTP calls to the EXA remote web-search API. Metrics include TTFA, precision/recall/F1 and exact accuracy for tool requests, false read rate, wasted calls, and cache-ready rate.

Why This Matters

Impact on research. The paper treats streaming speech not only as an input modality but as an opportunity for anticipatory computation, extending speculative execution ideas into the cascaded ASR-LLM-TTS setting — a pipeline that, unlike end-to-end speech systems, cannot initiate tool calls during generation. It also reports an accuracy gain, not just a latency gain, from injecting retrieved context into a small 3B model's prompt.

Real-world applications:

  • Voice assistants on phones and wearables where remote tool latency (web search, API calls) is directly perceived by users.
  • Smart speakers and in-car assistants where predictable response timing matters as much as raw speed.
  • Accessibility tools and hands-free assistants, where the delay between finishing a request and hearing an answer affects usability.
  • Any on-device agent with read-only, idempotent tool calls where speculative prefetching is safe.

Industry relevance. The results are measured on a shipped-class Android device with NPU-accelerated LLM inference, and the design explicitly balances latency, cache reuse, and wasted network calls — the cost of speculation in production. Because speculative calls are restricted to read-only tools, mispredictions do not alter user data such as alarms or calendars.

Future Directions

  • Predictor accuracy versus cost: The offline LLM-based router scored up to 59.9% F1 (full utterance) versus 39.4% for the rule-based full-utterance router, but those numbers exclude inference latency. Determining whether a partially run LLM predictor can reliably finish before ASR finalization remains open.
  • Reducing wasted calls: The adversarial set produced 21 speculative calls of which 21 were wasted, and the final configuration still logged 5 wasted calls in the ablation, so better self-correction detection could cut network cost.
  • Expanding beyond read-only tools: The safety argument rests entirely on speculation being limited to read-only tools; the paper does not report results for speculating on state-changing actions.
  • Broadening the evaluation set: The authors note that no public dataset exists for this task, and their constructed set of 100 normal utterances plus 22 hard negatives and 24 adversarial examples leaves room for wider, multilingual, and more varied benchmarking.

Target Audience

Researchers and engineers working on on-device speech assistants, cascaded ASR-LLM-TTS pipelines, and LLM tool calling; mobile/embedded ML practitioners concerned with latency, NPU resource contention, and network cost; and product teams building voice agents who need concrete evidence about how much remote tool latency can realistically be hidden on a phone.

Authors’ abstract

Tool-augmented speech assistants typically serialize automatic speech recognition, large language model inference, and external tool execution. As a result, tool latency is incurred only after the user has finished speaking and the LLM has identified the required tool calls. We present speculative tool execution for on-device cascaded voice agents, which predicts tool requests from partial ASR hypotheses and initiates tool execution while speech is still being received, thereby reducing end-to-end response latency. Our approach introduces a Predictor module that anticipates tool calls during speech recognition, executes them speculatively, and caches the results. The cached outputs are then injected into the LLM prompt, enabling faster responses. Additionally, to mitigate errors caused by user self-corrections during speech, we employ a rule-based validation mechanism that selectively injects only valid cached results. As a final safeguard, the LLM retains the ability to issue tool calls directly, ensuring that the latency of our framework is upper-bounded by the baseline serial execution pipeline in the worst case. We evaluate our method using live measurements from a fully implemented Android voice assistant. Our approach reduces the median time-to-first-audio from 5.79,s to 4.60,s and decreases the standard deviation from 3.49,s to 2.81,s, resulting in more predictable response latency.

Read the original paper