cs.LGOct 5, 2026

SchemaFill: Efficient LLM Tool Calling via Slot-Parallel Speculative Decoding

Authors: Zhi-Kai Chen, Song-Yan Li, De-Chuan Zhan, Han-Jia Ye

Organizations: School of Artificial Intelligence, Nanjing University, China · National Key Laboratory for Novel Software Technology, Nanjing University, China · Nanjing University, China

Abstract

LLM agents interact with external systems by generating structured tool calls. Given a user request, conversational context, and a catalog of tool schemas, a tool-calling model must select tools and generate their arguments, potentially producing multiple calls in a single response. Standard autoregressive decoding generates these calls token by token, incurring substantial latency for requests involving multiple calls or many argument fields. The explicit argument structure offers opportunities for parallel generation, but later argument values may depend on preceding fields and calls, so independently generated values can differ from the target model's output. We present SchemaFill, a framework for efficient LLM tool calling through slot-parallel speculative decoding. SchemaFill generates future slot values concurrently as candidates, without requiring advance knowledge of the actual call sequence or argument values. Candidates spanning multiple fields and calls are concatenated for verification by the target model under the actual output prefix. Only verified tokens are committed, and the target supplies corrections when candidates disagree. This applies target verification while exploiting parallelism across slots and calls. On Glaive and BFCL, SchemaFill achieves up to a 4.05×\times improvement in end-to-end throughput over autoregressive decoding. Code is available at https://github.com/Czzzk/SchemaFill.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. OoO-Spec: Out-of-Order Semantic Speculation for Fast Tool Calling

    Aug 1, 2026Zhiheng Zhang, Mujie Xu, Feiyu Sun +1Self-SpeculativeTool Invocation

  2. Harness Engineering in LLM Tool Use via Agent-Native Reusable Tool Primitives

    Sep 1, 2026Haibo Jin, Suijin Wang, Xucheng Yu +2Large Language Model Tool Use

  3. LLM Agents Already Know When to Call Tools -- Even Without Reasoning

    May 10, 2026Chung-En Sun, Linbo Liu, Ge Yan +2Large Language Model Tool UseLarge Language Model Agents