cs.LGSep 28, 2026

TokenCast: Forecasting Token Consumption During LLM Agent Execution

Authors: Chaoqian Ouyang, Ling Yue, Libin Zheng, Hanghui Guo, Shengxiang Xu, YiShu Wang, Ran Li, Jian Yin, +2 more

Organizations: Sun Yat-Sen University · Rensselaer Polytechnic Institute · Southeast University · Hong Kong University of Science and Technology

Abstract

When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates the input size of every subsequent call. The total consumption of a task is therefore hard to predict before execution and the prediction must be revised as the run unfolds. In this paper, we propose TokenCast, which learns a composable cost representation for each execution segment, recording its own consumption and the context growth it introduces. Composing adjacent segments yields a cumulative estimate that captures the extra input cost incurred when context from earlier segments is re-read by every later call. As execution unfolds, newly observed evidence refreshes the forecast, requiring no additional LLM calls and incurring a mean cumulative prediction time of 32.8 ms per run on SWE-bench Verified. Across 4 task suites and 6 agent models, TokenCast's mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations. In offline budget-control replay, TokenCast uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion. The code is available at https://github.com/DEFENSE-SEU/TokenCast.

Figures & tables

Appendix figures & tables26 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks

    Apr 24, 2026Longju Bai, Zhemin Huang, Xingyao Wang +5Coding AgentsArtificial Intelligence Agents

  2. TokTier: Exact Stateful CPU+GPU Tokenization for Agentic LLM Serving

    Jul 31, 2026Zhenyu Zhang, Zhichao CaoTime-To-First-TokenLarge Language Model Serving

  3. TokenPilot: Cache-Efficient Context Management for LLM Agents

    Jun 15, 2026Buqiang Xu, Zirui Xue, Dianmou Chen +12LLM Inference OptimizationLarge Language Model Memory