When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates the input size of every subsequent call. The total consumption of a task is therefore hard to predict before execution and the prediction must be revised as the run unfolds. In this paper, we propose TokenCast, which learns a composable cost representation for each execution segment, recording its own consumption and the context growth it introduces. Composing adjacent segments yields a cumulative estimate that captures the extra input cost incurred when context from earlier segments is re-read by every later call. As execution unfolds, newly observed evidence refreshes the forecast, requiring no additional LLM calls and incurring a mean cumulative prediction time of 32.8 ms per run on SWE-bench Verified. Across 4 task suites and 6 agent models, TokenCast's mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations. In offline budget-control replay, TokenCast uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion. The code is available at https://github.com/DEFENSE-SEU/TokenCast.
Figures & tables
Figure 1 : Four token-consumption prediction settings. At the call level, the input context supports a forecast before generation (1), and the output prefix updates it during generation (2). At the task level, total consumption is forecast before execution (3), and remaining consumption is updated as calls complete and context accumulates (4).
Figure 2 : Overview of TokenCast. Observed execution evidence supports call- and task-level forecasts, while the compositional path propagates predicted context growth into later-call costs.
Agent LLM
Prediction point
TokenCast (Ours)
TRAIL
EGTP
TIE
Self-Pred.
SWE-bench Verified
GPT-5.4
Task Start
144.0k
165.3k
160.9k
157.0k
152.0k
Call Start
64.5
71.1
70.3
69.5
70.9
In-call Update
38.9
78.9
74.6
77.4
80.2
Task Update
80.0k
115.0k
117.0k
119.0k
126.0k
Norm. Avg.
0.69
0.94
0.92
0.92
0.94
Table 1: MAE ↓ in tokens at four prediction points. Norm. Avg. is the macro-average of MAE normalized at each prediction point by the MAE of the corresponding history-median predictor. k denotes thousands of tokens. Best and second-best results are in bold and underlined, respectively. Marks are assigned using unrounded values.
Figure 3 : Normalized average MAE across four benchmarks and six agent LLMs. Each panel corresponds to one benchmark, and each bar group corresponds to one agent LLM. For each method, MAE is normalized by the corresponding history-median MAE at each prediction point and macro-averaged over the four prediction points. Lower is better.
Figure 4 : Trace completion under stopping limits on SWE-bench Verified. Panel (a) plots trace completion against mean tokens per run, and panel (b) plots it against replay-accounted wall time per run. Prediction overhead is included. The dashed curve is the fixed-budget baseline.
Figure 5 : Feature analysis at the four prediction points. For each point, the panels show the relative importance of selected features and contribution patterns for two representative features. Colors indicate the feature groups defined in Appendix B.1 .
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
Work
Prediction target
Evidence and estimator
Update point
TRAIL
Response length
Last-token hidden state and length bins
Before generation
EGTP
Response and remaining-response length
Hidden states and entropy pooling
Pre/mid generation
TIE
Response length distribution
Text embedding and log- t model
Before generation
Chimera
Remaining workflow output
Prompt and workflow features; quantile forest
Workflow request
Pythia
Workflow path and role output length
Historical trace profiles
Incoming request
Self-Prediction
Total task tokens
Agent inspection of task environment
Before execution
Appendix
Table 4 : Prediction targets and estimators in related work. Appendix C.3 describes our adaptations.
Model
Vendor
Access
Reasoning configuration
GPT-5.4
OpenAI
API
Low reasoning effort
Claude Opus 4.6
Anthropic
API
Adaptive thinking; low effort
Gemini 3.1 Pro
Google
API
Low thinking level
DeepSeek-V4-Pro
DeepSeek
API
Thinking enabled
Qwen3.8-27B
Alibaba
Self-hosted
Thinking enabled
Llama-3.2-3B-Instruct
Meta
Self-hosted
No reasoning mode
Appendix
Table 5: Agent LLMs used for trace collection, with access modes and reasoning configurations.
DeepSeek Harness
OpenHands
Model
Tokens
Calls
Correct
Tokens
Calls
Correct
SWE-bench Verified
GPT-5.4
245.4
21
40.7
314.5
20
36.3
Claude Opus 4.6
258.1
21
41.7
470.7
22
54.2
Gemini 3.1 Pro
333.5
21
43.5
253.5
21
47.5
DeepSeek-V4-Pro
280.9
20
34.3
336.8
28
33.8
Appendix
Table 6: Median tokens in thousands and calls among runs with complete usage, and correct runs as a percentage of all attempted runs, under each harness.
Figure 6 : Run-to-run variation in token consumption across repeated executions of the same task. The panels show the distribution across repeated runs and the within-task consumption spread as a function of task consumption, over 48 GPT-5.4 anchor tasks.
Prediction point
TokenCast
TRAIL
EGTP
TIE
Self-Pred.
History median
GPT-5.4
Task Start
144.0k
165.3k
160.9k
157.0k
152.0k
179.0k
Call Start
64.5
71.1
70.3
69.5
70.9
74.2
In-call Update
38.9
78.9
74.6
77.4
80.2
80.3
Task Update
80.0k
115.0k
117.0k
119.0k
126.0k
130.0k
Norm. Avg.
0.69
0.94
0.92
0.92
0.94
1.00
Appendix
Table 7: MAE on SWE-bench Verified. Norm. Avg. is the macro-average of MAE normalized at each prediction point by the history-median MAE. k denotes thousands of tokens. Best and second-best results are in bold and underlined, respectively. Marks are assigned using unrounded values.
Prediction point
TokenCast
TRAIL
EGTP
TIE
Self-Pred.
History median
GPT-5.4
Task Start
34.2k
37.8k
37.6k
36.4k
36.9k
44.9k
Call Start
34.5
37.3
36.5
34.0
39.8
44.0
In-call Update
31.3
39.4
38.5
36.8
39.9
42.0
Task Update
17.6k
23.1k
23.7k
23.3k
24.2k
25.0k
Norm. Avg.
0.75
0.89
0.88
0.85
0.91
1.00
Appendix
Table 8: MAE on Search-R1. Norm. Avg. is the macro-average of MAE normalized at each prediction point by the history-median MAE. k denotes thousands of tokens. Best and second-best results are in bold and underlined, respectively. Marks are assigned using unrounded values.
Prediction point
TokenCast
TRAIL
EGTP
TIE
Self-Pred.
History median
GPT-5.4
Task Start
8.0k
8.8k
8.5k
8.2k
7.4k
9.7k
Call Start
14.8
16.7
15.5
15.8
17.6
19.6
In-call Update
9.2
19.6
18.3
16.7
17.4
24.3
Task Update
4.3k
6.2k
6.5k
6.0k
5.8k
6.9k
Norm. Avg.
0.65
0.87
0.84
0.80
0.80
1.00
Appendix
Table 9: MAE on MMLU-Pro. Norm. Avg. is the macro-average of MAE normalized at each prediction point by the history-median MAE. k denotes thousands of tokens. Best and second-best results are in bold and underlined, respectively. Marks are assigned using unrounded values.
Prediction point
TokenCast
TRAIL
EGTP
TIE
Self-Pred.
History median
GPT-5.4
Task Start
32.2k
37.1k
36.0k
35.1k
32.7k
44.0k
Call Start
53.2
60.6
56.9
53.7
61.2
74.0
In-call Update
35.5
73.5
68.8
62.7
70.1
91.0
Task Update
13.9k
26.9k
24.2k
25.1k
28.2k
33.0k
Norm. Avg.
0.57
0.82
0.77
0.74
0.80
1.00
Appendix
Table 10: MAE on LongBench-v2. Norm. Avg. is the macro-average of MAE normalized at each prediction point by the history-median MAE. k denotes thousands of tokens. Best and second-best results are in bold and underlined, respectively. Marks are assigned using unrounded values.
Agent LLM
Prediction point
Mean target ( k )
TokenCast
TRAIL
EGTP
TIE
Self-Pred.
SWE-bench Verified
GPT-5.4
Task Start
475.6
30.3
34.8
33.8
33.0
32.0
Call Start
19.6
0.329
0.363
0.359
0.355
0.362
In-call Update
20.1
0.194
0.393
0.371
0.385
0.399
Task Update
286.3
27.9
40.2
40.9
41.6
44.0
Qwen3.8-27B
Task Start
611.5
31.4
37.9
34.2
35.2
36.1
Appendix
Table 11: WAPE (%) at four prediction points. Best and second-best results are in bold and underlined, respectively. Marks are assigned using unrounded values.
Figure 7 : Transfer to unseen agent LLMs with increasing amounts of target-model data. The panels report Call Start and remaining output-token prediction at Task Update. MAE is normalized by the target history median; lower is better.
Training harness
Test harness
Task Start
Call Start
In-call Update
Task Update
DeepSeek Harness
DeepSeek Harness
0.90
0.71
0.64
0.82
OpenHands
OpenHands
0.83
0.77
0.68
0.75
DeepSeek Harness
OpenHands
1.04
0.88
0.79
0.96
Appendix
Table 12: TokenCast MAE divided by the history-median MAE on the test harness. Lower is better. The two OpenHands rows use the same test runs.
Prediction point
vs. Self-Prediction ↓
vs. random split ↓
Task Start
−12.4%
+7.8%
Call Start
−15.7%
+1.9%
In-call Update
−44.9%
−2.6%
Task Update
−26.8%
+7.1%
Appendix
Table 13: Length extrapolation to tasks longer than those observed during training. MAE changes are reported relative to Self-Prediction and matched random splits.
Off
Low → Off
Prediction point
TokenCast
Self-Pred.
Direct
20 tasks
Call Start
37.1
46.8
50.7
37.9
In-call Update
22.5
45.7
33.0
23.6
Task Update
82.0k
121.0k
82.0k
76.0k
Appendix
Table 14: MAE under a GPT-5.4 reasoning-configuration shift on SWE-bench Verified at the three online prediction points. Results include evaluation within the off configuration, direct transfer from low to off, and adaptation with 20 target-configuration tasks.
Update interval
Predictions/run ↓
Time/run (ms) ↓
1 call
19.7
32.8±30.5
3 calls
6.9
11.9±10.4
5 calls
4.3
7.7±6.4
End
1.0
2.1±0.8
Appendix
Table 15: Prediction counts and cumulative overhead at online update intervals on SWE-bench Verified. Predictions/run includes the initial Task Start forecast and later Task Update refreshes.
Predictor
Norm. Avg. ↓
90% Cov. (%)
Model time (ms)
Ridge Regression
0.93
82.4
0.2
KNN
0.97
79.6
14.3
Random Forest
0.81
86.8
2.7
XGBoost
0.73
89.3
3.9
CatBoost
0.75
88.7
5.2
MLP
0.78
87.2
6.1
Appendix
Table 16: Base predictor comparison on SWE-bench Verified (GPT-5.4). Norm. Avg. is the normalized MAE averaged over the four prediction points. 90% Cov. is the empirical coverage of the 90% prediction interval. Best and second-best results in each column are in bold and underlined, respectively. Coverage is ranked high to low and other metrics low to high.
Strategy
Task Start ↓
Task Update ↓
90% Cov. (%)
Direct only
0.82
0.72
88.1
Compositional only
0.87
0.66
89.3
Direct–compositional average
0.81
0.64
89.8
Full correction pipeline
0.80
0.62
90.6
Appendix
Table 17: Effect of forecasting strategy on SWE-bench Verified (GPT-5.4). Task Start and Task Update columns report the normalized MAE at these two task-level prediction settings. Best and second-best results in each column are in bold and underlined, respectively. Coverage is ranked high to low and other metrics low to high.
Variant
Norm. Avg. ↓
Task-level ↓
Call-level ↓
Raw features
0.85
0.88
0.82
No composition
0.78
0.81
0.74
Drop g
0.76
0.79
0.72
Drop b
0.73
0.75
0.70
Full
0.69
0.71
0.68
Appendix
Table 18: Ablation of the segment representation on SWE-bench Verified (GPT-5.4). Task-level and Call-level columns report normalized MAE averaged over the two prediction points within each level. Best and second-best results in each column are in bold and underlined, respectively.
Configuration
Norm. Avg. ↓
90% Cov. (%)
Full
0.69
90.6
− Cross-fitting
0.73
88.9
− Cost weighting
0.72
89.4
− Correction
0.74
89.0
− Boundary vars
0.72
90.1
− Cross-fit. & cost-wt.
0.78
87.4
Appendix
Table 19: Component ablation on SWE-bench Verified (GPT-5.4). Each row removes one component from the full pipeline. Best and second-best results in each column are in bold and underlined, respectively. Coverage is ranked high to low and other metrics low to high.
Budget (k tokens)
Method
Trace-complete (%)
Execution (k/run)
Prediction (k/run)
Total (k/run)
Time (s/run)
Saving (%)
174
Fixed budget
30.2
154.7
0.0
154.7
113.7
–
TokenCast
30.2
101.1
0.0
101.1
76.5
34.6
TRAIL
29.5
91.0
12.5
103.5
73.0
33.1
EGTP
13.5
25.3
3.6
28.9
22.6
81.3
TIE
27.4
59.8
8.0
67.8
45.6
56.2
Self-Pred.
19.4
70.0
92.0
162.0
303.5
-4.7
Appendix
Table 20: Budget control on 288 GPT-5.4 runs from 144 SWE-bench Verified tasks. Tokens are in thousands per run, wall time is in seconds per run, and trace completion and savings are in percent. Only TokenCast matches the fixed-budget trace completion at every budget.
Method
Trace-complete (%)
Prediction tokens
Total tokens
Total time
TokenCast
59.9
0.0k
173.0k
122.7 s
TRAIL
59.5
20.7k
186.9k
124.5 s
EGTP
34.8
8.5k
75.6k
52.5 s
TIE
54.4
15.2k
139.3k
88.1 s
Self-Prediction
44.1
166.0k
300.1k
522.4 s
Appendix
Table 21: Budget-equal-weighted replay means. Trace completion is shown alongside overhead because lower processing totals can result from stopping more runs. Encoder-based baselines count locally processed input tokens; Self-Prediction counts additional LLM usage. These token counts therefore represent different resources.
Figure 8 : Trace completion and token consumption across seven budgets in the offline replay.
Figure 9 : Confirmed tokens over wall time for one run. The shaded span is call 12.
Figure 10 : Forecast updates during code repair. Left: task-level forecasts following verification outcomes. Right: current-call forecasts during edit call 12 as generation becomes visible.
Prediction point
Available evidence
Target
Forecast [90% interval]
Recorded
Task Start
Issue and configuration
T
194.638 [104.782, 333.912]
243.371
Call Start, 12
Edit request assembled
C12
13.368 [13.001, 13.912]
13.420
In-call, 12
795 bytes, old source span visible
C12
13.383 [13.201, 13.694]
13.420
In-call, 12
1,450 bytes, replacement text visible
C12
13.431 [13.407, 13.472]
13.420
Update, k=14
NumPy alias prevents reproduction
T
281.746 [191.304, 368.259]
243.371
Update, k=15
Compatibility workaround, assertions pass
T
252.381 [211.623, 298.576]
243.371
Appendix
Table 22 : Forecasts at selected points in the code-repair execution. Token quantities are in thousands. Recorded targets are shown retrospectively.
Figure 11 : Repeated input makes additional calls expensive. Left: task updates after a failed write and artifact preparation. Right: full-call Call Start forecasts from TokenCast and Self-Prediction.
Point
Newly available evidence
Sk
Forecast T [90% interval]
Task Start
Long-context question and output requirements
0.000
526.384 [264.731, 919.648]
Update, k=1
Answer generated, file creation fails
131.059
834.719 [525.193, 1204.762]
Update, k=2
Existing answer placeholder inspected
261.368
908.362 [643.881, 1268.997]
Update, k=3
Answer file successfully replaced
392.112
762.541 [652.967, 1041.863]
Update, k=5
Answer and patch read back
655.063
792.813 [765.208, 839.426]
Termination
Finish action recorded
787.465
—
Appendix
Table 23 : Long-context stages. All token quantities are in thousands. The same answer is carried through the subsequent file operations.
The wide adoption of AI agents in complex human workflows is driving rapid growth in LLM token consumption. When agents are deployed on tasks that require a significant amount of tokens, three questions naturally arise: (1) Where do AI agents spend the tokens? (2) Which models are more token-efficient? and (3) Can agents predict their token usage before task execution? In this paper, we present the first systematic study of token consumption patterns in agentic coding tasks. We analyze trajectories from eight frontier LLMs on SWE-bench Verified and evaluate models' ability to predict their own token costs before task execution. We find that: (1) agentic tasks are uniquely expensive, consuming 1000x more tokens than code reasoning and code chat, with input tokens rather than output tokens driving the overall cost; (2) token usage is highly variable and inherently stochastic: runs on the same task can differ by up to 30x in total tokens, and higher token usage does not translate into higher accuracy; instead, accuracy often peaks at intermediate cost and saturates at higher costs; (3) models vary substantially in token efficiency: on the same tasks, Kimi-K2 and Claude-Sonnet-4.5, on average, consume over 1.5 million more tokens than GPT-5; (4) task difficulty rated by human experts only weakly aligns with actual token costs, revealing a fundamental gap between human-perceived complexity and the computational effort agents actually expend; and (5) frontier models fail to accurately predict their own token usage (with weak-to-moderate correlations, up to 0.39) and systematically underestimate real token costs. Our study offers new insights into the economics of AI agents and can inspire future research in this direction.
Longju Bai, Zhemin Huang, Xingyao Wang +5
University of Michigan · Stanford University · Microsoft AI +3
LLM serving caches prompt KV state, yet most front ends still re-tokenize the full request on every call. Coding agents pay most: sessions repeatedly submit a long transcript after a small append, which can shift token boundaries near the end of the prior sequence. Across 153,951 calls the median append is ~1.4K characters; only 1.0-3.6% of calls start or rebuild a session, yet those carrymulti-million-character contexts. Fleet prompt-cache hit rate is 94.1%, and as it approaches 0.99, tokenization grows from 10% to 64% of time to first token (TTFT) in component measurements. TokTier is a stateful CPU+GPU tokenization service for this two-mode workload, under one contract: emitted token IDs are always identical to full reference tokenization. For session continuations it re-tokenizes a small window around the append and splices only when a per-request check finds a stable pre-tokenization boundary; failed checks widen the window or fall back to full reference tokenization. For calls without a reusable prefix it runs exact GPT-family regex pre-tokenization and BPE on a GPU. A sampled shadow verifier re-checks live traffic. Across 17 production tokenizer families, differential campaigns cover 1.5x10^10 split checks, a 12.4 TB real-text corpus, and 93,000+ replayed agent steps, with zero divergence. Incremental repair takes 0.5-1.1 ms from 100K to 3M characters, up to 437x faster than HF tokenization and 2.1x faster at 1M characters than the strongest cache-based baseline (Gigatoken) fully prewarmed. GPU tokenization encodes a 1M-character request in 0.87 ms, up to 491x below HF and 23.4x below the fastest published CPU method on the same protocol. With vLLM, median TTFT drops 16-34% and P99 TTFT 23% under recorded bursts. Under a 50 ms P99 objective, a four-core repair pool plus one GPU sustains 1,821 requests/s, where a 16-core stateless front end saturates at 40 requests/s.
As LLM agents are deployed in long-horizon sessions, context accumulation drives up inference costs. Existing approaches utilize text pruning or dynamic memory eviction to minimize token footprints; however, their unconstrained sequence mutations alter layouts, introducing prefix mismatches and cache invalidation. This reveals a critical trade-off between text sparsity and prompt cache continuity. To address this, we present TokenPilot, a dual-granularity context management framework. Globally, Ingestion-Aware Compaction acts as a framework harness to stabilize prompt prefixes and eliminate open-world environmental noise at the ingestion gate. Locally, Lifecycle-Aware Eviction monitors the ongoing residual utility of context segments, enforcing a conservative batch-turn schedule to offload content segments only when task relevance expires. Experiments on PinchBench and Claw-Eval under both isolated and continuous modes demonstrate that TokenPilot reduces costs by 61% and 56% in isolated mode, and 61% and 87% in continuous mode, while maintaining competitive performance compared to prior systems. TokenPilot has been integrated into LightMem2 at https://github.com/zjunlp/LightMem2.
Buqiang Xu, Zirui Xue, Dianmou Chen +12
1Zhejiang University · University of Electronic Science and Technology of China · 3Xi’an University of Electronic Science and Technology +1