cs.CLOct 1, 2026

Beyond Leaderboards: Tokenomics of Agentic Small Language Model Ensembles

Authors: Alexei N. Skurikhin, Emily M. Taylor, Nathan A. DeBardeleben

Organizations: Los Alamos National Laboratory Los Alamos, New Mexico, USA

Abstract

As large language models (LLMs) move from standalone assistants into agentic workflows, evaluation must extend beyond scalar leaderboard accuracy to account for operational reliability, cost, latency, and token efficiency. We use an agentic ensemble of small language models (SLMs) with an SLM-judge-mediated feedback loop as a case study for such beyond-leaderboard evaluation. On the 541-prompt IFEval benchmark, the best ensemble achieves 97.34% strict prompt accuracy, exceeding the strongest standalone LLM baseline, gpt-5.4, by 5.81 percentage points while operating in a lower-cost regime. We then analyze the tokenomics and operational behavior behind this gain, including cost per sample, token composition, useful-output goodput, feedback-loop recovery, latency decomposition, and performance across instruction categories and constraint counts. Our results show that agentic SLM ensembles can trade additional test-time tokens and orchestration overhead for improved instruction-following fidelity, motivating multi-dimensional evaluation protocols for future agentic AI systems.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems

    May 28, 2026Lorenz Kutschka, Bernhard GeigerAgentic BenchmarksKeyed Javascript Object Notation Record Yields

  2. Small Language Models are the Future of Agentic AI

    Jun 2, 2025Peter Belcak, Greg Heinrich, Shizhe Diao +5Production Agentic Systems

  3. TokenCast: Forecasting Token Consumption During LLM Agent Execution

    Sep 28, 2026Chaoqian Ouyang, Ling Yue, Libin Zheng +7