As large language models (LLMs) move from standalone assistants into agentic workflows, evaluation must extend beyond scalar leaderboard accuracy to account for operational reliability, cost, latency, and token efficiency. We use an agentic ensemble of small language models (SLMs) with an SLM-judge-mediated feedback loop as a case study for such beyond-leaderboard evaluation. On the 541-prompt IFEval benchmark, the best ensemble achieves 97.34% strict prompt accuracy, exceeding the strongest standalone LLM baseline, gpt-5.4, by 5.81 percentage points while operating in a lower-cost regime. We then analyze the tokenomics and operational behavior behind this gain, including cost per sample, token composition, useful-output goodput, feedback-loop recovery, latency decomposition, and performance across instruction categories and constraint counts. Our results show that agentic SLM ensembles can trade additional test-time tokens and orchestration overhead for improved instruction-following fidelity, motivating multi-dimensional evaluation protocols for future agentic AI systems.
Figures & tables
Figure 1. Strict prompt accuracy ( μ±σ ) on the IFEval benchmark ( N=541 ). For ensembles, identifiers denote the judge model in parentheses followed by the ensemble members. Horizontal dashed lines and bold labels denote the performance winners in each category, highlighting a gain of 5.81 percentage points for the proposed ensemble model over the strongest single-model baseline. strict prompt accuracy prompt loose and instruction level accuracy
Figure 2
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Large Language Model (LLM)
Small Language Model (SLM)
gpt-5.4
gpt-5.4-mini
gemini-3.1-pro-preview
gemini-3.1-flash-lite-preview
claude-sonnet-4-6
claude-haiku-4-5
llama-4-maverick
llama-4-scout
mistral-large-3
mistral-small-4
Ensemble composition: [Ens-n] (Judge model) ensemble member names
Appendix
Table 1. Evaluated large and small language models (LLMs and SLMs) and ensembles of SLMs.
Figure 7. Accuracy ( μ±σ ) breakdown by IFEval instruction general categories. The breakdown is partitioned into (a) and (b). We show the top two standalone baselinesand two leading ensembles. The ensemble indices ( [Ens-n] ) match the ensembles in Figure 1. strict accuracy per category
Know Center Research GmbH / Sandgasse 34, A-8010 Graz · Signal Processing and Speech Communication, Graz University of Technology / Inffeldgasse 16c, A-8010 Graz · Graz Center for Machine Learning // A-8010 Graz