cs.LGAug 11, 2026

Lifecycle-Optimal Tokenization: Vocabulary Size as a Deployment-Regime-Dependent Infrastructure Parameter

Authors: Rima MittalAnkit GubraniSatyanarayana Kakollu

Abstract

Tokenizer vocabulary size is a foundational design choice in large language model (LLM) infrastructure, yet it is typically fixed at training time based on convention rather than deployment analysis. We show that the cost-optimal vocabulary is not a constant but a function of the serving regime. We formalize total deployment cost as Clifecycle(V)=Ctrain(V)+λCinfer(V,B)C_{lifecycle}(V) = C_{train}(V) + λ\cdot C_{infer}(V, B), where λλ is inference volume and BB is the serving batch size. Through controlled experiments on two GPU families spanning the memory-bound to compute-bound regimes (A10G, ridge \approx 117 FLOP/byte; A100, ridge \approx 183 FLOP/byte), we demonstrate: (1) the inference-optimal vocabulary shifts 16x with serving batch, from 32k at B=1B=1 to 524k at B=64+B=64+, driven by amortization of the V×dV \times d unembedding matrix read; (2) at 1.3-2.3B model scale, quality (bits per byte, BPB) is optimized at V=65V=65k, confirming scale-dependent vocabulary preference; (3) the lifecycle-optimal vocabulary diverges from training-optimal by up to 16x for production deployments. Quality is approximately invariant across the optimal range (<<2% BPB spread), making vocabulary a pure systems optimization with no quality penalty in the measured range. Our results provide actionable capacity planning guidance: on-device deployments (B=1B=1) should use V32V \approx 32k; datacenter serving (B64B \geq 64, λ10λ\geq 10) should use V131V \approx 131-262k.

Explore similar work

CardsList
  1. Compute Optimal Tokenization

    May 2, 2026Tomasz Limisiewicz, Artidoro Pagnoni, Srini Iyer +6Token CompressionLarge Models