cs.AIAug 2, 2026

Perspectives on Tsallis Statistics for Artificial Intelligence

Authors: Kleyton da CostaBernardo Modenesi

Organizations: University College London, London, United Kingdom · Division of Biostatistics & School of Computing, University of Utah, United States

Abstract

Tsallis statistics generalizes Boltzmann-Gibbs statistical mechanics through a single real parameter qq that controls the weight assigned to rare and frequent events. Originally proposed to describe physical systems with long-range correlations, multifractal geometry, and heavy-tailed fluctuations, the framework has become a recurring ingredient in modern artificial intelligence (AI): it underlies sparse attention mechanisms (\textsc{sparsemax} and αα-\textsc{entmax}), maximum-entropy reinforcement learning with controllable exploration, robust and heavy-tailed probabilistic models, and a family of generalized loss functions and regularizers. This paper offers a structured perspective on where Tsallis statistics meets AI. We first review the mathematical core: qq-entropy and its variational (maximum-entropy) foundation, the qq-exponential and qq-logarithm, the qq-central limit theorem, qq-Gaussian distributions, and their dynamical origin in superstatistics, emphasizing the properties that matter for machine learning. We then survey applications across softmax generalization, reinforcement learning, sequential and graph neural models, generative and probabilistic modeling, loss design, and optimization, extracting the recurring design pattern in each case: a tunable interpolation between dense/uniform and sparse/peaked behavior governed by qq. We further argue that the heavy-tailed weight spectra and gradient-noise statistics empirically observed in deep networks are themselves nonextensive signatures, placing modern learning dynamics within the scope of qq-statistics. Finally, we discuss methodological pitfalls, the relationship to information geometry and qq-exponential families, and open directions, arguing that qq should be treated as a learnable inductive bias rather than a fixed hyperparameter.

Explore similar work

CardsList