cs.LGMay 8, 2026

Approximation Error Upper and Lower Bounds for Hölder Class with Transformers

Authors: Xin HeYuling JiaoXiliang LuJerry Zhijian Yang

Abstract

We explore the expressive power of Transformers by establishing precise approximation error upper and lower bounds for Hölder class. Specifically, a new approximation upper bound is derived for the standard Transformer architecture equipped with Softmax operators, ReLU activation functions, and residual connections. We prove that a Transformer network composed of at most O(εd0/α)\mathcal{O}(\varepsilon^{-{d_{0}}/α}) blocks can approximate any bounded Hölder function with d0d_{0}-dimensional input and smoothness α(0,1]α\in(0,1] under any accuracy ε>0\varepsilon>0. In the case of approximation lower bounds, leveraging the VC-dimension upper bound, we are the first to rigorously prove that Transformers demand for at least Ω(εd0/(4α))Ω(\varepsilon^{-{d_{0}}/({4α})}) blocks to achieve the ε\varepsilon approximation accuracy. As a final step, we extend the derived results for standard Transformers to a general regression task and establish the corresponding excess risk rates demonstrating Transformers' empirical effectiveness in real-world settings.

Explore similar work

CardsList