Random Feature Gaussian Process Attention: Linear-Time Probabilistic Attention with Calibrated Uncertainty
Organizations: Department of Computer Engineering, Sharif University of Technology · School of Computing and Data Science, The University of Hong Kong
Abstract
Transformers provide a state-of-the-art modeling framework, yet poor calibration limits their reliability in safety-critical applications. A promising direction addresses this issue by interpreting attention as a Gaussian process (GP) posterior, which enables principled uncertainty calibration but incurs cubic complexity in sequence length due to the inversion of the kernel; although decoupled GP variants reduced the cost to quadratic, the computation remains prohibitive in practice. In this paper, we propose the plug-and-play random Fourier feature Gaussian process attention (RFF-GPA) module, which represents the attention as a GP with a stationary kernel approximated by random Fourier features. This low-rank approximation results in linear-time complexity for approximating the posterior mean and variance, making it far more scalable compared to previous work. Empirical results on multiple real-world datasets show that our attention module improves calibration while maintaining predictive accuracy, and simultaneously reduces computational complexity to linear in the sequence length.
Figures & tables
| METHOD | IMAGE | TEXT | ||||
|---|---|---|---|---|---|---|
| F-MNIST | CIFAR-10 | SVHN | 20NG | Hyperp. | SST-2 | |
| METRIC | ACC | |||||
| MLE | 0.889 0.003 | 0.715 0.008 | 0.909 0.007 | 0.654 0.015 | 0.744 0.015 | 0.775 0.009 |
| MLE+Temp | 0.889 0.003 | 0.715 0.008 | 0.909 0.007 | 0.654 0.015 | 0.744 0.015 | 0.775 0.009 |
| MCD | 0.885 0.001 | 0.713 0.002 | 0.902 0.001 | 0.682 0.004 | 0.785 0.010 | 0.781 0.006 |
| SNGP | 0.903 0.002 | 0.727 0.007 | 0.913 0.004 | 0.680 0.009 | 0.774 0.015 | 0.789 0.004 |
| METHOD | (clean) | |||||
|---|---|---|---|---|---|---|
| METRIC | ACC | |||||
| MLE | 0.715 0.008 | 0.681 0.010 | 0.637 0.009 | 0.605 0.009 | 0.564 0.010 | 0.503 0.008 |
| SNGP | 0.727 0.007 | 0.693 0.006 | 0.647 0.001 | 0.613 0.001 | 0.572 0.002 | 0.508 0.003 |
| SGPA | 0.665 0.023 | 0.613 0.023 | 0.556 0.024 | 0.530 0.018 | 0.495 0.014 | 0.441 0.012 |
| RFF-CGP (ours) | 0.737 0.003 | 0.664 0.003 | 0.616 0.004 | 0.588 0.004 | 0.548 0.001 | 0.490 0.002 |
| RFF-GPA (ours) | 0.759 0.001 | 0.702 0.001 | 0.653 0.002 | 0.620 0.002 | 0.576 0.003 | 0.514 0.002 |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Name | Description | Size | Citation/source |
|---|---|---|---|
| Fashion-MNIST | Grayscale clothing-image classification with classes; patch size gives sequence length . | k train / k test | Xiao, Rasul, and Vollgraf (2017) |
| CIFAR-10 | Color natural-image classification with classes; patch size gives sequence length . | k train / k test | Krizhevsky and Hinton (2009) |
| SVHN | Real-world street-view house-number digit classification with classes; sequence length . | 73k train / k test | Netzer et al. (2011) |
| CIFAR-10-C | Distribution-shift benchmark built from CIFAR-10 with common image corruptions at multiple severities. | corruptions severities k images | Hendrycks and Dietterich (2019) |
| 20 Newsgroups | Topic classification over newsgroup posts with classes; text is truncated/padded to tokens. | 18k documents | Lang (1995) |
| Hyperpartisan | Binary news-bias classification using the by-article split; text is truncated/padded to tokens. | Small by-article split | Kiesel et al. (2019) |