SO(3)-RoPE for Spherical Transformers
Organizations: Institute of Computer Science & CIDAS, University of Göttingen · Max Planck Institute for Dynamics and Self-Organization, Göttingen · MIT CSAIL
Abstract
Spherical data arise in many scientific applications. Often spherical transformers disregard the geometry of the underlying spherical domain, causing distortions and coordinate singularities near the poles. We introduce SO(3)-RoPE, a relative positional embedding that incorporates spherical geometry into transformer attention through unitary SO(3) representations. Our formulation is SO(3)-equivariant and compatible with FlashAttention, retaining efficiency of vanilla transformers. On shallow water dynamics prediction over a rotating sphere, our SO3ViT outperforms an S2Transformer baseline with lower errors and reduced runtime.
Figures & tables
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Model | Depth | Width | Params | Memory | Inference Time | ||
| S2Transformer | 4 | 128 | 0.53 M | 0.031 | 0.040 | 8.01 GB | 141.4 ms |
| S2Transformer | 4 | 256 | 2.12 M | 0.031 | 0.040 | 15.96 GB | 198.8 ms |
| S2Transformer | 4 | 512 | 8.43 M | 0.027 | 0.035 | 31.92 GB | 282.1 ms |
| S2Transformer | 8 | 256 | 4.22 M | 0.031 | 0.039 | 21.99 GB | 471.0 ms |
| SO3ViT | 4 | 128 | 0.98 M | 0.045 | 0.058 | 0.74 GB | 6.1 ms |
| SO3ViT | 4 | 256 | 3.54 M | 0.034 | 0.043 | 1.28 GB | 9.3 ms |