The Ball and the Box: Two Geometries of Computation in Superposition
Organizations: University of New South Wales · University of Sydney · University of Cambridge
Abstract
Neural representations can encode more features than they have dimensions, a phenomenon known as superposition. We study the dimension needed to compute Boolean gates from such representations. For a single threshold layer with a Gaussian random dictionary and uniformly random sparse Boolean inputs, we derive sharp dimension thresholds under two error criteria. A vanishing expected error count can require more dimensions than correctness of every output with high probability. Shared reads explain the gap: rare realizations can produce many errors at once. The expected-count threshold has ball geometry, while joint reliability has box geometry when a gate is evaluated on every feature tuple. Optimizing shared readout weights and biases gives explicit thresholds for conjunction, disjunction, and majority. For pairwise conjunction, the analysis also describes the transition near the threshold, in agreement with exact simulations.
Figures & tables
| Midpoint bias | Optimized bias | |
|---|---|---|
| Count errors, (ball) | ||
| Forbid errors, (box) | ||
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Meaning |
|---|---|
| dictionary, independent columns | |
| , , , | active set (uniform -subset), input, sparsity, number of inactive features |
| , | stored vector and feature reads; ( ), (inactive reads) |
| , , , | , independent Gaussian samples of sizes and , ( Theorem 3.1 ) |
| deterministic interference scale | |
| , | on and off ; interference vector of a tuple |
| Ratio | [95% interval] | MC mean | MC s.e. | Quadrature | |
|---|---|---|---|---|---|
| 0.60 | 164032 | 0.0000 [0.0000, 0.0009] | 2.462e+06 | 4.199e+05 | 2.485e+06 |
| 1.00 | 273387 | 0.1548 [0.1440, 0.1662] | 312.3 | 127.7 | 427.3 |
| 1.20 | 328064 | 0.8252 [0.8133, 0.8365] | 10.46 | 8.232 | 6.742 |
| 1.35 | 369072 | 0.9727 [0.9672, 0.9772] | 0.06288 | 0.01097 | 0.3274 |
| 1.50 | 410081 | 0.9976 [0.9955, 0.9987] | 0.002693 | 0.0008277 | 0.0178 |
| 1.75 | 478427 | 0.9995 [0.9982, 0.9999] | 0.0004904 | 0.0003452 | 0.0002069 |
| Ratio | 0.8 | 0.9 | 1 | 1.1 | 1.2 | 1.35 | 1.5 |
|---|---|---|---|---|---|---|---|
| , half-slack | 0 | 21 | 668 | 2262 | 3387 | 3980 | 4081 |
| , oracle | 1 | 141 | 1573 | 3305 | 3952 | 4092 | 4096 |
| , half-slack | 1 | 154 | 1258 | 2919 | 3758 | 4065 | 4093 |
| , oracle | 11 | 571 | 2456 | 3788 | 4077 | 4095 | 4096 |
| , half-slack | 0 | 0 | 640 | 3298 | 4009 | 4092 | 4096 |
| , oracle | 0 | 0 | 1308 | 3944 | 4094 | 4096 | 4096 |
| Ratio | succ. | succ. | succ. | |||
|---|---|---|---|---|---|---|
| 0.9 | 33 | 0 | 0 | |||
| 1 | 1208 | 1327 | 1392 | |||
| 1.1 | 3191 | 3896 | 4091 | |||
| 1.2 | 3922 | 4089 | 4096 | |||
| 1.35 | 4084 | 4096 | 4096 | |||
| Campaign, rule | Limit | Simulation [95% interval] | |||
|---|---|---|---|---|---|
| E1, half-slack | 64 | 2.458 | 0.155 | 0.143 [0.133, 0.154] | |
| E1, half-slack | 256 | 2.342 | 0.154 | 0.139 [0.129, 0.150] | |
| E1, half-slack | 1024 | 2.275 | 0.154 | 0.155 [0.144, 0.166] | |
| E1, half-slack | 1024 | 1.538 | 0.147 | 0.149 [0.139, 0.160] | |
| E2, half-slack | 64 | 2.458 | 0.155 | 0.149 [0.138, 0.160] | |
| E2, half-slack | 256 | 2.342 | 0.154 | 0.141 [0.131, 0.152] |
| Criterion | Protects | Midpoint bias | Optimized bias | Source |
|---|---|---|---|---|
| a typical gate | (any fixed bias) | Corollary M.1 | ||
| a typical input | Theorem 6.1 (a,b) | |||
| the total additive loss | Theorem 6.1 (c,d) | |||
| correct on every support | every input, chosen after | — | Proposition 8.1 | |
| Reliability | Expected count | |||
|---|---|---|---|---|
| Gate | (a) half-slack | (b) midpoint | (c) midpoint | (d) optimized |
| Recovery, | ||||
| Target | ||||
|---|---|---|---|---|
| Midpoint bias | ||||
| Half-slack bias | ||||
| Best deterministic common bias |