Small transformers track Bayesian evidence for latent common causes via a context-invariant mechanism
Organizations: Department of Linguistics, University of Tübingen
Abstract
We present an in-depth investigation of how a form of Bayesian reasoning about common causes can emerge as a cross-contextual generalization in small, tractable transformers. Incrementing on recent work, our set-up (i) disentangles causal mechanisms in the model from the causal structure of the true data-generating process, (ii) orients more towards natural language prediction by considering inference of latent common causes, and (iii) considers whether and how Bayesian evidence accumulation for latent common causes can be implemented in representations and mechanisms that allow for cross-context generalization to novel test cases.
Figures & tables
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Layers / heads per layer | 2 / 1 |
|---|---|
| / head dimension | 16 / 16 |
| MLP hidden size, activation | 16, GELU |
| Dropout | none |
| Biases | QKV and output projections |
| Positional encoding | sinusoidal, max length 1000 |
| Vocabulary | 19 (1 shared delimiter 3 contexts 6) |
| Optimizer | AdamW, , weight decay |
|---|---|
| Learning rate | , constant (no warmup or decay) |
| Gradient clipping | norm , no accumulation |
| Batch size | 384 sequences (92,160 tokens) |
| Training budget | 20k steps (7.68M sequences) |
| Validation set | 4,800 fixed sequences |
| Evaluation / checkpoint interval | every 10 steps |
| SparseYour attention | Total | |||
|---|---|---|---|---|
| Layer 1 | 31 | 8 | 13 | 10 |
| Layer 2 | 14 | 7 | 2 | 5 |