Scaling Collider Event Generation with Residual-Quantized Tokens
Organizations: Weizmann Institute of Science, Rehovot, Israel
Abstract
Full detector simulation and reconstruction of collider events are projected to become major bottlenecks at the High-Luminosity Large Hadron Collider, motivating the development of fast, ML-based surrogates. At the same time, LLMs have driven fast progress in generative discrete modeling: autoregressive transformers trained on tokenized data now represent the state of the art across a range of generative tasks. We extend the discrete modeling paradigm by introducing a particle-level generative model trained on residual-quantized full-event data. We demonstrate the ability of this model family to perform conditional generation from detector-stable particles; we study its scaling behavior across a range of dataset and model sizes, characterize the effects of repeated data exposure and demonstrate that token-level loss systematically predicts downstream physical fidelity. These results provide an empirical framework for scalable collider full-event generation based on residual-quantized representations.
Figures & tables
| Dataset | – [GeV] | Type | Training [evts] | Testing [evts] |
|---|---|---|---|---|
| – | Out-of-distribution | |||
| – | In distribution | |||
| QCD | 470–600 | In distribution | ||
| QCD | 600–800 | In distribution | ||
| QCD | 1000–1400 | Out-of-distribution |
| Dataset | Usage | Perplexity | Usage | Perplexity | Usage | Perplexity | Usage | Perplexity |
|---|---|---|---|---|---|---|---|---|
| QCD + | 58.5% | 4042 | 99.9% | 5575 | 99.9% | 4894 | 99.9% | 4723 |
| 58.2% | 3338 | 99.3% | 5382 | 99.4% | 4797 | 99.6% | 4662 | |
| QCD 1000 | 58.5% | 3897 | 99.9% | 5652 | 100.0% | 4981 | 100.0% | 4871 |
| Name | Layers | Heads | Total Params | |||
|---|---|---|---|---|---|---|
| 10M | 256 | 8 | 8 | 32 | 512 | M |
| 22M | 384 | 10 | 8 | 48 | 768 | M |
| 41M | 512 | 12 | 8 | 64 | 1,024 | M |
| 97M | 768 | 14 | 12 | 64 | 1,536 | M |
| 187M | 1,024 | 16 | 16 | 64 | 2,048 | M |
| 1B | 2,048 | 24 | 32 | 64 | 4,096 | M |
| Data fraction | Unique events | Epochs | PF tokens | Total tokens |
|---|---|---|---|---|
| 3.3% | M | M | ||
| 33.3% | B | B | ||
| 100% | B | B |
| Law | Functional form | MAPE [%] | |||||
|---|---|---|---|---|---|---|---|
| Chinchilla | 3.55 | 0.9695 | |||||
| Skaling | 1.17 | 0.9961 | |||||
| Prescriptive | 2.47 | 0.9690 | |||||
| Custom | 0.79 | 0.9967 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| mean | 1.13 | 0.00 | 0.206 | 0.331 | 0.34 |
|---|---|---|---|---|---|
| std. | 0.83 | 1.29 | 0.691 | 0.702 | 5.73 |
| Model | Transformer | Code proj. | Token head | Card. head | Embedding tables | |
|---|---|---|---|---|---|---|
| 10M | 5.244 | 0.004 | 2.106 | 0.169 | 2.305 | 9.827 |
| 22M | 14.747 | 0.007 | 3.155 | 0.302 | 3.457 | 21.668 |
| 41M | 31.460 | 0.009 | 4.204 | 0.468 | 4.609 | 40.750 |
| 97M | 82.578 | 0.013 | 6.302 | 0.899 | 6.914 | 96.706 |
| 187M | 167.775 | 0.017 | 8.400 | 1.461 | 9.218 | 186.871 |
| 1B | 1006.638 | 0.035 | 16.792 | 5.018 | 18.436 | 1046.918 |
| A100 | H200 | |||||
|---|---|---|---|---|---|---|
| Model | evt/s | ms/evt | evt/s | ms/evt | ||
| 10M | 1,024 | 126.5 | 7.9 | 2,048 | 265.1 | 3.8 |
| 22M † | 512 | 15.1 | 66.1 | 1,024 | 26.8 | 37.4 |
| 41M | 256 | 62.5 | 16.0 | 512 | 132.5 | 7.5 |
| 97M | 256 | 38.2 | 26.2 | 256 | 72.7 | 13.8 |
| 187M | 128 | 22.0 | 45.4 | 256 | 48.8 | 20.5 |
| Chinchilla | Skaling | Prescriptive | Custom | |
| – | – |
| Residual | Unit | 10M | 22M | 41M | 97M | 187M | 1B | RQ-VAE | Noise |
|---|---|---|---|---|---|---|---|---|---|
| Event residuals | |||||||||
| – | 0.0212 | 0.0154 | 0.0102 | 0.00538 | 0.00526 | 0.00262 | 0.00144 | 0.00132 | |
| GeV | 17.3 | 9.98 | 6.69 | 3.81 | 2.97 | 1.81 | 0.363 | 0.271 | |
| GeV | 17.1 | 10.3 | 6.82 | 3.77 | 3.09 | 1.79 | 0.251 | 0.347 | |
| GeV | 22.9 | 13.2 | 8.60 | 4.62 | 3.78 | 2.09 | 0.439 | 0.491 | |
| – | 0.0195 | 0.0131 | 0.00940 | 0.0130 | 0.00629 | 0.00513 | 0.00793 | 0.00509 | |