Searching for BSM Experimental Signatures with Large Lagrangian Models
Authors: Ibrahim Elsharkawy, Victoria Knapp-Perez, Wahid Bhimji, Aishik Ghosh
Organizations: Department of Physics, University of Toronto and Vector Institute, Toronto, ON, Canada · NERSC, Lawrence Berkeley National Laboratory, Berkeley, California, USA · Department of Physics and Astronomy, University of California, Irvine, CA 92697 · Halluminate, San Francisco, California, USA, 94107 · Georgia Institute of Technology, Atlanta, GA 30332 · Lawrence Berkeley National Laboratory, Berkeley, CA 94720
The search for physics Beyond the Standard Model (BSM) is generally limited not by the supply of theory descriptions but by the lack of discriminating experimental observations. A case in point is dark matter, where the overwhelming gravitational evidence only goes so far in distinguishing between models within a vast theory space. Exploring the space of testable model signatures may help identify overlooked experimental observables and indicate the utility of future experiments. A challenge is designing a search through model signatures outside what is found in the literature. Our primary contribution is hAIthem, a framework that combines the self-guided exploration of reinforcement learning (RL) with the broad literature-derived knowledge of LLMs. We build an RL agent that learns to find which portions of a theory's high-dimensional parameter space are not excluded under some subset of constraints by playing a Battleship-style "game" against a suite of phenomenology tools. The agent is built as a Large Lagrangian Model (LLaM), an autoregressive transformer that reads a tokenized Lagrangian, is pretrained at scale (here on ~1 billion tokens from ~10,000 Lagrangians), and is fine-tuned in a live environment. The framework then constructs a decision tree that separates RL-found regions using observables computed with established tools, and passes the remaining degenerate regions to a set of LLM agents that compete to produce realistic signatures. In this proof of concept, RL-search outperforms an evolutionary-algorithm baseline, finding more viable regions with greater physical diversity. In a restricted space of single dark scalar multiplet models, we find that hAIthem proposes interesting combinations of previously studied observables, such as the application of a halo-independent kinematic ratio to paleo-detectors.
Figures & tables
Figure 1 : Illustration of the hAIthem framework. Symbolic Lagrangians are built and fed to the Large Lagrangian Model along with an encoding of what exclusions we consider. The Large Lagrangian Model searches through parameter space over B tunrs, where each round it probes parameter space N times. The non-excluded regions it finds (non-excluded under our set up, not in nature) are then used to build a decision tree first with observables we compute with community built tools. Remaining degenerate nodes are fed to a set of LLM-agents who compete in producing a realistic signature.
Figure 2 : The similarities between AlphaGo and AlphaZero (left) and hAIthem (right). Both search optimal moves conditioned on the current play turn and past actions. AlphaGo and hAIthem are similar in being pretrained at scale and fine-tuned with RL and, like AlphaZero , hAIthem is not trained on human-generated data. Major differences include the non-discrete nature of the parameter space and the lack of a competitor.
Probe
Observable
Cut
Package
Relic density
Ωh2
τ=1⟶Ωh2<0.118 or Ωh2>0.126 [ 25 , 34 ]
micrOMEGAs
Direct detection
σSI,σSDp
r>1 at 90% CL [ 34 ]
micrOMEGAs
Invisible Higgs
BR(h→inv)
>0.11 [ 5 , 44 ]
micrOMEGAs
Fermi-LAT
⟨σv⟩c⋆
annihilation-channel and mDM dependent bound [ 23 , 34 ]
micrOMEGAs
IceCube
μsolar
conservative bbˉ bound as a function of mDM [ 19 , 17 , 34 ]
micrOMEGAs
LHC pair production
rmaxLHC
rmaxLHC>1
SModelS
Table 1 : Observables that define viability. The relic-density cut is always active with an episode-dependent allowable error given by τ . Each remaining constraint is independently active with probability 1/2 per episode, with the resulting mask provided to the model.
Figure 3 : Viability Computation Pipeline
Figure 4 : The Large Lagrangian Model. The Lagrangian, viability configuration and search history are each encoded into tokens (green) that a shared transformer body processes jointly. H policy heads read the autoregressive tokens, each outputting its own per-dimension Beta conditionals. The value head reads a learned token to estimate the per-turn value. Two auxiliary heads are added and used during pretraining.
small
medium
Hidden width d
256
512
Attention Blocks
4
12
Attention Heads
8
16
Total parameters
∼4.8 M
∼44.0 M
Table 2 : Two model-size configurations. We describe training and model hyperparameters in Appendix B.4 .
Figure 5 : Tokenization of an example Lagrangian composed of a real scalar, Dirac fermion, a complex scalar, and a Majorana fermion. Each dark-sector field is embedded into its own field token from its gauge-eigenbasis quantum numbers, while the symmetry content and per-field summaries (dashed) are embedded into a single global Lagrangian token. LSM is fixed across all models and never encoded.
Figure 6 : A depiction of autoregressive sampling. Each parameter is drawn from a model forward pass and fed back as context for the next forward pass for the next parameter. The model chooses the beta distribution or differential evolution style policy once per probe and fixes the choice for all parameter dimensions.
Figure 7 : The two-stage training workflow. The policy is first pretrained offline (Section II.5 ) on ∼ 10,000 distinct Lagrangians and O(1B) output tokens from earlier RL runs, using the supervised objectives of Equation ( 4 ). The resulting weights warm-start PPO fine-tuning (Section II.4 ), which trains on the live micrOMEGAs and SModelS environment.
Quantity
Count
Distinct Lagrangians
10,843
Boards (Lagrangians + viability criteria, computed with cut augmentation)
≈5.6×107
Probes evaluated
111,046,650
Output tokens (probes × parameter dimension d )
1,308,007,956
Table 3 : Pretraining dataset summary. Board counts include cut augmentation (i.e., we include all combinations of cuts that can be on or off per Lagrangian). The relic-density allowed width τ provides another continuous direction of augmentation. The number of probes evaluated and output tokens are independent of augmentation (i.e., we did in fact evaluate 100M probes). d here is the Lagrangian parameter dimension.
Figure 8 : Two (very) small example decision trees, for (upper) a Lagrangian composed of a Complex Scalar Doublet and (lower) a Lagrangian composed of a Real Scalar Singlet. Here, viable means passes all experimental constraints in Section II .
Figure 9 : Example LLaM-small-RL search episodes at budget B=50 for three benchmark Lagrangians (see numbering of Appendix D.1 ), shown early in the episode (left) and at the final turn (right). Grey points are excluded probes, red dots are current turn’s probes, gold stars are the cumulative viable points found, and the shaded ellipses are the policy’s beta distribution heads. The search acts in the full d -dimensional parameter space. We display the three parameters along which its viable points spread the most. Policy heads need not be in the right location when Differential Evolution is chosen by the LLaM . More examples can be found in Appendix D.1 , and interactive visualizations can be found online.
Search Method
Nv
Lv
R100
Nσ
W
LLaM-small (pretrained)
6,105
9/21
1.5
14
0
LLaM-small (RL)
20,744
14/21
0.9
29
3
LLaM-medium (pretrained)
3,007
13/21
0.8
17
0
LLaM-medium (RL)
51,802
13/21
0.5
23
12
Differential Evolution
8,591
11/21
0.1
24
0
Table 4 : Search benchmark composed of 21 representative Lagrangians. Here, summed over all Lagrangians, all budgets, and all three repeats. Nv is the total number of viable points found, Lv the number of Lagrangians where at least one viable point was found, R100 is the number of regions found via DBSCAN per 100 viable points (a measure of how diverse the viable points are) [ 84 ] , Nσ the distinct signature classes discovered, and W is the win rate, i.e., the number of Lagrangians on which the policy found the most viable points. See Appendix D.1 for information on the Lagrangians used along with per-Lagrangian results. Bold marks the best performance per column.
Figure 10 : A Benchmark of search performance. Total number of viable points found as a function of search budget, summed over the 21 benchmark Lagrangians.
Budget
Search Method
Nv
Lv
R100
Nσ
W
B=5
LLaM-small (pretrained)
4.3±2.9
2.7±1.7
0.0±0.0
1.7±0.5
1.3±1.2
LLaM-small (RL)
7.3±2.1
3.0±0.0
0.0±0.0
1.7±0.5
2.0±0.8
LLaM-medium (pretrained)
3.0±2.8
2.0±1.4
0.0±0.0
2.0±1.4
0.3±0.5
LLaM-medium (RL)
222±183
2.3±0.9
1.8±1.7
3.7±1.9
2.0±1.4
Differential Evolution
3.7±0.5
3.0±0.8
0.0±0.0
2.3±0.9
1.0±0.8
B=10
LLaM-small (pretrained)
7.3±3.7
3.3±1.2
0.0±0.0
3.0±0.8
0.7±0.5
Table 5 : Search benchmark composed of 21 representative Lagrangians, resolved by search budget: B=5 , 10 , 25 , and 50 turns (that is 640, 1280, 3200, and 6400 probes respectively), with mean and error computed over three repeats. Columns are as in Table 4 . Bold marks the best performance per column within each budget block. See Appendix D.1 for per-Lagrangian results
Figure 11 : Diversity benchmark. Distinct experimental-signature classes discovered (that is, the number of decision-tree leaf nodes; solid, left axis) and DBSCAN -found regions per 100 viable points (striped, right axis) for each search method [ 84 ] .
Figure 12 : Mean number of viable points per benchmark Lagrangian as a function of the Lagrangian parameter-space dimension d . The LLaM becomes more effective as the dimension grows, while differential evolution collapses, as expected from the curse of dimensionality. We only plot bins of Lagrangian parameter-space dimension d for which we have sufficient statistics.
Lagrangian
d
Nv
Nσ
NR
Real Scalar Singlet
3
877
3
3
Real Scalar Triplet
3
0
0
0
Complex Scalar Doublet
5
20,068
4
4
Complex Scalar Doublet + dark U(1)+
7
0
0
0
Complex Scalar Doublet + dark U(1)−
7
24
4
4
Complex Scalar Singlet + dark U(1)± (secluded)
8
1,752
18
19
Table 6 : All 11 distinct Lagrangians in our single-dark-multiplet scan, split by spin and ordered by parameter-space dimension d . Nv is the total number of viable points found. Nσ is the number of leaves of each model’s signature tree plotted in Appendix D.2 . NR is the total number of separated parameter-space regions (computed with DBSCAN as in Section III ) [ 84 ] . Models with Y=0 and Dark-U(1) are identical under a field redefinition between Q=±1 Dark-U(1) field charge. 11 11 11 Due to the self-conjugate condition, real scalars cannot carry hypercharge, and thus their doublets (which require Y=±21 for electric neutrality) are excluded. For the same reason, real scalars cannot carry a charge under a Dark-U(1) and thus those models are excluded. 12 12 12 Since our search space allows for different Zn charges for the different components in a complex scalar field, for each complex scalar model, we have two variants. One in which the complex component has a Higgs portal and one in which it is secluded. See Section II .
Figure 13 : Pie chart of exclusions. Each point that is excluded by some set of constraints is tallied, and we compute the ratio of points that were excluded by that combination from the total exclusion count.
Table 7 : Types of measurements proposed by the LLM agent over global and model trees for all Nrepeat=5 repeats. n is the number in that category. The bottom row is the total per split-type for the final aggregated Trees in Figure 14 and in Appendix D.2 . Percentages for categories are percentages over category type. Percentages for Totals are percentages over LLM split type.
Figure 15 : Diagram of Dark Bremsstrahlung possible with the scalar + dark U(1) models we consider. This process results in a dimuon + MET final state near the Higgs mass.
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 16 : Direct-detection Observables σSI and σSDp bounds interpolated from [ 11 , 12 , 15 , 38 , 26 ] . SuperCDMS combines the Si-HV, Si-iZIP, Ge-HV, and Ge-iZIP constraints.
Figure 17 : ⟨σv⟩ bounds per channel interpolated from [ 22 , 65 , 23 ] . The WW channel is non-constraining below ∼mW , similarly for other channels.
Figure 18 : IceCube bbˉ channel bound on μsolar interpolated from [ 17 , 18 ]
τ
Ωh2 range
1
[0.118,0.126]
10
[0.088,0.169]
50
[0.024,0.629]
Appendix
Table 8 : Example Relic-band τ and the acceptance range it maps to. 17 17 17 Given by Ω±(τ)=ΩloΩhi⋅(ΩloΩhi)±τ/2 with Ωlo=0.118 , Ωhi=0.126
Figure 19 : The three production channel classes. (A) Z′ -mediated Drell–Yan for DM charged under U(1)′ . (B) Electroweak Drell–Yan via SM gauge bosons for DM in non-singlet SU(2)L reps. (C) Higgs-portal production (here gluon fusion through a top loop) for SM-singlet DM with no U(1)′ .
#
Lorentz
(SU(2)L,Y)
U(1)′ ?
Channel(s)
1
complex scalar
singlet, Y=0
yes
A
2
Dirac fermion
singlet, Y=0
yes
A
3
real scalar
singlet, Y=0
no
C
4
complex scalar
singlet, Y=0
no
C
5
Majorana / Dirac fermion
singlet, Y=0
no
—
6
complex scalar
doublet, Y=21
no
B
Appendix
Table 9: Dominant production DM channel in the LHC in the search space of Section II .
Table 10 : Representative scalar dark-matter Lagrangians from the agent’s search space. H is the Higgs doublet, s is a real scalar, S is a complex scalar, and Z′ is the U(1)′ gauge boson with charge q . Kinetic and mass terms are not shown except for U(1)′ . αi are coupling parameters the agent learns to scan through and gZ′ , ϵ the dark U(1)′ ’s coupling and kinetic mixing terms.
SU(2)L
Y
Real scalar
Complex scalar
Majorana
Dirac
Singlet
0
✓
✓
✓
✓
Doublet
21
×
✓
×
✓
Triplet
0
✓
✓
✓
✓
Appendix
Table 11 : Spin and quantum numbers for a single dark multiplet we consider. Electrical neutrality requires Y=0 for singlets and Y=±21 for doublets and we restrict to Y=0 for triplets. Real scalars or Majorana fermions cannot carry hypercharge, and similarly cannot carry a dark U(1)’ charge. [ 113 ] .
Type
Examples
Maximum Allowed Range
DM or mediator mass
MDM , MZ′
1 GeV to 10 TeV
Dimensionless Coupling
a2 , λi , gZ′
10−2 to 4π
Kinetic Mixing Parameter
ϵ
10−6 to 10−1
Appendix
Table 12 : Continuous parameters sampled by the agent and their full ranges. Each episode further restricts every type to a randomly drawn sub-range of the full range below. 22 22 22 Our EFT scale range is Λ∈[10−1,103] GeV for the two non-renormizable operators we consider. No points are accepted that do not satisfy mDM≤Λ . This is not relevant for our scan of scalar models who contain only renormalizable operators.
Experiment
Observable
Threshold
Software
XLZD
σSI
bound on σSI as a function of mDM [ 14 ]
micrOMEGAs
LZ full exposure
σSI
bound on σSI as a function of mDM [ 30 ]
micrOMEGAs
DarkSide-20k
σSI
bound on σSI as a function of mDM [ 15 ]
micrOMEGAs
SuperCDMS
σSI
bound on σSI as a function of mDM≤10 GeV [ 26 ]
micrOMEGAs
PICO-500
σSDp
bound on σSDp as a function of mDM [ 39 ]
micrOMEGAs
CTA
⟨σv⟩c⋆
annihilation-channel dependent bound [ 22 ]
micrOMEGAs
Appendix
Table 13 : Projected experiments scored as testability targets. Bounds are interpolated log-log in mDM when needed. HL-LHC and FCC-ee Invisible Higgs thresholds are flat in mDM . The HL-LHC collider testability flag is computed by rescaling the SModelS rmax to projected HL-LHC luminosity. How we interpolate bounds from literature values is given in App. A.1 .
Figure 20 : The probe-history encoder. M learned queries cross-attend, with a learned recency bias, over the embeddings of every probe taken so far, and the M outputs join the input token set as history tokens.
Figure 21 : ( Upper ) Effect of the ν window on a two-parameter toy parameter space. 1σ contours of the broadest (cyan) and sharpest (black) densities each window permits, for a broad and a sharp head at the start, middle and end of an episode. ( Lower ) Per-head concentration windows over an episode with T=50 turns per episode. Each head’s hard ν window slides from broad (exploration) to sharp (exploitation), with a stagger for a division of labor across heads.
small
medium
Token hidden dimension dtok
256
512
Body depth
4
12
Body attention heads
8
16
MLP width
4dtok=1024
4dtok=2048
Total parameters
4.8 M
44.0 M
Appendix
Table 14: LLaM-small and LLaM-medium hyperparameters. Parameters shared between the two sizes are given in Table 15 .
Shared hyperparameters
Architecture
PPO and Training Parameters
Policy heads H
4
Learning rate (AdamW, RL and Pre)
1×10−4
History query tokens M
64
LR warmup episodes for RL
1000
Attention head dim
32
Clip ϵ
0.1
Max Lagrangian parameter dim dmax
128
Probes per turn N
128
Appendix
Table 15 : Hyperparameters shared between the two model sizes.
Observable
Channel Dependence
Bin Width in log10
Overflow Bin Edges in log10
σSI for XLZD, LZ, and DarkSide
N/A
1
±2
σSD for PICO
N/A
1
±2
⟨σv⟩ for CTA
channel ( bbˉ , WW , ττ )
1
±1.5
⟨σv⟩ for Fermi-LAT (15 yr)
channel ( bbˉ , WW , ττ )
1
±1.5
Upward muon flux for IceCube-Gen2
N/A
1
±2
BR( h→ inv): HL-LHC
N/A
0.5
±3
Appendix
Table 16 : Observables that are used to split nodes in the decision tree. Each μe=log10(Opred/Olim) is a continuous ratio of the experiment’s projected limit in log space. We discretize this μe into fixed-width bins with overflow bins. All values below the lower overflow bin are merged into a single “far-below-reach” bin, all values above the upper overflow bin into a single “far-above-reach” bin.
#
Field content
parameter dim d
nf
U(1) ′
1
Z3 : Real Scalar Singlet
3
1
–
2
Z3 : Majorana Singlet
3
1
–
3
Z3 : Majorana Triplet
3
1
–
4
Z3 : Dirac Doublet
4
1
–
5
Z3 : Dirac Triplet
4
1
–
6
Z3 : Complex Scalar Doublet
5
1
–
Appendix
Table 17 : The 21 Lagrangian classes of the benchmark with nf number of fields, ordered by parameter-space dimension d .
LLaM-small (pretrained)
LLaM-small (RL)
LLaM-medium (pretrained)
LLaM-medium (RL)
Differential Evolution
#
Nv
R100
Nσ
Nv
R100
Nσ
Nv
R100
Nσ
Nv
R100
Nσ
Nv
R100
Nσ
1
0
–
0
418
0.2
2
1
0.0
1
0
–
0
0
–
0
2
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
3
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
4
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
5
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
Appendix
Table 18 : Per-Lagrangian benchmark results, summed over all budgets and three repeats, where row numbers refer to Table 17 . Bold marks the policy with the most viable points on each Lagrangian (ties bolded jointly).
LLaM-small (pretrained)
LLaM-small (RL)
LLaM-medium (pretrained)
LLaM-medium (RL)
Differential Evolution
#
Nv
R100
Nσ
Nv
R100
Nσ
Nv
R100
Nσ
Nv
R100
Nσ
Nv
R100
Nσ
1
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
2
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
3
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
4
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
5
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
Appendix
Table 19 : Per-Lagrangian benchmark results at budget B=5 (mean and error over three repeats), where row numbers refer to Table 17 . Bold marks the policy with the most viable points (ties bolded jointly).
LLaM-small (pretrained)
LLaM-small (RL)
LLaM-medium (pretrained)
LLaM-medium (RL)
Differential Evolution
#
Nv
R100
Nσ
Nv
R100
Nσ
Nv
R100
Nσ
Nv
R100
Nσ
Nv
R100
Nσ
1
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
2
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
3
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
4
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
5
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
Appendix
Table 20 : Per-Lagrangian benchmark results at budget B=10 (mean and error over three repeats), where row numbers refer to Table 17 . Bold marks the policy with the most viable points (ties bolded jointly).
LLaM-small (pretrained)
LLaM-small (RL)
LLaM-medium (pretrained)
LLaM-medium (RL)
Differential Evolution
#
Nv
R100
Nσ
Nv
R100
Nσ
Nv
R100
Nσ
Nv
R100
Nσ
Nv
R100
Nσ
1
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
2
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
3
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
4
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
5
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
Appendix
Table 21 : Per-Lagrangian benchmark results at budget B=25 (mean and error over three repeats), where row numbers refer to Table 17 . Bold marks the policy with the most viable points (ties bolded jointly).
LLaM-small (pretrained)
LLaM-small (RL)
LLaM-medium (pretrained)
LLaM-medium (RL)
Differential Evolution
#
Nv
R100
Nσ
Nv
R100
Nσ
Nv
R100
Nσ
Nv
R100
Nσ
Nv
R100
Nσ
1
0
–
0
139±197
0.2±0.0
0.7±0.9
0.3±0.5
0.0±0.0
0.3±0.5
0
–
0
0
–
0
2
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
3
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
4
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
5
0
–
0
0
–
0
0
–
0
0
–
0
0
–
0
Appendix
Table 22 : Per-Lagrangian benchmark results at budget B=50 (mean and error over three repeats), where row numbers refer to Table 17 . Bold marks the policy with the most viable points (ties bolded jointly).
Figure 22 : Example LLaM-small-RL search episodes at budget B=50 for benchmark Lagrangians (see numbering of Table 17 ), shown early in the episode (left) and at the final turn (right). Grey points are excluded probes, red dots are current turn’s probes, gold stars are the cumulative viable points found, and the shaded ellipses are the policy’s beta distribution heads. The search acts in the full d -dimensional parameter space. We display the three parameters along which its viable points spread the most. Policy heads need not be in the right location when Differential Evolution is chosen by the LLaM.
Figure 23 : Example LLaM-small-RL search episodes at budget B=50 for benchmark Lagrangians (see numbering of Table 17 ), shown early in the episode (left) and at the final turn (right). Grey points are excluded probes, red dots are current turn’s probes, gold stars are the cumulative viable points found, and the shaded ellipses are the policy’s beta distribution heads. The search acts in the full d -dimensional parameter space. We display the three parameters along which its viable points spread the most. Policy heads need not be in the right location when Differential Evolution is chosen by the LLaM.
Figure 24 : Three Decision Trees from the One-Multiplet Search Space. None meet the condition for the LLM agent to be applied.
Figure 25 : Complex Scalar Singlet Decision Tree.
Figure 26 : Decision Tree for a Complex Scalar Singlet with Dark U(1)′ and a secluded complex component.
Figure 27 : Decision Tree for a Complex Scalar Singlet with Dark U(1)′ .
Machine Learning in Science, University of Tübingen, Tübingen, Germany · Tübingen AI Center, Tübingen, Germany · Boehringer Ingelheim, Biberach, Germany +1