We develop a theoretical foundation for designing group-equivariant neural networks that align the choice of symmetries with the underlying probability distributions of the data. Utilising the general structure of fibre decompositions on the domain under group equivariant maps and its relation to that of the likelihood ratio, we present a theoretical framework for identifying group actions that maintain optimal classification performance via the Neyman-Pearson lemma. This provides a unified methodology for improving classification accuracy especially in fundamental applications where one has knowledge of the inherent symmetries of the distributions and how they are broken by measurement. As an application to jet classification at the Large Hadron Collider, we find that there can be performance gains when one utilises smaller permutation symmetries within the constituents. This work offers insights and practical guidelines for constructing more effective group equivariant architectures in diverse machine-learning contexts.
Figures & tables
Sym.
Uniform
Normal
AUC
ϵˉ0(ϵ1=0.95)
AUC
ϵˉ0(ϵ1=0.95)
E(3)
0.974 ± 0.001
0.868 ± 0.007
0.920 ± 0.003
0.652 ± 0.010
O(3)
0.981 ± 0.000
0.905 ± 0.002
0.998 ± 0.000
0.994 ± 0.000
O(2)
1.000 ± 0.000
1.000 ± 0.000
0.999 ± 0.000
0.999 ± 0.000
Table 1: The mean and standard deviation of the background rejection at 95% signal acceptance ϵˉ0(ϵ1=0.95)=1−ϵ0(ϵ1=0.95) and the AUC over ten training runs for the two toy scenarios with added noise. The best performing values are highlighted in bold.
Task
Particle type
Value
Quark/Gluon tagging
Negatively charged mesons and leptons
1
Positively charged mesons and leptons
2
Neutral mesons and the photon
3
Rest
0
Top tagging
Bottom
1
Rest
0
Table 2: Scalar node attribute for particle type utilised in the jet classification experiments.
Cont. Sym.
Discrete Sym.
ΔΓ=30 GeV
ΔΓ=50 GeV
AUC
ϵˉ0(ϵ1=0.95)
AUC
ϵˉ0(ϵ1=0.95)
O(1,3)
Sn
0.972 ± 0.001
0.891 ± 0.008
0.979 ± 0.000
0.915 ± 0.002
⊗tSmt
0.977 ± 0.000
0.933 ± 0.002
0.983 ± 0.000
0.937 ± 0.001
O(1,1) l⊗ O(2) t
Sn
0.977 ± 0.001
0.933 ± 0.006
0.984 ± 0.001
0.941 ± 0.003
⊗tSmt
0.978 ± 0.000
0.938 ± 0.002
0.984 ± 0.000
0.942 ± 0.002
Table 3: The mean and standard deviation of the background rejection at 95% signal acceptance ϵˉ0(ϵ1=0.95)=1−ϵ0(ϵ1=0.95) and the AUC over five training runs for top tagging datasets with different mass windows ΔΓ . The best performing values are highlighted in bold.
Cont. Sym.
Discrete Sym.
AUC
ϵˉ0(ϵ1=0.95)
O(1,3)
Sn
0.875 ± 0.000
0.450 ± 0.002
⊗tSmt
0.879 ± 0.001
0.463 ± 0.002
O(1,1) l⊗ O(2) t
Sn
0.900 ± 0.003
0.531 ± 0.010
⊗tSmt
0.901 ± 0.001
0.532 ± 0.005
Table 4: The mean and standard deviation of the background rejection at 95% signal acceptance ϵˉ0(ϵ1=0.95)=1−ϵ0(ϵ1=0.95) and the AUC over five training runs for quark vs gluon classification. The best performing values are highlighted in bold.