MADGRAV: a multilevel anomaly-detection pipeline for gravitational-wave searches applied to LIGO data
Authors: Gianluca Inguglia, Huw Haigh, Ulyana Dupletsa, Alessandro Longo
Organizations: Marietta Blau Institute for Particle Physics, Austrian Academy of Sciences, Vienna, Austria · Universit`a degli Studi di Urbino ‘Carlo Bo’, I-61029 Urbino, Italy · INFN, Sezione di Firenze, I-50019 Sesto Fiorentino, Firenze, Italy
We present the results of \textbf{MADGRAV}, a deep-learning-based search for high-mass compact binary coalescences, applied to the data collected by the LIGO interferometers during the third observing run and during the first and second part of the fourth observing run. The \textbf{MADGRAV} pipeline consists of a series of sequential convolutional neural networks that perform anomaly detection, glitch classification, coherence testing, and signal ranking. Data from the Hanford and Livingston LIGO detectors are studied (both individually and in coherence) by way of 1 second Q-transform windows. Of the candidates that survive every stage of the pipeline, 48 reach the significance threshold, and we report 47 gravitational wave detections characterised by a false alarm rate below 1yr−1 with a probability of astrophysical origin pastro>0.9. Of the 47 detections, 44 are shared with the minimally modelled coherent WaveBurst search. The observed total source-frame masses, extracted from official gravitational wave transient catalogues, are in the 14−236M⊙ range with a median of 69M⊙, and a median SNR of 16. We note that the recovered fraction of confident detections rises with mass: for LIGO detectors network SNR >10 the pipeline recovers 8.1% of confident catalog events below 30M⊙, 39.8% between 30 and 100M⊙, and 53.3% above 100M⊙, corresponding to 33.3%, 45.5% and 53.3% of the events detected by coherent WaveBurst in the same bins. These results suggest that anomaly detection pipelines can serve as an independent detection channel complementary to matched filtering in the high-mass high-SNR regime.
Figures & tables
Figure 1 : Detection flow of MADGRAV . The shaded band is the frozen per-detector front end (whitening, 1 s Q-tile, autoencoder score σ , glitch arm g ); orange boxes are the network-level stages, from the σnet trigger through the likelihood-ratio ranking of Eq. ( 4 ) and the gated specialist networks to the false-alarm rate; grey boxes are non-learned stages. Solid arrows follow a zero-lag candidate; dashed arrows indicate the time-slid background.
Figure 2 : Injection recovery over the source-frame total mass—network SNR plane, one panel per run. Colour is the recovery probability at FAR<1yr−1 ; white curves are the 50% (solid) and 90% (dashed) contours. Grey circles are GWTC confident H1–L1 coincident events, blue squares cWB and orange stars MADGRAV detections. The grid used contains 16 logarithmic mass bins and 13 SNR levels ( ρnet=5 – 80 ) with ≥40 injections per cell.
ρnet
5
6
7
8
10
12
15
20
25
30
40
60
80
O3a
0.00
0.00
0.00
0.01
0.07
0.18
0.39
0.55
0.72
0.84
0.92
0.92
0.89
O3b
–
0.00
0.00
0.01
0.05
0.15
0.36
0.53
0.69
0.82
0.93
0.96
0.96
O4a
0.00
0.00
0.00
0.00
0.04
0.12
0.32
0.52
0.73
0.83
0.92
0.96
0.97
O4b
0.00
0.00
0.00
0.01
0.06
0.18
0.41
0.59
0.79
0.88
0.95
0.97
0.97
Table 1 : Recovery fraction at FAR<1yr−1 versus injected H1+L1 network SNR. A dash marks a level absent for that run. Values are run averages over the injected mass mix, not per-cell efficiencies.
Run
Comparable
Found
Predicted
Difference
O3a
37
9
7.8±1.8
+0.69σ
O3b
31
4
4.0±1.1
−0.03σ
O4a
76
15
11.6±2.2
+1.54σ
O4b
86
18
23.2±2.7
−1.88σ
Total
230
46
46.5±4.1
−0.13σ
Table 2 : Detections found versus predicted, evaluated at each event own (Mtot,ρnet) cell. Uncertainties are Poisson-binomial.
Figure 3 : Left (a): Source-frame total mass versus H1+L1 network matched-filter SNR for the confident events of GWTC-2.1, GWTC-3, GWTC-4.0, and GWTC-5.0 that were observed by Hanford and Livingston simultaneously (grey), the 120 identified by cWB at FAR<1yr−1 (blue squares), and the MADGRAV detections at calibrated FAR<1yr−1 (orange stars). GW240406_062847, an event we detected, has no published source-frame total mass [ 5 ] and is not shown in the plot. Right (b): Calibrated FAR with 90% Poisson intervals versus source-frame total mass for the 46 detections (47 less GW240406_062847). Arrows denote 90% upper limits for the 9 events with no louder background family in their full time-slide set (266–623 yr per observing run).
Event
Run
Mtot [ M⊙ ]
ρnet
ρH1+L1
σnet
N
FAR [yr -1 ]
FARincl
pastro
cWB
GW190408_181802
O3a
43
14.6
14.4
5.8
30
0.454
0.454
1.00
✓
GW190412
O3a
37
19.8
19.4
6.1
3
0.041
0.041
1.00
✓
GW190513_205428
O3a
54
12.5
11.9
5.9
6
0.083
0.083
1.00
GW190519_153544
O3a
106
15.9
15.7
7.5
20
0.276
0.276
1.00
✓
GW190521_074359
O3a
76
25.9
25.9
8.1
10
0.151
0.151
1.00
✓
GW190602_175927
O3a
116
13.2
12.6
8.8
19
0.288
0.288
1.00
✓
Table 3 : The 47 MADGRAV detections at calibrated FAR<1yr−1 . Mtot is the published source-frame total mass and ρnet the published network matched-filter SNR [ 12 , 10 , 3 , 5 ] , over the full observing network; ρH1+L1 is the H1+L1-only value of Eq. ( 9 ), the quantity this coincident search could reach and the axis of Fig. 3(a) . The two differ only where Virgo also observed. σnet is the MADGRAV network excess significance. Events with no louder background family ( N=0 ) carry a 90% upper limit. pastro is a per-run Poisson-mixture (FGMC) fit on the foreground-excluded FAR axis. The last column marks candidates also identified by cWB at FAR<1yr−1 .
Figure 4 : Left: volume-averaged efficiency of the search at the detection criterion (calibrated FAR<1yr−1 ) versus detector-frame total mass, per observing run. Right: the corresponding comoving sensitive volume–time ⟨VT⟩ versus source-frame total mass (see text): per-injection horizons at physical network SNR 5 under the run-median release spectra, comoving Vmax=∫dVc/(1+z) (flat Λ CDM, H0=67.9 , Ωm=0.3065 ), and the source-frame rebin implied by each injection redshift. Both panels average over a population uniform in comoving volume with (1+z)−1 rate dilation. Points whose injection support falls below Ninjeff=300 [ 24 ] after the source-frame rebin are not shown (330–400 M⊙ in every run: the bank detector-frame ceiling of 400M⊙ starves that source-frame bin at the horizons, z∼0.3 –0.5); the 260–330 M⊙ points are lower bounds for the same reason.
Figure 5 : Ratio of MADGRAV comoving ⟨VT⟩ to the LVK pipelines’ at matched FAR<1yr−1 , per run, versus source-frame total mass. Both sides are injection-set integrals on the same population and exposure: the numerator from the campaign of Sec. 3 , the denominator from the GWTC-3 and GWTC-5.0 sensitivity-injection releases [ 18 , 19 ] reweighted to it [ 22 ] . Points with fewer than Ninjeff=300 reweighted injections are dropped, which sets both limits: the releases lose support above 230M⊙ , and below 20M⊙ the source-frame rebin migrates volume down-mass on the MADGRAV side alone.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Model
k
latent dim.
detected
GW190521
all
frozen
128
65,536
47/47
yes
83
bottleneck
32
16,384
47/47
yes
84
bottleneck
20
10,240
47/47
yes
76
bottleneck
13
6,656
47/47
yes
83
bottleneck
1
512
24/47
yes
38
Appendix
Table 4 : Catalog events above the matched threshold: the 47 detections, the near-miss GW190521, and all 136 scorable GWTC events ( Mtot>10M⊙ , H–L coincident, ρHL>10 ).
Figure 6 : Trigger-level injection recovery versus network SNR per detector-frame mass stratum at the production operating point. Black: frozen model; colours: bottleneck widths k .
We study machine-learning detection of controlled beyond-General-Relativity (beyond-GR) deviations in gravitational-wave signals, using both synthetic aLIGO-PSD noise and real LIGO H1 detector strain. Three deviation families are applied to General-Relativistic inspiral-merger-ringdown waveforms: amplitude modulation, phase modulation, and frequency modulation, each parameterized by a dimensionless strength coefficient β. A hybrid classifier combining a one-dimensional convolutional neural network with ten hand-crafted waveform statistics is trained on GR and modified waveforms and tested on a deviation type excluded from training. The central result is a quantitative detectability curve as a function of β. Using the real GW150914 strain as a template and real H1 detector noise, we find a detection threshold at β≈0.25, with accuracy rising smoothly from chance at β≤0.2 to perfect classification at β≥0.5. The threshold value is specific to the quadratic-in-time modulation form adopted here and should not be interpreted as a generic constraint on beyond-GR parameters. We nevertheless argue that the negative result at small β is informative: it establishes a quantitative limit on machine-learning-only beyond-GR searches in real detector noise, in the absence of matched-filter signal extraction.
Muhammad Adnan Shahzad
Department of Computer Science and Software Engineering (CSSE), Concordia University, Montreal, QC, Canada
Gravitational Waves (GWs) represent the newest window of astronomy, furthering our understanding of compact objects like black holes and neutron stars in the Universe. The signal from two merging neutron stars is especially interesting since it brings the prospect of concordant electromagnetic and neutrino emissions. Such multi-messenger observations have a transformational impact on fundamental physics, nuclear matter, astrophysics, and gravity. It was first witnessed in 2017 with the detection of the binary neutron star (BNS) merger GW170817. However, searching for BNS signals in real-time in the LIGO-Virgo-KAGRA (LVK) GW detectors presents a computational challenge, as the data streaming out must be matched against ∼ million reference waveforms, which requires up to a thousand CPU cores. We present a different approach using neural networks to learn the presence of a signal in the data. Our algorithm, called Aframe, was deployed in the LVK's fourth observing run and was the first artificial intelligence (AI)-enabled search to detect multiple binary black holes (BBHs) live. In this work, we demonstrate that the approach extends to the lower-mass BNS regime, and is the first AI-enabled search that achieves sensitivity comparable to matched-filter pipelines at lower computational and latency costs. The challenge of the longer-duration BNS signals is addressed by heterodyning the data, following which the network architecture used for BBHs is sufficient to distinguish signal versus background. We also show that this analysis requires a single non-flagship GPU for online deployment. Furthermore, the design and adoption of inference-as-a-service tools allow rapid offline analysis using a distributed pool of GPU resources. Hence, aside from the use case of rapid online data analysis, we also establish the use of Aframe for efficient archival data analysis.
Bhavya Gupta, Deep Chatterjee, William Benoit +7
1LIGO Laboratory, 185 Albany St, MIT, Cambridge, MA 02139, USA · Department of Physics, MIT, Cambridge, MA 02139, USA · School of Physics and Astronomy, University of Minnesota, Minneapolis, MN 55455, USA
In the last decade, kilometre-scale interferometric gravitational-wave detectors have observed hundreds of compact binary mergers, the majority of which are binary black holes. However, the data are noise-dominated, and multiple independent search algorithms (pipelines) are used to enhance sensitivity and improve robustness. Rather than the standard approach of selecting the most significant pipeline output, we combine the outputs from all pipelines using a conformal prediction-based framework to provide statistically rigorous confidence estimates for candidate events. While combining pipelines improves sensitivity and ranking robustness, it requires a principled statistical framework that remains valid as data properties evolve across observing runs. A key challenge is distribution shifts between simulated datasets used for training and calibration and the real, unlabelled, observations used for testing, which can invalidate coverage guarantees and bias confidence estimates. In this work, we address this challenge by incorporating likelihood-ratio reweighting into our conformal prediction framework to account for covariate shift. Using mock datasets containing simulated signals, we demonstrate that weighted conformal prediction restores well-calibrated coverage under covariate shift and increases the confidence of events near the detection threshold, recovering true signals that would otherwise be missed.
Ann-Kristin Malz, Gregory Ashton, Nicolo Colombo
Department of Physics, Royal Holloway, University of London · School of Mathematical Sciences, University of Southampton · Department of Computer Science, Royal Holloway University of London