dIon: Fragmentation-Based Invariance for Self-Supervised Learning of Tandem Mass Spectra
Authors: Alfred Nilsson, Joel Lapin, Samuel H. Payne, Mathias Wilhelm, Lukas Käll
Organizations: Science for Life Laboratory, KTH Royal Institute of Technology Stockholm, Sweden · Computational Mass Spectrometry, TUM School of Life Sciences Technical University of Munich, Freising, Germany · Biology Department, Brigham Young University Provo, Utah 84602, United States · Computational Mass Spectrometry, TUM School of Life Sciences Technical University of Munich, Freising, Germany Munich Data Science Institute, Technical University of Munich Garching, Germany
We introduce a novel invariance for peptide tandem mass spectrometry data, unlocking self-supervised representation learning that improves de novo sequencing of peptides. This invariance exploits the physical relationship between precursor properties (mass and charge) and fragment-ion evidence, without requiring peptide sequence labels. We introduce dIon, which adapts the DINO framework with two latent prediction tasks, both recovering a clean teacher representation: one from a spectrum mixture, using the precursor as a selection query, and one from a partial spectrum with the precursor withheld. The first associates precursor information with fragment-ion evidence; the second prevents representational collapse onto that information alone. Mechanistic probes support both effects, and ablations show that the full objective performs best. Under identical end-to-end training, dIon initialization improves de novo peptide precision over training from scratch by 5.5 and 8.4 percentage points on the held-out MassIVE-KB and Kingdoms test sets, and by 2.3 and 4.8 percentage points with a larger supervised training corpus. The resulting models surpass fully supervised state-of-the-art de novo sequencing models on the diverse, multi-species Kingdoms corpus under the same greedy-decoding protocol. Without peptide labels, dIon learns strong native peptide-similarity geometry compared with other learned models; with limited peptide-supervised adaptation, it achieves the best retrieval and pair-discrimination performance across all representation benchmarks.
Figures & tables
Figure 1: Dual latent-prediction objective. (a) The student sees the mixture Ai∪Bi with the precursor (M,c) of A and matches the teacher output qi for the clean view Ai ( Lmix ). (b) The student sees a partial view Pj of A with the precursor withheld and matches the same target ( Lpart ). The teacher is an EMA of the student, and SK denotes Sinkhorn–Knopp balancing.
Peptide precision
Training / initialization
Kingdoms
MSKB-final
MSKB-final training
Scratch, end-to-end
0.449
0.747
dIon init., end-to-end
0.533
0.802
dIon-de-novo-labeled-v1 training
Scratch, end-to-end
0.554
0.819
Table 1: De novo sequencing on the Kingdoms de novo (all 70 species) and MSKB-final test sets. Within each training regime, scratch and dIon initialization use an identical end-to-end training protocol.
Bacterial
NineSpecies
Kingdoms
Retrieval
Pairs
Retrieval
Pairs
Retrieval
Pairs
Method
mAP
AUROC
mAP
AUROC
mAP
AUROC
Binned spectral angle
0.915
0.943
0.975
0.954
0.905
0.871
Peptide-supervised
Casanovo 4.0
0.881
0.898
0.932
0.859
0.838
0.746
GLEAMS
0.916
0.948
0.976
0.962
0.903
0.876
Table 2: Representation geometry across test benchmarks (Kingdoms: all 70 species).
Correct-source accuracy
Probe
First
Second
Both
Mixture A∪B , precursor of A / B
0.997
0.929
0.926
Partial views P1 / P2 of A , null precursor
0.987
0.984
0.978
Table 3: Source recovery by dIon on the NineSpecies validation set, by cosine similarity of the global representation.
Figure 2: Validation probes during pretraining of the 100-epoch ablation checkpoints. Strict is defined in Section 4.3 .
Appendix figures & tables34 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Use
Train
Validation
Test
Split
PXD010000
SSL pretraining
9,234,199
467,723
—
Acquisition project / run
PXD010613
Retrieval
—
51,770
518,804
Run held out; exact sequence filtered; development from PXD010000 validation
NineSpecies
Retrieval
—
6,119
182,393
Reserved test; test peptides excluded from development
Kingdoms
Retrieval
—
658,738
4,112,590
Acquisition batch; run-disjoint where available; 70 species
MSKB-final
De novo
1,989,996
10,004
200,000
Casanovo split; canonical peptidoform
DNL-v1
De novo training
5,275,467
170,140
345,414
Global peptidoform separation
Appendix
Table 4: Datasets used in the principal experiments, with spectrum counts per split and the split criterion.
Corpus
Full-charge
Charge-2–4
Split definition
PXD010613
518,804
499,617
PXD010613 runs are held out from PXD010000 development data; exact modified sequences overlapping PXD010000 train/validation are removed.
NineSpecies
182,393
181,852
Reserved fixed paper test; development excludes normalized peptide backbones appearing in this test.
Kingdoms
4,112,590
4,087,062
For 10 multi-batch species, validation holds out one deterministic acquisition batch; test uses remaining files plus all other species. Run-disjoint where batch structure exists, not universally project-disjoint.
Appendix
Table 5: Held-out retrieval spectra. The charge-2–4 column is the matched subset used for external-reference comparisons.
Source
Train
Validation
Test
Charges
MSKB-final
1,974,956
25,002
199,992
2–5
PXD010000 bacterial corpus
863,005
48,273
48,073
1–10
ProteomeTools Part III (PXD021013)
788,303
37,086
37,426
1–7
ProteomeTools Part I (PXD004732)
508,466
12,662
12,939
2–7
MassIVE-KB v2.0.15
416,182
17,891
17,472
1–8
ProteomeTools Part II (PXD010595)
382,417
16,881
16,576
2–6
Appendix
Table 6: Sources of DNL-v1, with spectra per split and precursor-charge range.
Component
Setting
Transformer blocks
15
Model dimension d
1024
Attention heads
8 (128 dimensions per head)
Feed-forward network
Linear(1024, 2048) → ReLU → Linear(2048, 1024)
Normalization
Post-residual LayerNorm (attention and FFN branches)
Independently sampled per view from the same batch
Appendix
Table 8: View construction and objective weighting for the selected dIon configuration. Peak fractions are uniform random subsets of the anchor spectrum.
Figure 3: End-AA probes during pretraining of the 100-epoch ablation checkpoints, complementing Figure 2 . The dual objective leads both probes in the second half of pretraining.
Peptide precision
Pretraining objective
Frozen encoder
End-to-end
Dual objective
0.675
0.782
Pure mixture
0.638
0.778
Mixture-free control
0.610
0.764
No pretraining
0.582
0.744
Appendix
Table 10: De novo transfer from the 100-epoch ablation checkpoints on the MSKB-final development- validation split. The last row is the control without pretraining: a randomly initialized frozen encoder, and a model trained from scratch.
Figure 4: Unsupervised embedding quality evaluation for the objective ablation: normalized RankMe ( Garrido et al., 2023 ) and self-cluster ( Tsitsulin et al., 2023 ) . These are coarse checks for pathological geometry and are not monotonic measures of representation quality, so they are not used to rank the objectives. In particular, the pure mixture objective shows no dimensional collapse: its effective rank is not low. Its failure is representational collapse onto precursor information, which these diagnostics cannot detect.
Correct-source accuracy
Separation
Readout
Precursor of A
Precursor of B
Both
Margin
Backbone cosine (selected)
0.997
0.929
0.926
0.767
Backbone Euclidean
0.999
0.847
0.846
18.860
Projection head, Jensen–Shannon
0.998
0.869
0.867
0.004
Appendix
Table 11: Precursor-conditioned source selection by dIon on the NineSpecies validation set. The mixture A∪B of an anchor A and a competing spectrum B is held fixed and only the precursor changes. A trial is correct when the representation lies closer to the clean representation of the conditioned source than to that of the other source.
Correct-source accuracy
Separation
Readout
View 1
View 2
Both
Mean
Margin
Backbone cosine (selected)
0.987
0.984
0.978
0.986
0.498
Appendix
Table 12: Source recovery by dIon from independent precursor-withheld partial views, NineSpecies validation set. Two independently sampled 60% peak subsets of each anchor are embedded under the null precursor state. A view is correct when it lies closer to the clean representation of its own anchor A than to that of a different-peptide spectrum B .
Figure 5: Conditioning-mass sweep on one measured spectrum containing the co-isolated peptides TACTYNNIPLER (A, targeted) and TFYGYNDMADAK (B). The peaks are fixed and only the conditioning mass varies. A positive margin means the representation is closer to A’s reference than to B’s. Dashed lines mark the true masses of A and B, and the dotted line the first 13 C isotope of A.
Bacterial
NineSpecies
Kingdoms
Retrieval
Pairs
Retrieval
Pairs
Retrieval
Pairs
Method
mAP
AUROC
mAP
AUROC
mAP
AUROC
Binned spectral angle
0.913
0.941
0.972
0.954
0.908
0.871
Random Transformer
0.714
0.791
0.823
0.704
0.710
0.637
Peptide-supervised
Casanovo 4.0
0.880
0.896
0.933
0.858
0.838
0.746
Appendix
Table 13: Native representation geometry across test benchmarks over all precursor charges. Unlike the charge-2–4 subsets of Table 2 , these populations retain slightly more spectra and exclude GLEAMS, which embeds only charges 2–4. All metrics compare only spectra with the same precursor charge and precursor mass within 10 ppm.
Retrieval
Static pairs
Method
Broad mAP
AP
AP
AUROC
Binned spectral angle
0.527
0.913
0.934
0.941
Random Transformer
0.126
0.714
0.804
0.791
Casanovo 4.0
0.416
0.880
0.900
0.896
InstaNovo-FM
0.377
0.847
0.876
0.862
dIon
0.492
0.893
0.926
0.927
Appendix
Table 14: Complete native-geometry diagnostics on the bacterial test set.
Retrieval
Static pairs
Method
Broad mAP
AP
AP
AUROC
Binned spectral angle
0.784
0.972
0.939
0.954
Random Transformer
0.157
0.823
0.726
0.704
Casanovo 4.0
0.524
0.933
0.862
0.858
InstaNovo-FM
0.417
0.877
0.784
0.762
dIon
0.611
0.922
0.868
0.864
Appendix
Table 15: Complete native-geometry diagnostics on the NineSpecies V2 test set.
Retrieval
Static pairs
Method
Broad mAP
AP
AP
AUROC
Binned spectral angle
0.499
0.908
0.879
0.871
Random Transformer
0.083
0.710
0.661
0.637
Casanovo 4.0
0.324
0.838
0.780
0.746
InstaNovo-FM
0.323
0.806
0.753
0.708
dIon
0.442
0.824
0.767
0.748
Appendix
Table 16: Complete native-geometry diagnostics on the Kingdoms test set.
Retrieval
Static pairs
Method
Broad mAP
AP
AP
AUROC
Binned spectral angle
0.537
0.915
0.938
0.943
Casanovo 4.0
0.424
0.881
0.903
0.898
InstaNovo-FM
0.384
0.847
0.880
0.864
GLEAMS
0.475
0.916
0.946
0.948
dIon
0.506
0.895
0.929
0.928
Appendix
Table 17: Complete charge-2–4 matched diagnostics on the bacterial test set.
Retrieval
Static pairs
Method
Broad mAP
AP
AP
AUROC
Binned spectral angle
0.786
0.975
0.939
0.954
Casanovo 4.0
0.526
0.932
0.863
0.859
InstaNovo-FM
0.419
0.874
0.785
0.763
GLEAMS
0.746
0.976
0.951
0.962
dIon
0.612
0.926
0.869
0.865
Appendix
Table 18: Complete charge-2–4 matched diagnostics on the NineSpecies V2 test set.
Retrieval
Static pairs
Method
Broad mAP
AP
AP
AUROC
Binned spectral angle
0.503
0.905
0.879
0.871
Casanovo 4.0
0.327
0.838
0.780
0.746
InstaNovo-FM
0.326
0.805
0.753
0.708
GLEAMS
0.440
0.903
0.895
0.876
dIon
0.445
0.821
0.767
0.748
Appendix
Table 19: Complete charge-2–4 matched diagnostics on the Kingdoms test set.
Task AUROC
Representation / initialization
SQA
Chimericity
Oxidized Met
Frozen transfer
Binned + precursor metadata
0.737
0.708
0.654
Random Transformer
0.753
0.760
0.642
Casanovo 4.0
0.807
0.807
0.795
InstaNovo-FM
0.810
0.827
0.835
Appendix
Table 20: Auxiliary binary spectrum-level tasks on held-out test sets (AUROC). Frozen representations and end-to-end adaptation are reported as separate regimes.
Component
Setting
Decoder blocks
9
Model dimension
1024
Attention heads
8
Feed-forward width
1024
Dropout
0.25
Encoder memory dimension
1024
Appendix
Table 21: De novo peptide decoder configuration.
Setting
Value
Objective
Teacher-forced token cross-entropy
Label smoothing
None
Optimizer
Adam
Learning rate
1×10−4
Weight decay
None
Gradient clipping
2.0
Appendix
Table 22: De novo supervised training settings, held fixed across initialization conditions.
Matched
Decoy
Model
Coverage
Emitted
All spectra
Tryptic
Reversed
Shuffled
Modified calls
This work
dIon init., DNL-v1 training
1.000
0.332
0.332
0.319
0.0025
0.0021
0.239
External reference
Casanovo 5.2.1
0.900
0.324
0.292
0.311
0.0020
0.0024
0.326
InstaNovo 1.2.2
1.000
0.320
0.320
0.305
0.0020
0.0017
0.291
Appendix
Table 23: Complete unlabeled yeast results. Coverage is the fraction of submitted spectra receiving a prediction, tryptic restricts matches to tryptic peptides, the shuffled decoy preserves each protein’s amino-acid composition, and modified calls are predictions carrying at least one modification.
Figure 6: Peptide precision against coverage on the Kingdoms de novo test set. Each model ranks its predictions by its own confidence score, and coverage is the fraction of all test spectra accepted, so the right endpoint is the full-coverage precision of Table 1 . dIon and scratch are fine-tuned end to end on DNL-v1 with 200 peaks.
Amino acids
Training / initialization
Precision
Recall
Precision @ 100% coverage
MSKB-final training
Scratch, end-to-end
0.610
0.612
0.449
Scratch, end-to-end (1,000 peaks)
0.615
0.617
0.454
dIon init., end-to-end
0.702
0.702
0.533
dIon init., end-to-end (1,000 peaks)
0.709
0.708
0.538
Appendix
Table 24: Complete Kingdoms de novo results, including amino-acid metrics and the models fine-tuned with a 1,000-peak cap for both initializations. Within each training regime, scratch and dIon initialization use an identical end-to-end training protocol.
DNL-v1, 200 peaks
External reference
Species
Spectra
dIon init.
Scratch
Casanovo 5.2.1
InstaNovo 1.2.2
Akkermansia muciniphila
27,278
0.584
0.535
0.487
0.546
Arabidopsis thaliana (callus)
100,000
0.617
0.573
0.528
0.586
Arabidopsis thaliana (root)
100,000
0.620
0.574
0.510
0.581
Arabidopsis thaliana (sprout)
100,000
0.652
0.605
0.537
0.620
Bacillus subtilis
17,126
0.559
0.527
0.506
0.546
Appendix
Table 25: Per-species peptide precision at full coverage on the Kingdoms de novo test set (deterministic run-disjoint split, at most 100k spectra per species).
Amino acids
Training / initialization
Precision
Recall
Precision @ 100% coverage
MSKB-final training
Scratch, end-to-end
0.895
0.895
0.747
dIon init., end-to-end
0.927
0.926
0.802
dIon init., end-to-end (1,000 peaks)
0.935
0.934
0.822
DNL-v1 training
Appendix
Table 26: Complete MSKB-final de novo results, including amino-acid metrics and the dIon-initialized models fine-tuned with a 1,000-peak cap. Within each training regime, scratch and dIon initialization use an identical end-to-end training protocol.
Amino acids
Model
Precision
Recall
Peptide precision
Precision–coverage AUC
dIon, frozen encoder
0.225
0.227
0.106
0.328
Appendix
Table 27: Oracle-precursor decoding of real wide-window DIA MS2 scans: 17,781 queries, each pairing one of 9,973 scans with the precursor of a DIA-NN-identified peptide in it. Metrics use the matcher of Appendix E.7 .
Setting
Value
Decoding order
Reverse peptide order
Search
Single beam (greedy)
Maximum peptide length
100
Minimum peptide length
6
Precursor mass tolerance
50 ppm
Isotope-error range
0–1
Appendix
Table 28: De novo decoding protocol, identical across all reported models.
Figure 7: Partial views retaining 30% or 60% of peaks, validation probes. The canonical run’s mini-de-novo points are an offline reconstruction from its saved checkpoints.
Figure 8: Gram-anchored refinement continued from epoch 279 of the canonical run, validation probes. The canonical run’s mini-de-novo points are an offline reconstruction from its saved checkpoints.
Peptide precision
Training corpus / encoder
Frozen encoder
End-to-end
DNL-v1 training
dIon
0.555
0.842
dIon, Gram-refined
0.712
0.842
MSKB-final training
dIon
0.487
0.802
Appendix
Table 29: De novo peptide precision on the MSKB-final test set with and without Gram-anchored refinement of the pretrained encoder.
Figure 9: Collapse diagnostics of the three development runs and, in panel (b), the 100-epoch dual-objective ablation run, as per-epoch means. For the dual-objective run the entropy is that of the sharpened teacher distribution before Sinkhorn–Knopp balancing. The transient-collapse run was stopped early.
Figure 10: iBOT-style masked-token screen against the canonical run, validation probes. The canonical run’s mini-de-novo points are an offline reconstruction from its saved checkpoints.
De novo peptide sequencing from tandem mass spectrometry is pivotal in proteomics, enabling identification of novel peptides without reference databases. While recent Transformer-based encoder-decoder models have achieved remarkable performance, we uncover a critical pathology in their inference dynamics. Through comprehensive feature scaling experiments, we demonstrate that existing auto-regressive peptide decoders tend to over-rely on generated-sequence priors while progressively under-utilizing fine-grained physical evidence from the input mass spectrum. This phenomenon leads to suboptimal results, where generated peptide sequences are biologically plausible yet not faithful to the input spectrum. To rectify this, we propose MemNovo, a training-free and plug-and-play mechanism that re-balances peptide and spectral contributions at inference time. MemNovo alleviates the information bottleneck by establishing a persistent spectral memory bank and injecting retrieved features directly into the final decoding stage via an ultra-conservative residual connection. Theoretical analysis confirms that this mechanism restores the mutual information between the decoder state and the raw spectrum. Extensive experiments on the Nine Species benchmark with two representative baselines, Casanovo and InstaNovo, demonstrate that MemNovo consistently improves both amino acid precision and peptide precision, achieving up to 39.1% relative improvement in peptide precision for Casanovo and up to 3.9% for InstaNovo, with negligible computational overhead.
Dongxin Lyu, Jingbo Zhou, Hongxin Xiang +2
Westlake University Hangzhou, Zhejiang, China · Hunan University Changsha, Hunan, China · Shanghai Artificial Intelligence Laboratory Shanghai, China +1
Tandem mass spectrometry provides a high-throughput framework for identifying and quantifying proteins in complex biological samples. In computational proteomics, predicting peptide MS/MS spectra is a critical task, enabling downstream applications such as large-scale peptide identification and quantification. While deep learning architectures have substantially improved prediction accuracy, three evaluation challenges obscure the true progress of the field. First, inconsistent data preprocessing and incompatible model output spaces hinder fair model comparison. Second, flawed data splitting strategies can permit hidden sequence leakage and inflate reported performance. Third, existing evaluations typically lack comprehensive cross-species benchmarking and systematic assessment of model robustness to influential experimental conditions. To address these challenges, we propose PepSpecBench, a unified benchmark for peptide MS/MS spectrum prediction. PepSpecBench standardizes data preprocessing across complementary public datasets, enforces a strict backbone-disjoint splitting strategy to eliminate sequence leakage, and evaluates diverse architectures within a shared fragment-ion representation space. It further introduces a comprehensive multi-species evaluation suite and physically grounded metadata perturbation probes to assess model robustness and instrument awareness. We uncover previously unrecognized performance discrepancies and robustness limitations across six representative models, providing actionable insights for future model design, evaluation and practical deployment.
Zhiwen Yang, Pan Liu, Yifan Li +2
The Hong Kong University of Science and Technology (Guangzhou). · Yangzhou University. · The Hong Kong University of Science and Technology.
Molecular structure elucidation from tandem mass spectra (MS/MS) is a central inverse problem in analytical chemistry. Most existing approaches to MS/MS identification remain tied to reference libraries or predefined candidate sets, whereas de novo methods aim to generate structures directly from spectra. A common de novo route predicts a molecular fingerprint from the spectrum and then decodes structures from it, enabling decoder pretraining on large molecule-only corpora. However, this paradigm creates a training-inference mismatch: the decoder is trained on oracle fingerprints computed from molecules, but at inference it is queried with a noisy spectrum-induced fingerprint posterior that is typically collapsed to a single thresholded fingerprint. We introduce MS-GPT, which recasts fingerprint-mediated de novo elucidation as spectrum-induced posterior querying of a conditional molecule-language model. MS-GPT conditions a molecule-language model on fingerprints and formulas, then converts the spectrum-induced posterior into a band of fingerprint queries near the oracle-fingerprint manifold through active-bit density calibration. Candidates sampled across this band are pooled and ranked by generation-frequency consensus. A lightweight LoRA adapter further mitigates domain-specific posterior bias while preserving the pretrained molecular prior. On NPLIB1 and MassSpecGym, MS-GPT sets a new state of the art, reaching Top-1/Top-10 exact-match accuracy of 29.8%/41.1% and 23.9%/28.7%, respectively. Candidate-pool scaling shows that efficient autoregressive molecular generation continues to improve recall with a little additional inference cost. The source code and model checkpoints are available at https://github.com/VIKI623/MS-GPT.