Skip to main content
Mass Spectrometry· 15-page report· 1 figure

MS Fragmentation Rule Mining

Discover fragmentation patterns from 1,267 GNPS spectra with FDR-corrected statistical enrichment analysis.

What this research found

Most peaks in a tandem mass spectrum go unexplained, because the catalogue of known fragmentation rules is small and largely hand-assembled. Mining 1,267 natural-product spectra for statistical associations between molecular substructures and neutral losses yielded 1,541 rules that survive false-discovery correction, spanning 96 substructures and 136 distinct losses with enrichment factors up to 52.82-fold. The rules cluster into nine fragmentation families and were packaged as a searchable compendium plus a prediction tool, although precision on held-out spectra remains 2.5%.

  • 1,541 substructure-to-neutral-loss rules survived Benjamini-Hochberg correction out of 21,009 tests (141 structural features against 149 loss bins), all at Q < 0.05 with odds ratios above 3.0, covering 96 unique substructures and 136 unique neutral losses.
  • Rule strength varies enormously. Confidence — the probability of a loss given the substructure — ranges from 1.7% to 92.9%, and enrichment from 1.11-fold to 52.82-fold. The single strongest rule links MACCS structural key 23 to a 100.06 Da loss at 92.9% confidence and 52.8-fold enrichment.
  • The rules organise into nine fragmentation families. A bipartite network of 232 nodes and 1,541 edges yields a Louvain modularity of 0.486, forms one connected component, and averages 13.3 connections per node.
  • Simple combinatorial bond breaking accounts for very little of what spectra actually contain: 51,241 theoretical fragments matched 6,381 experimental peaks, explaining an average of just 4.8% of spectral intensity, since one or two bond cuts cannot represent rearrangements, multi-step losses or charge migration.
  • Applying the rules on an 80/20 split of the same library raised F1-score by 127%, but precision stayed at 2.5%, meaning the rule set proposes far more losses than a spectrum shows.
  • The predictor is sparse in practice. At a 30% confidence threshold it returned a single rule for caffeine — an 84.06 Da loss at 50% confidence and 16.1-fold enrichment — and no predictions at all for quercetin or cholesterol.

How it was done

1,267 tandem mass spectra from the GNPS NIH Natural Products Library passed quality control requiring a valid structure annotation and at least five fragment peaks. Each molecule was fragmented combinatorially in RDKit by cutting one or two non-aromatic, non-small-ring bonds, and theoretical fragment masses were matched to experimental peaks within 0.01 Da or 20 ppm in protonated-molecule mode, which succeeded for 1,265 molecules and produced 8,186 neutral loss events. Each of the 166 MACCS structural keys was then tested against each 0.02 Da loss bin using Fisher's exact test with Benjamini-Hochberg correction, retaining only associations at Q < 0.05 with an odds ratio above 3.0 and discarding trivial losses below 19 Da such as water and ammonia. Surviving rules were assembled into a bipartite substructure-loss network partitioned by the Louvain community detection algorithm, then compiled into a rule compendium and wrapped in a tool that takes a molecular structure string and returns candidate losses ranked by confidence and enrichment.

Data sources

  • GNPS NIH Natural Products Library — 1,267 quality-filtered MS/MS spectra with structure annotations
  • MACCS structural keys — 166-bit molecular fingerprints computed with RDKit

Limitations

Validation used an 80/20 split of the same spectral library rather than an independent database such as MassBank or NIST, and precision remained at 2.5%. Scope is narrow in other ways too: only protonated-molecule ionisation was analysed, the rules derive from natural products and may not transfer to synthetic compounds, and MACCS keys are coarse descriptors reported as bit indices rather than named chemical moieties.

How this research was produced

K-Dense Web planned and ran this mass spectrometry investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.

Share:
Mass Spectrometry

MoNA MS/MS ClassyFire Classifier

Classify MoNA MS/MS spectra into ClassyFire chemical taxonomy with hierarchical label evaluation.

Drug Discovery

Natural Products Chemical Space

Cluster 715K+ compounds from the COCONUT database to identify bacterial-derived drug candidates using ML dimensionality reduction.

Genomics

Longevity Gene Analysis

Map model organism longevity genes to human orthologs and evaluate their validation through GWAS studies.

Run this kind of analysis on your own question

Try K-Dense Web free and see how an AI co-scientist accelerates your research.