Skip to main content
Mass Spectrometry· 18-page report· 1 figure

MoNA MS/MS ClassyFire Classifier

Classify MoNA MS/MS spectra into ClassyFire chemical taxonomy with hierarchical label evaluation.

What this research found

Most tandem mass spectra collected in untargeted metabolomics are never matched to a structure, so assigning a spectrum to a broad chemical class is a useful fallback. This work tests how far that can be pushed using nothing but the raw fragmentation peak list — no molecular formula, no precursor mass, no retention time — with a classical tree ensemble and a split that guarantees no test compound was ever seen during training. On 739 held-out compounds from the MassBank of North America library, the correct ClassyFire superclass was predicted 50.3% of the time with a macro-averaged F1 of 0.373, against 23.3% for always guessing the largest class.

  • Across 739 held-out compounds the classifier reached 50.3% accuracy and a macro-averaged F1 of 0.373, roughly 2.2 times the 23.3% majority-class baseline. Scored per spectrum rather than per compound, which favours heavily replicated molecules, accuracy rises to 55.5% and macro-F1 to 0.412.
  • Results vary sharply by class. Phenylpropanoids and polyketides (F1 0.651), lipids and lipid-like molecules (0.617), organic acids and derivatives (0.548) and nucleosides and nucleotides (0.531) come out best — the categories whose members share pronounced, recurring fragmentation motifs.
  • Three superclasses present in the test set were never recovered at all, scoring an F1 of zero: lignans, organic nitrogen compounds, and organosulfur compounds. Each has only a handful of training compounds, and balanced class weighting did not compensate for that scarcity.
  • The errors track chemical adjacency rather than noise. Benzenoids and organoheterocyclic compounds account for 50 misclassifications between them, 30 one way and 20 the other, and organic acids with organoheterocyclics for a further 31. Both pairs share aromatic fragments and small neutral losses, and organoheterocyclics, the broadest superclass, acts as an attractor for the rest.
  • Assembling a clean training set costs most of the library. Of 102,580 records in the source export, 20,403 spectra survived the requirement of positive-mode tandem acquisition from a protonated precursor with a valid structure identifier and a class annotation, covering 3,695 unique compounds across 16 superclasses.

How it was done

Records were streamed from a dated MassBank of North America positive-mode export and kept only where the spectrum was tandem, positive-mode, from a protonated precursor, and carried both a valid structure identifier and a ClassyFire annotation. Each surviving spectrum had peaks below 1% of its base peak discarded, was rescaled so the base peak equals one, and was binned onto a 1 Da grid from 50 to 1000 Da, giving 950 features that were then TF-IDF weighted so that fragment masses appearing everywhere count for less than distinctive ones, with the weighting fit on training spectra only. Compounds were grouped by the 14-character InChIKey skeleton so that stereoisomers and protonation variants of one molecule stay together, then split 80/20 with a fixed seed and verified to share nothing across the boundary. Extremely randomised trees were selected over a random forest by grouped cross-validation and tuned, and each compound's prediction was formed by averaging the class probabilities across all of its spectra.

Data sources

  • MassBank of North America, LC-MS/MS Positive Mode export generated 6 July 2026 — 102,580 records, filtered to 20,403 spectra over 3,695 compounds and 16 superclasses, with parsed collision energies from 0 to 80 eV
  • ClassyFire and the ChemOnt taxonomy, Djoumbou Feunang et al., Journal of Cheminformatics 8:61 (2016) — structure-derived superclass labels

Limitations

The inputs are deliberately impoverished: precursor mass, collision energy, molecular formula and fragmentation trees are all withheld, and the 1 Da binning discards the sub-Dalton accurate-mass information that formula-based tools rely on, so these figures are a baseline rather than a ceiling. The library also mixes instruments and collision energies spanning 0 to 80 eV, so one molecule can present markedly different fragmentation patterns, and the coarse superclass labels themselves embed the aromatic-versus-heterocyclic ambiguity that drives the leading confusions.

How this research was produced

K-Dense Web planned and ran this mass spectrometry investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.

Share:
Mass Spectrometry

MS Fragmentation Rule Mining

Discover fragmentation patterns from 1,267 GNPS spectra with FDR-corrected statistical enrichment analysis.

Drug Discovery

Natural Products Chemical Space

Cluster 715K+ compounds from the COCONUT database to identify bacterial-derived drug candidates using ML dimensionality reduction.

Genomics

Longevity Gene Analysis

Map model organism longevity genes to human orthologs and evaluate their validation through GWAS studies.

Run this kind of analysis on your own question

Try K-Dense Web free and see how an AI co-scientist accelerates your research.