Skip to main content
Chemistry· 18-page report· 1 figure

METLIN Retention-Time QSAR Prediction

Predict METLIN chromatographic retention times with QSAR models against observed values.

What this research found

Chromatographic retention time is recorded for free in every liquid-chromatography mass-spectrometry run and encodes properties, such as hydrophobicity and polar surface area, that mass spectrometry alone cannot resolve — which makes it a useful extra filter when identifying unknown metabolites. A gradient-boosted tree model was trained to predict it from molecular structure alone using 78,629 curated compounds from the METLIN small-molecule retention-time dataset. On a held-out fifth of the compounds it reached a mean absolute error of 45.99 seconds, a median absolute error of 25.09 seconds, and an R² of 0.842, with 79.2% of compounds predicted within 60 seconds of the measured value.

  • On the 15,726 held-out compounds the model reached a mean absolute error of 45.99 seconds and a median of 25.09 seconds, with an R² of 0.842. Tightening or relaxing the tolerance gives 56.6% of compounds within 30 seconds, 79.2% within 60 seconds, and 92.4% within 120 seconds.
  • The gap between mean and median error comes almost entirely from one subpopulation. The 395 poorly retained compounds eluting within the first 300 seconds have a subgroup mean absolute error of 342.02 seconds and only 5.3% inside the 60-second window, and they are over-predicted almost uniformly, with a mean residual of −335.7 seconds.
  • Excluding those near-void analytes and scoring only the 15,331 well-retained compounds above 400 seconds drops the mean absolute error to 38.42 seconds and lifts the within-60-second fraction to 81.1%, which quantifies what an applicability-domain cut-off actually buys.
  • What the model learned matches established reversed-phase chemistry. Physicochemical descriptors carry 74.6% of the total split gain despite being only 211 of 2,259 features, with the Wildman–Crippen octanol–water partition estimate alone accounting for 6.25% — more than three times the next feature — followed by partial-charge- and surface-area-weighted polarity terms.
  • Combining structural and physicochemical representations beat either on its own. In cross-validation, gradient boosting on the concatenated features gave a mean absolute error of 50.40 seconds against 56.66 for descriptors alone, 58.52 for the fingerprint alone, and 63.65 for the strongest ridge-regression baseline.

How it was done

The raw 80,038-entry release was curated down to 78,629 compounds: 81 structures failed to parse, 1,299 were dropped using a published 2024 exclusion list of potentially erroneous entries, and 29 rows went in resolving duplicate structures, where replicates agreeing within 30 seconds were collapsed to a median and discordant ones discarded. Each molecule was encoded as a 2048-bit radius-2 Morgan fingerprint concatenated with 211 standardised physicochemical descriptors. Compounds were split 80/20 at the structure level with a fixed seed, and a LightGBM regressor trained on an absolute-error objective was chosen over ridge regression and single-representation baselines by three-fold cross-validation, then tuned over learning rate and leaf count before being evaluated exactly once on the untouched hold-out set. Split-gain attribution and residual analysis were then used to interpret the fit and to map where it breaks down.

Data sources

  • METLIN small-molecule retention-time dataset, Domingo-Almenara et al., Nature Communications 10:5811 (2019) — 80,038 authentic compounds measured under one reversed-phase gradient, retention times from 0.3 to 1471.7 seconds
  • Khrisanfov, Matyushin & Samokhin, Journal of Chromatography A 1745:465761 (2025) — exclusion list of 1,299 potentially erroneous dataset entries
  • RDKit 2026.03.3 — Morgan fingerprints and physicochemical descriptors

Limitations

The model predicts retention only for the single reversed-phase gradient the dataset was acquired on, and transferring it to another column or gradient requires projection or calibration. Its two-dimensional representation cannot distinguish stereoisomers or conformers that share a flat structure, and its applicability domain is bounded on the low-retention side, where predicted times should not be used to reject candidate identifications.

How this research was produced

K-Dense Web planned and ran this chemistry investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.

Share:
Chemistry

nmrshiftdb2 NMR Structure Elucidation

Elucidate molecular structures from nmrshiftdb2 NMR shifts using a multi-stage candidate funnel.

Chemistry

SAMPL6 First-Principles pKa Prediction

Predict SAMPL6 pKa values from first principles and correlate predictions against experimental benchmarks.

Chemistry

QM9 HOMO-LUMO Gap Prediction

Predict QM9 HOMO–LUMO gaps with ML regression and validate predicted-versus-true parity plots.

Run this kind of analysis on your own question

Try K-Dense Web free and see how an AI co-scientist accelerates your research.