What this research found
Chromatographic retention time is recorded for free in every liquid-chromatography mass-spectrometry run and encodes properties, such as hydrophobicity and polar surface area, that mass spectrometry alone cannot resolve — which makes it a useful extra filter when identifying unknown metabolites. A gradient-boosted tree model was trained to predict it from molecular structure alone using 78,629 curated compounds from the METLIN small-molecule retention-time dataset. On a held-out fifth of the compounds it reached a mean absolute error of 45.99 seconds, a median absolute error of 25.09 seconds, and an R² of 0.842, with 79.2% of compounds predicted within 60 seconds of the measured value.
- On the 15,726 held-out compounds the model reached a mean absolute error of 45.99 seconds and a median of 25.09 seconds, with an R² of 0.842. Tightening or relaxing the tolerance gives 56.6% of compounds within 30 seconds, 79.2% within 60 seconds, and 92.4% within 120 seconds.
- The gap between mean and median error comes almost entirely from one subpopulation. The 395 poorly retained compounds eluting within the first 300 seconds have a subgroup mean absolute error of 342.02 seconds and only 5.3% inside the 60-second window, and they are over-predicted almost uniformly, with a mean residual of −335.7 seconds.
- Excluding those near-void analytes and scoring only the 15,331 well-retained compounds above 400 seconds drops the mean absolute error to 38.42 seconds and lifts the within-60-second fraction to 81.1%, which quantifies what an applicability-domain cut-off actually buys.
- What the model learned matches established reversed-phase chemistry. Physicochemical descriptors carry 74.6% of the total split gain despite being only 211 of 2,259 features, with the Wildman–Crippen octanol–water partition estimate alone accounting for 6.25% — more than three times the next feature — followed by partial-charge- and surface-area-weighted polarity terms.
- Combining structural and physicochemical representations beat either on its own. In cross-validation, gradient boosting on the concatenated features gave a mean absolute error of 50.40 seconds against 56.66 for descriptors alone, 58.52 for the fingerprint alone, and 63.65 for the strongest ridge-regression baseline.
How it was done
The raw 80,038-entry release was curated down to 78,629 compounds: 81 structures failed to parse, 1,299 were dropped using a published 2024 exclusion list of potentially erroneous entries, and 29 rows went in resolving duplicate structures, where replicates agreeing within 30 seconds were collapsed to a median and discordant ones discarded. Each molecule was encoded as a 2048-bit radius-2 Morgan fingerprint concatenated with 211 standardised physicochemical descriptors. Compounds were split 80/20 at the structure level with a fixed seed, and a LightGBM regressor trained on an absolute-error objective was chosen over ridge regression and single-representation baselines by three-fold cross-validation, then tuned over learning rate and leaf count before being evaluated exactly once on the untouched hold-out set. Split-gain attribution and residual analysis were then used to interpret the fit and to map where it breaks down.
Data sources
- METLIN small-molecule retention-time dataset, Domingo-Almenara et al., Nature Communications 10:5811 (2019) — 80,038 authentic compounds measured under one reversed-phase gradient, retention times from 0.3 to 1471.7 seconds
- Khrisanfov, Matyushin & Samokhin, Journal of Chromatography A 1745:465761 (2025) — exclusion list of 1,299 potentially erroneous dataset entries
- RDKit 2026.03.3 — Morgan fingerprints and physicochemical descriptors
Limitations
The model predicts retention only for the single reversed-phase gradient the dataset was acquired on, and transferring it to another column or gradient requires projection or calibration. Its two-dimensional representation cannot distinguish stereoisomers or conformers that share a flat structure, and its applicability domain is bounded on the low-retention side, where predicted times should not be used to reject candidate identifications.
How this research was produced
K-Dense Web planned and ran this chemistry investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.


