What this research found
A publicly deposited targeted serum metabolomics dataset of 224 participants was re-analysed to ask how well circulating small molecules separate colorectal cancer from health. Holding the 76 colorectal-polyp samples aside left 64 cancer and 84 healthy sera, and two classifiers built on 113 named metabolites reached out-of-fold areas under the ROC curve of 0.807 for an L1-regularized logistic regression and 0.813 for a random forest. The metabolites ranked highest — depleted histidine alongside elevated glutamate, proline, hydroxyproline, kynurenate, ornithine, conjugated bile acids and hippurate — match alterations reported independently for colorectal cancer.
- Both classifiers reached moderate and closely matched discrimination on out-of-fold predictions: AUROC 0.813 for the random forest and 0.807 for the Lasso, with areas under the precision–recall curve of 0.806 and 0.803 against a prevalence baseline of 0.432. Their near-superimposable curves suggest the separability is largely additive rather than driven by higher-order interactions.
- At a threshold tuned to 90% specificity, an operating point suited to a screening adjunct, the Lasso detected 64.1% of cancers against 53.1% for the random forest. At the default 0.5 threshold the Lasso was better balanced (accuracy 0.757, sensitivity 0.734, specificity 0.774) while the random forest was more conservative (specificity 0.857, sensitivity 0.563).
- Histidine was by far the strongest single feature and the dominant marker of the healthy class, with a mean Lasso coefficient of −2.27 and non-zero selection in 100% of folds. It was significantly depleted in cancer (log2 fold change −0.28, false-discovery rate 0.007).
- Sixteen of the 113 metabolites reached significance at a 5% false-discovery rate, and the same 16 held up under both Welch's t-test and the rank-based Mann–Whitney test. The largest effects were all increases in cancer: glycocholate (log2 fold change +1.61), glycochenodeoxycholate (+1.26) and hippuric acid (+1.07).
- Of the 30 profiled metabolites carrying a directional report in the prior colorectal cancer literature, 20 (67%) moved in the previously reported direction. Every one of the ten discordant cases was statistically non-significant with a near-zero fold change.
How it was done
The named-metabolite table from Metabolomics Workbench study ST000284 — 224 serum samples by 113 metabolites measured by liquid chromatography–tandem mass spectrometry — came from the repository's public API. The 76 polyp samples were held aside rather than folded into either class, since polyps are an intermediate state on the adenoma–carcinoma continuum, leaving a 148-participant contrast at 0.432 cancer prevalence. Abundances went through half-minimum imputation, sample-wise median normalization, log2 transformation and per-metabolite auto-scaling; principal-component analysis before and after showed the first five components falling from 88% to 49% of variance as dominant technical scale was removed. An L1-regularized logistic regression, its penalty tuned by an inner loop inside each training fold, and a 100-tree random forest were evaluated by outer stratified 5-fold cross-validation, so no sample contributed to the model that scored it. Metabolites were ranked by mean model weight across folds and by univariate statistics computed on the pre-scaling data.
Data sources
- Metabolomics Workbench study ST000284 — targeted serum metabolic profiling for colorectal cancer detection, University of Washington (224 samples × 113 named metabolites, CC BY 4.0)
- Zhu et al., Journal of Proteome Research 13:4120 (2014) — the originating profiling study
- Cross et al., Cancer 120:3049 (2014) — prospective study of serum metabolites and colorectal cancer risk
- Geijsen et al., International Journal of Cancer 145:1221 (2019) — discovery–replication plasma metabolite study
- Curated directional annotations for 30 previously reported colorectal cancer serum or plasma markers
Limitations
The cohort is single-centre and modest at 148 participants, the targeted 113-metabolite panel is narrower than contemporary untargeted platforms, and clinical covariates such as stage, tumour location, age, sex, diet and medication were not modelled. Being cross-sectional case–control data, it cannot establish whether these metabolic changes precede diagnosis or arise as a consequence of established disease.
How this research was produced
K-Dense Web planned and ran this metabolomics investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.


