Skip to main content
Drug Discovery· 20-page report· 10 figures

Natural Products Chemical Space

Cluster 715K+ compounds from the COCONUT database to identify bacterial-derived drug candidates using ML dimensionality reduction.

What this research found

The COCONUT natural products database was mined for bacterial-derived compounds worth putting into an antimicrobial screening library. Because the ChEMBL bioactivity data needed for a supervised model was unreachable, the work pivoted to unsupervised clustering of chemical space, which split the bacterial subset cleanly into drug-like and structurally complex families. Fifty candidates were then drawn from both ends deliberately: 25 highly drug-like leads and 25 large, complex scaffolds of the kind that has historically yielded antibiotics.

  • Bacterial sources make up a small slice of the database. Filtering 715,822 COCONUT compounds by organism keywords left 24,911 entries, or 3.48% of the original, of which 24,910 passed structure validation with no duplicates removed.
  • The bacterial subset splits in two. K-Means on ten molecular descriptors found an optimal cluster count of 2 with a silhouette score of 0.3910: a drug-like cluster of 18,780 compounds (75.4%, mean molecular weight 396 Da, mean quantitative drug-likeness score of 0.44) and a complex cluster of 6,130 (24.6%, mean 978 Da, mean score 0.09).
  • Most bacterial natural products fail conventional drug-likeness filters. Only 44.8% (11,151 of 24,910) satisfy Lipinski's Rule of Five, 13.7% carry pan-assay interference substructures, and just 39.7% pass both checks.
  • The two selected candidate groups sit at opposite extremes. Group A averages 322.09 Da and a drug-likeness score of 0.927 with 100% Lipinski compliance, while Group B averages 1,929.73 Da and a score of 0.047, with 17.36 rings and 10.88 aromatic rings per molecule and no Lipinski compliance at all.
  • Property ranges across the subset are very wide: molecular weight runs from 1.008 to 4,899.957 Da around a mean of 539.124, rotatable bonds reach 150, and drug-likeness scores span 0.006 to 0.943.

How it was done

The December 2025 COCONUT release was filtered to bacterial-origin entries using eleven organism keywords including streptomyces, bacillus, and mycobacterium, then validated and deduplicated. Thirty-six features were computed per compound, covering physicochemical properties (molecular weight, LogP, polar surface area, hydrogen-bond donors and acceptors), structural descriptors (ring counts, aromatic rings, rotatable bonds), and drug-likeness measures (Lipinski compliance, quantitative drug-likeness, interference-substructure flags), alongside biosynthetic pathway annotations from NPClassifier. Ten descriptors fed a principal component analysis retaining over 80% of variance in ten components, with K-Means for clustering, DBSCAN for outlier detection, and t-SNE and UMAP for visualisation. Candidates were drawn separately from each cluster — the highest-scoring Lipinski-compliant molecules from the drug-like side, and the highest ring-count-by-aromatic-ring complexity scores from the other — and written up as a 20-page manuscript with 11 figures and 27 verified citations.

Data sources

  • COCONUT natural products database, December 2025 release — 715,822 compounds, 24,910 bacterial-derived retained
  • NPClassifier — biosynthetic pathway annotations for the bacterial subset
  • ChEMBL bioactivity data — intended as supervised training labels, unavailable during the run

Limitations

Nothing here was tested at the bench: the ranking rests on computed rather than measured properties, with no target docking or activity prediction behind it. The unsupervised route was itself a fallback, adopted because the bioactivity labels for a supervised model could not be retrieved.

Figures from this analysis

PCA scree plot showing explained variance
PCA feature loadings biplot
PCA scatter plot with K-Means cluster labels
UMAP scatter plot with K-Means cluster labels
PCA scatter plot with DBSCAN cluster labels
Heatmap of normalized cluster feature profiles
PCA visualization of selected candidates in chemical space
Property distribution comparison (MW, LogP, QED) between Group A and Group B

Outputs produced

  • coconut_features.csv

    Feature-enriched dataset with molecular descriptors (24,910 compounds × 36 features)

    42 MB
  • chemical_space_analysis_results.csv

    Full dataset with cluster labels and PCA coordinates (24,910 compounds)

    50 MB
  • prioritized_candidates.csv

    Final prioritized candidate list (50 compounds: 25 drug-like, 25 complex scaffolds)

  • final_research_paper.pdf

    Comprehensive formal research paper with Abstract, Methods, Results, Discussion, figures, and tables

  • coconut_csv-12-2025.csv

    Raw COCONUT database (715,822 compounds)

    664 MB
  • coconut_cleaned.csv

    Cleaned bacterial natural products dataset (24,910 compounds)

    41 MB
  • cluster_statistics.csv

    Cluster characteristics and mean feature values

  • data_prep_summary.txt

    Data preparation summary report

How this research was produced

K-Dense Web planned and ran this drug discovery investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.

Share:
Drug Discovery

EGFR pIC50 QSAR Prediction Model

Build an EGFR pIC50 QSAR model for kinase inhibitor potency prediction with graphical abstract reporting.

Drug Discovery

AqSolDB Aqueous Solubility QSAR Model

Model aqueous solubility on AqSolDB with QSPR methods and report predictive performance.

Genomics

Longevity Gene Analysis

Map model organism longevity genes to human orthologs and evaluate their validation through GWAS studies.

Run this kind of analysis on your own question

Try K-Dense Web free and see how an AI co-scientist accelerates your research.