What this research found
The COCONUT natural products database was mined for bacterial-derived compounds worth putting into an antimicrobial screening library. Because the ChEMBL bioactivity data needed for a supervised model was unreachable, the work pivoted to unsupervised clustering of chemical space, which split the bacterial subset cleanly into drug-like and structurally complex families. Fifty candidates were then drawn from both ends deliberately: 25 highly drug-like leads and 25 large, complex scaffolds of the kind that has historically yielded antibiotics.
- Bacterial sources make up a small slice of the database. Filtering 715,822 COCONUT compounds by organism keywords left 24,911 entries, or 3.48% of the original, of which 24,910 passed structure validation with no duplicates removed.
- The bacterial subset splits in two. K-Means on ten molecular descriptors found an optimal cluster count of 2 with a silhouette score of 0.3910: a drug-like cluster of 18,780 compounds (75.4%, mean molecular weight 396 Da, mean quantitative drug-likeness score of 0.44) and a complex cluster of 6,130 (24.6%, mean 978 Da, mean score 0.09).
- Most bacterial natural products fail conventional drug-likeness filters. Only 44.8% (11,151 of 24,910) satisfy Lipinski's Rule of Five, 13.7% carry pan-assay interference substructures, and just 39.7% pass both checks.
- The two selected candidate groups sit at opposite extremes. Group A averages 322.09 Da and a drug-likeness score of 0.927 with 100% Lipinski compliance, while Group B averages 1,929.73 Da and a score of 0.047, with 17.36 rings and 10.88 aromatic rings per molecule and no Lipinski compliance at all.
- Property ranges across the subset are very wide: molecular weight runs from 1.008 to 4,899.957 Da around a mean of 539.124, rotatable bonds reach 150, and drug-likeness scores span 0.006 to 0.943.
How it was done
The December 2025 COCONUT release was filtered to bacterial-origin entries using eleven organism keywords including streptomyces, bacillus, and mycobacterium, then validated and deduplicated. Thirty-six features were computed per compound, covering physicochemical properties (molecular weight, LogP, polar surface area, hydrogen-bond donors and acceptors), structural descriptors (ring counts, aromatic rings, rotatable bonds), and drug-likeness measures (Lipinski compliance, quantitative drug-likeness, interference-substructure flags), alongside biosynthetic pathway annotations from NPClassifier. Ten descriptors fed a principal component analysis retaining over 80% of variance in ten components, with K-Means for clustering, DBSCAN for outlier detection, and t-SNE and UMAP for visualisation. Candidates were drawn separately from each cluster — the highest-scoring Lipinski-compliant molecules from the drug-like side, and the highest ring-count-by-aromatic-ring complexity scores from the other — and written up as a 20-page manuscript with 11 figures and 27 verified citations.
Data sources
- COCONUT natural products database, December 2025 release — 715,822 compounds, 24,910 bacterial-derived retained
- NPClassifier — biosynthetic pathway annotations for the bacterial subset
- ChEMBL bioactivity data — intended as supervised training labels, unavailable during the run
Limitations
Nothing here was tested at the bench: the ranking rests on computed rather than measured properties, with no target docking or activity prediction behind it. The unsupervised route was itself a fallback, adopted because the bioactivity labels for a supervised model could not be retrieved.
Figures from this analysis
Outputs produced
- 42 MB
coconut_features.csv
Feature-enriched dataset with molecular descriptors (24,910 compounds × 36 features)
- 50 MB
chemical_space_analysis_results.csv
Full dataset with cluster labels and PCA coordinates (24,910 compounds)
prioritized_candidates.csv
Final prioritized candidate list (50 compounds: 25 drug-like, 25 complex scaffolds)
Comprehensive formal research paper with Abstract, Methods, Results, Discussion, figures, and tables
- 664 MB
coconut_csv-12-2025.csv
Raw COCONUT database (715,822 compounds)
- 41 MB
coconut_cleaned.csv
Cleaned bacterial natural products dataset (24,910 compounds)
cluster_statistics.csv
Cluster characteristics and mean feature values
Data preparation summary report
How this research was produced
K-Dense Web planned and ran this drug discovery investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.


