Skip to main content
Drug Discovery· 51-page report· 1 figure

EGFR pIC50 QSAR Prediction Model

Build an EGFR pIC50 QSAR model for kinase inhibitor potency prediction with graphical abstract reporting.

What this research found

EGFR is one of the most heavily prosecuted oncology drug targets, and its long medicinal-chemistry record makes it a good test of how well a potency model generalizes to chemistry it has never seen. Bioactivity records for the human receptor were curated from ChEMBL into 10,667 unique compounds, then split so that no ring scaffold appeared in both training and test sets. On the 2,133 held-out compounds the selected random forest reached a Spearman rank correlation of 0.808 with a root-mean-squared error of 0.746 log units, ranking compounds reliably while systematically under-predicting the potency of the most active novel chemotypes.

  • On the scaffold-disjoint held-out set of 2,133 compounds the model reached RMSE 0.746 and MAE 0.561 in pIC50 units, R² of 0.646, and Spearman rank correlation of 0.808. Predictions land within 0.5 log units of the measurement for 55.9% of test compounds, within 1.0 unit for 83.1%, and within 1.5 units for 94.4%.
  • Ranking works far better than absolute calibration. The fitted predicted-versus-measured line has slope 0.64 rather than 1, the signature of regression toward the mean: weak compounds are over-predicted and potent ones under-predicted, with mean residual climbing from +0.35 above pIC50 7 to +0.54 above 8 and +0.88 above 9.
  • The random forest beat XGBoost consistently but narrowly under scaffold-grouped cross-validation (CV RMSE 0.827 ± 0.035 against 0.838 ± 0.034), and the narrow fold-to-fold spread indicates stable behaviour across scaffold folds.
  • The worst misses are potent compounds on scaffolds absent from training. Among the ten most potent held-out compounds — all sub-nanomolar at pIC50 ≥ 9.8 — seven were predicted within about 0.8 log units, while two on distinct frameworks were under-predicted by 1.86 and 1.79 units.
  • Chemical diversity in the curated set is heavily concentrated: 3,838 unique Bemis–Murcko scaffolds cover 10,667 compounds, 2,542 of those scaffolds appear only once, and the single 4-anilinoquinazoline core accounts for 575 compounds.
  • The report argues this pattern makes the model suitable for virtual screening, hit triage, and library prioritization, but recommends pairing point predictions with an applicability-domain flag — for instance the variance across individual trees — so that out-of-domain potent compounds are surfaced rather than silently down-ranked.

How it was done

19,378 IC50 activity records with a defined pChEMBL value were pulled for the human EGFR single-protein target from the ChEMBL 37 release, and the harmonized potency labels were verified against an independent recomputation from the reported nanomolar values (mean absolute deviation 0.0022). All 10,734 distinct input structures were desalted to the largest covalent fragment, neutralized, and tautomer-canonicalized, and replicate measurements were averaged, leaving 10,667 unique compounds spanning pIC50 4.0 to 11.0. Whole Bemis–Murcko scaffold groups were assigned to partitions to give 8,534 training and 2,133 test compounds with zero ring-scaffold overlap, so held-out scores measure extrapolation to new chemotypes rather than interpolation among close analogues. Compounds were encoded as 2048-bit radius-2 Morgan fingerprints, a random forest and an XGBoost regressor were tuned by scaffold-grouped 5-fold cross-validation, and the lower-error model was retrained on the full training set and evaluated once on the untouched test set.

Data sources

  • ChEMBL 37 (released 1 May 2026) — target CHEMBL203, human EGFR; 19,378 IC50 records curated to 10,667 compounds
  • RDKit 2026.03.3 — structure standardization, Bemis–Murcko scaffold assignment, and Morgan fingerprints

Limitations

IC50 values aggregated across heterogeneous assays, cell lines, and wild-type versus mutant EGFR constructs carry roughly 0.5 to 1.0 log units of inter-laboratory variability, which puts a floor under the achievable error. The model also ignores mutation status, uses a single two-dimensional fingerprint representation, and — because the scaffold split is deliberately conservative — should be read as a lower bound relative to deployment on chemistry close to the training set.

How this research was produced

K-Dense Web planned and ran this drug discovery investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.

Share:
Drug Discovery

Natural Products Chemical Space

Cluster 715K+ compounds from the COCONUT database to identify bacterial-derived drug candidates using ML dimensionality reduction.

Drug Discovery

AqSolDB Aqueous Solubility QSAR Model

Model aqueous solubility on AqSolDB with QSPR methods and report predictive performance.

Genomics

Longevity Gene Analysis

Map model organism longevity genes to human orthologs and evaluate their validation through GWAS studies.

Run this kind of analysis on your own question

Try K-Dense Web free and see how an AI co-scientist accelerates your research.