Skip to main content
Cheminformatics· 19-page report· 8 figures

EGFR Chemotype–Mutation Potency Matrix

Integrate TCGA-LUAD, ChEMBL, and PDB data to rank EGFR inhibitor chemotypes against six sensitizing and resistance mutations with explicit go/no-go gates.

What this research found

Which chemical scaffolds should a preclinical EGFR resistance panel be built around? Three unrelated public resources — lung adenocarcinoma genomics from TCGA, EGFR bioactivity records from ChEMBL, and the osimertinib-bound EGFR kinase-domain crystal structure — were combined into a chemotype-by-mutation potency matrix and run through an explicit five-gate go/no-go framework. The 4-anilinoquinazoline scaffold came out best-in-class against sensitizing driver mutations but was disqualified for the resistance setting, where it loses 24.2-fold potency against T790M.

  • The 4-anilinoquinazoline core behind gefitinib and erlotinib is the most prominent and best-sampled EGFR chemotype, with 3,169 compounds on record. A pyrrolo[2,3-d]pyrimidine (7-deazapurine) series is more potent on average, at a mean biochemical pActivity of about 9.2, but is far more sparsely characterised.
  • Against sensitizing drivers the 4-anilinoquinazoline series reaches pActivity as high as 10.4, yet it loses 24.2-fold potency against the T790M gatekeeper mutation, a pActivity drop of 1.38. Cellular and biochemical measurements of the same chemotype-mutation pairs agree only moderately, at r = 0.55.
  • Six mechanistically motivated alleles were nominated for the panel from live prevalence across 585 TCGA lung adenocarcinoma cases: L858R, the exon-19 E746_A750 deletion, T790M, C797S, L861Q, and S768I.
  • Pocket geometry underwrites the resistance calls rather than leaving them as threshold artefacts. The Cys797 sulfur sits 3.35 Å from the inhibitor's acrylamide beta-carbon, a covalent trajectory that the C797S substitution abolishes, while the gatekeeper Thr790 lies 4.68 Å from the ligand and Met793 forms hinge hydrogen bonds.
  • Verdicts were resolved by indication instead of collapsed into one score, which preserves the finding that a scaffold can be simultaneously best-in-class and disqualified: GO for the 4-anilinoquinazoline series against sensitizing mutations, NO-GO against resistance, and NO-GO on both counts for the pyrrolo[2,3-d]pyrimidine alternative pending dedicated mutant-panel profiling. Composite scores were 8.97 and 0.29.

How it was done

Somatic mutation and RNA-seq data for lung adenocarcinoma were retrieved from the NCI Genomic Data Commons at Data Release 45.0 to rank candidate alleles by prevalence; EGFR bioactivity records were pulled from ChEMBL 37 for target CHEMBL203 and grouped into scaffold families; and the osimertinib-bound EGFR kinase domain from the Protein Data Bank supplied binding-pocket geometry. Potencies were aggregated on a log scale into a chemotype-by-mutation matrix, 71.4% of whose cells could be populated, and each candidate series was then passed through a five-gate go/no-go framework whose thresholds were fixed numerically before being applied. Every dataset was fetched through public interfaces, and the write-up ran to a 19-page paper with seven figures, five tables, and 38 references checked against publisher metadata.

Data sources

  • NCI Genomic Data Commons — TCGA lung adenocarcinoma somatic mutations and RNA-seq, Data Release 45.0 (585 cases)
  • ChEMBL 37 — EGFR bioactivity records for target CHEMBL203 (3,169 compounds in the leading chemotype)
  • Protein Data Bank entry 4ZAU — osimertinib-bound EGFR kinase domain

Limitations

Assay heterogeneity across sources, a 69% share of bioactivity records with unclassified mutation status, and the fact that reported T790M shifts come from double-mutant compound contexts all constrain the matrix; the C797S biochemical cell rests on three records from two compounds and is flagged as not robustly characterised. The work is target-discovery cheminformatics and makes no patient-treatment or dosing recommendations.

Figures from this analysis

How this research was produced

K-Dense Web planned and ran this cheminformatics investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.

Share:
Cheminformatics

FGFR2 Inhibitor SAR Mining from Patent WO2020231990

A reproducible six-step pipeline that extracts 4-amino-pyrrolopyrimidine compounds from an image-only patent table, fuses ligand descriptors with FGFR2 co-crystal pocket geometry, and trains an explainable Random Forest model to surface the structural drivers of kinase activity.

Drug Discovery

Natural Products Chemical Space

Cluster 715K+ compounds from the COCONUT database to identify bacterial-derived drug candidates using ML dimensionality reduction.

5.8 weeks → 45 min406× cheaper than doing it alone
Genomics

Longevity Gene Analysis

Map model organism longevity genes to human orthologs and evaluate their validation through GWAS studies.

4.7 weeks → 75 min658× cheaper than doing it alone

Run this kind of analysis on your own question

Start a session and see how an AI co-scientist accelerates your research. Pay-as-you-go, no subscription required.