Skip to main content
Cheminformatics· 19-page report· 8 figures

EGFR Chemotype–Mutation Potency Matrix

Integrate TCGA-LUAD, ChEMBL, and PDB data to rank EGFR inhibitor chemotypes against six sensitizing and resistance mutations with explicit go/no-go gates.

What this research found

Which chemical scaffolds should a preclinical EGFR resistance panel be built around? Three unrelated public resources — lung adenocarcinoma genomics from TCGA, EGFR bioactivity records from ChEMBL, and the osimertinib-bound EGFR kinase-domain crystal structure — were combined into a chemotype-by-mutation potency matrix and run through an explicit five-gate go/no-go framework. The 4-anilinoquinazoline scaffold came out best-in-class against sensitizing driver mutations but was disqualified for the resistance setting, where it loses 24.2-fold potency against T790M.

  • The 4-anilinoquinazoline core behind gefitinib and erlotinib is the most prominent and best-sampled EGFR chemotype, with 3,169 compounds on record. A pyrrolo[2,3-d]pyrimidine (7-deazapurine) series is more potent on average, at a mean biochemical pActivity of about 9.2, but is far more sparsely characterised.
  • Against sensitizing drivers the 4-anilinoquinazoline series reaches pActivity as high as 10.4, yet it loses 24.2-fold potency against the T790M gatekeeper mutation, a pActivity drop of 1.38. Cellular and biochemical measurements of the same chemotype-mutation pairs agree only moderately, at r = 0.55.
  • Six mechanistically motivated alleles were nominated for the panel from live prevalence across 585 TCGA lung adenocarcinoma cases: L858R, the exon-19 E746_A750 deletion, T790M, C797S, L861Q, and S768I.
  • Pocket geometry underwrites the resistance calls rather than leaving them as threshold artefacts. The Cys797 sulfur sits 3.35 Å from the inhibitor's acrylamide beta-carbon, a covalent trajectory that the C797S substitution abolishes, while the gatekeeper Thr790 lies 4.68 Å from the ligand and Met793 forms hinge hydrogen bonds.
  • Verdicts were resolved by indication instead of collapsed into one score, which preserves the finding that a scaffold can be simultaneously best-in-class and disqualified: GO for the 4-anilinoquinazoline series against sensitizing mutations, NO-GO against resistance, and NO-GO on both counts for the pyrrolo[2,3-d]pyrimidine alternative pending dedicated mutant-panel profiling. Composite scores were 8.97 and 0.29.

How it was done

Somatic mutation and RNA-seq data for lung adenocarcinoma were retrieved from the NCI Genomic Data Commons at Data Release 45.0 to rank candidate alleles by prevalence; EGFR bioactivity records were pulled from ChEMBL 37 for target CHEMBL203 and grouped into scaffold families; and the osimertinib-bound EGFR kinase domain from the Protein Data Bank supplied binding-pocket geometry. Potencies were aggregated on a log scale into a chemotype-by-mutation matrix, 71.4% of whose cells could be populated, and each candidate series was then passed through a five-gate go/no-go framework whose thresholds were fixed numerically before being applied. Every dataset was fetched through public interfaces, and the write-up ran to a 19-page paper with seven figures, five tables, and 38 references checked against publisher metadata.

Data sources

  • NCI Genomic Data Commons — TCGA lung adenocarcinoma somatic mutations and RNA-seq, Data Release 45.0 (585 cases)
  • ChEMBL 37 — EGFR bioactivity records for target CHEMBL203 (3,169 compounds in the leading chemotype)
  • Protein Data Bank entry 4ZAU — osimertinib-bound EGFR kinase domain

Limitations

Assay heterogeneity across sources, a 69% share of bioactivity records with unclassified mutation status, and the fact that reported T790M shifts come from double-mutant compound contexts all constrain the matrix; the C797S biochemical cell rests on three records from two compounds and is flagged as not robustly characterised. The work is target-discovery cheminformatics and makes no patient-treatment or dosing recommendations.

Figures from this analysis

How this research was produced

K-Dense Web planned and ran this cheminformatics investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.

Share:
Cheminformatics

FGFR2 Inhibitor SAR Mining from Patent WO2020231990

A reproducible six-step pipeline that extracts 4-amino-pyrrolopyrimidine compounds from an image-only patent table, fuses ligand descriptors with FGFR2 co-crystal pocket geometry, and trains an explainable Random Forest model to surface the structural drivers of kinase activity.

Drug Discovery

Natural Products Chemical Space

Cluster 715K+ compounds from the COCONUT database to identify bacterial-derived drug candidates using ML dimensionality reduction.

Genomics

Longevity Gene Analysis

Map model organism longevity genes to human orthologs and evaluate their validation through GWAS studies.

Run this kind of analysis on your own question

Try K-Dense Web free and see how an AI co-scientist accelerates your research.