Skip to main content
Cancer Biology· 6 figures

HNSCC Treatment Response Biomarkers

Build ML classifier for head and neck cancer treatment response using 270 patient samples with AUC=0.76.

What this research found

Head and neck squamous cell carcinoma (HNSCC) responds unpredictably to treatment, so tumour expression profiles from 270 patients were searched for a signature separating those whose disease progressed from those whose did not. No single gene survived multiple-testing correction — 0 of 31,330 genes reached a false discovery rate below 0.05 — yet a random forest built on the 20 top-ranked genes classified held-out patients with a ROC-AUC of 0.7558. Whatever signal exists here is combinatorial and weak rather than attributable to any individual marker.

  • No individual gene distinguished the two outcome groups after correction. Zero of 31,330 genes reached a false discovery rate below 0.05, 563 (1.8%) passed an uncorrected p-value of 0.01, and fold changes among the top 20 spanned only 0.03 to 0.34 on a log2 scale.
  • A random forest on those 20 genes reached a ROC-AUC of 0.7558, accuracy of 0.6852, sensitivity of 0.7037 and specificity of 0.6667 on 54 held-out patients, correctly calling 19 of 27 who progressed and 18 of 27 who did not.
  • Cross-validated performance was lower than the held-out figure at a ROC-AUC of 0.7019 ± 0.0550, and the gap against a mean training score of 0.9968 indicates the model overfits its training data.
  • The three most informative features were LOC645946, an uncharacterised locus, at an importance of 0.078; SCYE1, also known as EMAP-II, a pro-inflammatory cytokine with context-dependent effects on tumour angiogenesis, at 0.071; and C9ORF126, now annotated as NMRK1, an enzyme in NAD+ salvage biosynthesis, at 0.065.
  • Outcome labels were derived from progression-free survival rather than radiological assessment: 137 patients with no progression event during follow-up were treated as complete responders and 133 with a progression event as progressive disease, a near-even 50.7% to 49.3% split.

How it was done

Head and neck cancer expression data was retrieved from the Gene Expression Omnibus under accession GSE65858 — 270 tumours across 31,330 genes, already log2-normalised. No India-specific HNSCC dataset carrying treatment response data was found in the repository, so this general cohort was used instead. Patients were labelled by progression-free survival event, Welch's t-test was run across every gene with Benjamini-Hochberg correction, and the 20 genes with the lowest raw p-values were carried forward as model features. A random forest was tuned by grid search over 24 hyperparameter combinations under 5-fold cross-validation, then evaluated on a held-out set of 54 patients, and the highest-ranked features were annotated for biological plausibility. Supporting outputs included a volcano plot, a hierarchically clustered heatmap of the 20 genes across all 270 samples, and a written paper.

Data sources

  • Gene Expression Omnibus accession GSE65858 — 270 head and neck cancer tumours profiled across 31,330 genes on the Illumina HumanHT-12 V4.0 platform (GPL10558), with 137 non-progressors and 133 progressors

Limitations

The 20 features were chosen by differential expression across the entire dataset, including the samples later held out for testing, so the reported test metrics are an upper bound rather than an unbiased estimate of generalisation. The cohort is also not India-specific as originally intended, and treatment response was inferred from progression events rather than directly measured.

Figures from this analysis

Confusion matrix of test set predictions (PNG, 300 DPI)
Feature importance bar plot for top 10 biomarkers (PNG, 300 DPI)
Hyperparameter optimization visualization
Heatmap of top 20 genes with hierarchical clustering (PNG)
Volcano plot of differential expression (PNG, 300 DPI)

Outputs produced

  • clinical_data.csv

    Clinical metadata for 270 HNSCC samples with CR/PD classification

    0.081 MB
  • expression_data.csv

    Normalized gene expression matrix (log2-transformed)

    57 MB
  • dataset_summary.txt

    Comprehensive dataset summary and validation

  • dea_results.csv

    Complete differential expression analysis results for all genes

    4.1 MB
  • top_20_biomarkers.csv

    Top 20 candidate biomarker genes for ML modeling

  • best_model.joblib

    Trained Random Forest classifier (optimized via grid search)

  • X_train.csv

    Training features (20 biomarker genes)

  • y_train.csv

    Training labels (CR=1, PD=0)

How this research was produced

K-Dense Web planned and ran this cancer biology investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.

Share:
Cancer Biology

Doxorubicin Mechanism of Action

Elucidate doxorubicin's mechanism in MCF7 cells through differential expression, pathway enrichment, and drug connectivity.

Cancer Biology

Renal Cancer Drug Response Stratification

Stratify 97 ccRCC patients by predicted response to standard therapies and identify drug repurposing candidates.

Cancer Biology

PDAC KRAS Mutation Survival Analysis

Analyze KRAS variant distribution and survival outcomes in 2,336 pancreatic adenocarcinoma patients from MSK cBioPortal data.

Run this kind of analysis on your own question

Try K-Dense Web free and see how an AI co-scientist accelerates your research.