What this research found
Cytomegalovirus (CMV) leaves a lasting imprint on the T-cell receptor repertoire, and a 2017 study showed that serostatus can be read from blood immunosequencing using public CMV-associated receptor sequences — given a discovery cohort of about 640 subjects. This re-analysis asked how much of that signal survives at 200 subjects under an evaluation designed to exclude data leakage. The public-clonotype features fell to chance (AUROC 0.51–0.53), while seven simple diversity and clonality statistics reached AUROC 0.766.
- Seven global repertoire summary statistics predicted CMV serostatus at AUROC 0.766 and AUPRC 0.736 against a 0.50 prevalence baseline, with sensitivity 0.31 at the threshold achieving 90% specificity.
- The three de novo public-clonotype models sat at chance on held-out subjects — burden score AUROC 0.525, L1 logistic regression 0.510, L2 0.532 — even though in-fold training separation was near-perfect at roughly 0.97, the signature of selection on noise.
- Only 9.4 public clonotypes per fold cleared p < 10⁻⁴ out of roughly 850,000 candidates screened. Of the 764 clonotypes ever entering a fold's feature set, just 6 recurred in all five folds; the most reproducible was CASSPGLAGVYNEQFF.
- The predictive signal is a shift in clone-size distribution consistent with CMV memory inflation: seropositive subjects had lower Simpson diversity (0.9971 vs 0.9993, p = 1 × 10⁻¹²) and higher clonality (0.086 vs 0.048, p = 5 × 10⁻¹¹), with no difference in unique clone count or CDR3 length.
- The diversity-based model is far cheaper to run. Its AUROC stayed at roughly 0.75–0.77 across a 60-fold reduction in sequencing depth down to 5,000 templates per subject, while public-clonotype discovery yielded zero significant clonotypes at depths of 25,000 templates or below.
How it was done
A balanced 200-subject subset (100 seropositive, 100 seronegative) was drawn from the Emerson 2017 discovery repertoires, matched exactly on sex and pair-for-pair on age, with sequencing depth showing no association with serostatus. Three feature families were compared: a burden score counting CMV-associated clones present, a sparse binary matrix of the top 200 associated receptor sequences, and seven whole-repertoire summaries covering Shannon entropy, Simpson diversity, clonality, CDR3 length, and clonal expansion. Evaluation used subject-level five-fold cross-validation in which every label-dependent step, including the hypergeometric enrichment test used to pick candidate clonotypes, ran inside training folds only. Each repertoire was then multinomially down-sampled to fixed depths from 5,000 to 100,000 templates and the whole pipeline re-run at each depth.
Data sources
- Emerson et al., Nature Genetics 49:659 (2017) — TCRβ discovery cohort, 785 subjects, accessed via the pre-processed DeepRC distribution
- Pooled vocabulary of 23,666,730 unique CDR3β amino-acid sequences across the 200 selected repertoires
Limitations
The 1:1 design gives a clean effect estimate but does not reflect real-world CMV seroprevalence of 60–90%, so achievable precision would differ in deployment. The pre-processed data retain amino-acid sequence and template count but not V and J gene calls, and the cohort comes from a predominantly North American donor population whose HLA distribution may not transfer, since CMV-associated public clonotypes are HLA-restricted.
How this research was produced
K-Dense Web planned and ran this immunology investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.


