What this research found
Adults given a Tdap booster mount antibody recalls against pertussis toxin that range from a slight decline to more than a 60-fold rise. Using the third public CMI-PB prediction challenge, day-0 measurements alone — antibody titres, plasma cytokines, blood cell frequencies, and demographics, with gene expression deliberately excluded — were used to predict whether a subject's day-14 fold-rise would land above or below the cohort median. A random forest reached AUROC 0.841 against a 49.5% high-responder prevalence, and pre-existing anti-pertussis-toxin IgG was the dominant predictor, acting inversely.
- The random forest reached AUROC 0.841 ± 0.081 and AUPRC 0.838 ± 0.087 across 50 outer folds, about 0.34 above the 0.495 no-skill precision-recall line. An elastic-net logistic regression reached AUROC 0.771 ± 0.074 and AUPRC 0.741 ± 0.091.
- Baseline anti-pertussis-toxin IgG was the strongest predictor in both models and pointed the other way: the largest logistic coefficient at −3.13, and the top forest importance at 0.105, more than double the next feature. Subjects already carrying high titres have the least room to boost — a ceiling effect.
- Beyond antibody titres the signal was cellular. Baseline B-cell frequency was a top-four feature in both models with a positive coefficient of +1.66, alongside central-memory CD4 (+1.39) and effector-memory CD8 (+1.29) T-cell frequencies. Among plasma proteins, the Th2 cytokine interleukin-4 carried a negative coefficient of −1.20.
- Whole-cell-primed adults boosted more strongly than acellular-primed adults on unadjusted comparison — median log2 fold-rise 2.23 versus 1.69, and 53.6% versus 45.5% high responders — yet the priming indicator carried almost no independent weight once molecular and cellular features were included, suggesting priming acts through the measured immune state rather than as a separate axis.
- The booster reliably recalled anti-pertussis-toxin IgG but with wide variation: the median subject rose about 4.2-fold, and log2 fold-rise spanned roughly −1.5 to +6. Labelling was insensitive to the normalization choice, with raw and normalized ratios giving identical labels and a Spearman correlation of 1.0.
How it was done
Longitudinal tables covering 172 adults were assembled — 118 across three training years plus 54 in a held-out challenge year — and each subject's day-0 and day-14 specimens were matched to the planned time points. The endpoint was the day-14 to day-0 ratio of normalized anti-pertussis-toxin IgG, split at the cohort median of 4.2387 into 55 high and 56 low responders; the 7 training subjects missing either time point were dropped rather than imputed. Because assay panels differed by year, features were harmonized to the cross-year intersection, giving 84 baseline predictors: 3 demographic, 31 antibody titres, 30 plasma cytokines, and 20 blood cell-population frequencies. An elastic-net logistic regression and a random forest were compared under nested cross-validation — 10 repeats of stratified five-fold outer splits, with imputation, scaling, and an inner five-fold grid search confined to each outer training split — then refit on all 111 labeled subjects to score the withheld challenge cohort.
Data sources
- CMI-PB 3rd public prediction challenge, 2024 build — 172 adults, sampled at days −30, −14 and 0 pre-boost and days 1, 3, 7 and 14 post-boost
- Luminex antibody titres against pertussis toxin, pertactin, filamentous haemagglutinin, fimbriae 2/3, and diphtheria and tetanus control antigens
- Olink plasma cytokine panel and PBMC cell-subset frequencies
- Shinde et al., PLOS Computational Biology 21:e1012927 (2025) — challenge design and harmonization guidance
Limitations
The labeled training set is small at 111 predominantly young adults from a few academic cohorts, and harmonizing to the cross-year intersection discarded features measured in only some years, so residual batch effects may persist. Cytokine and cell-frequency panels were missing for roughly 37% of subjects, meaning median imputation probably understates their contribution, and the median-dichotomized endpoint discards magnitude information.
How this research was produced
K-Dense Web planned and ran this immunology investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.


