What this research found
A widely used mass-cytometry benchmark — 265,627 human bone-marrow cells measured across 32 protein markers, 104,184 of them assigned by an expert to 14 populations — was clustered from marker expression alone and each cluster mapped back to a population by majority vote. Abundant, phenotypically distinct cell types were recovered at 93.0% accuracy with population frequencies almost exactly reproduced (Pearson r = 0.998). The clustering could not separate two overlapping CD34+ stem and progenitor compartments, which differ in essentially one marker.
- Mapped predictions matched the manual gates on 93.0% of labeled cells (micro-F1 0.930), while chance-corrected partition agreement was lower — adjusted Rand index 0.369, normalized mutual information 0.688 — as expected when 30 clusters are fit to 14 target populations.
- Recovered population frequencies tracked the manual ones with a mean absolute error of 0.41 percentage points, a maximum of 1.21, and a correlation of 0.998. Pre-B cells came within 0.001 percentage points, at 5.889% manual versus 5.890% recovered.
- Seven abundant populations reached F1 of 0.90 or better, led by monocytes at 0.987 and mature B cells at 0.974. Macro-F1 was far lower at 0.602 because four rare populations — CD34+CD38lo HSCs, pro B cells, plasma B cells, and CD34+CD38+CD123+ HSPCs — never formed the majority of any cluster and so scored zero.
- The hardest pair to distinguish was CD34+CD38+CD123− hematopoietic stem/progenitor cells versus CD34+CD38lo hematopoietic stem cells, with a mutual-confusion score of 0.996. Of 916 gated stem cells, 912 landed in progenitor-dominated clusters.
- The failure traces to a lack of contrast rather than missing information. Of 32 markers only CD38 separates the two compartments appreciably (median 3.409 versus 1.197 arcsinh units), followed distantly by HLA-DR at 1.005 and CD45 at 0.601, while eight lineage markers differ by less than 0.01 — one informative dimension among 31 near-identical ones counts for little in an equally weighted Euclidean distance.
How it was done
Thirty raw cytometry files from two healthy donors were read and reconstructed into a single table of 265,627 cells, and the 32 lineage markers were arcsinh-transformed with the conventional mass-cytometry cofactor of 5, then z-scored so no high-variance marker dominated the distance metric. Mini-batch k-means at k = 30 clustered every cell with the manual labels withheld; setting k above the 14 target populations was deliberate, giving rare or heterogeneous types a chance at their own cluster. Each cluster was then assigned to the population holding the majority of its labeled members, and agreement was scored with adjusted Rand index, normalized mutual information, and per-population precision, recall, and F1. Every population pair was ranked by mutual confusion, and the worst pair was diagnosed by comparing median expression across all 32 markers. A UMAP embedding of 20,000 labeled cells was produced for visual inspection only.
Data sources
- Levine et al., Cell 162:184 (2015) — 32-marker bone-marrow mass cytometry benchmark, 265,627 cells from two healthy donors, 104,184 manually gated into 14 populations
- HDCytoData Bioconductor resource — analysis-ready distribution of the benchmark
Limitations
One clustering family at one resolution with a single random seed was tested, so the figures characterise this configuration rather than the best achievable. Majority-vote mapping cannot by construction predict a population that never dominates a cluster, and the reference is operator-dependent manual gating — most consequentially for the CD38 threshold that separates the two populations found hardest to resolve.
How this research was produced
K-Dense Web planned and ran this immunology investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.


