Skip to main content
Nanomedicine· 7 figures

Nanoparticle Protein Corona Prediction

Build ML models to predict protein corona composition on gold nanoparticles using physicochemical properties.

What this research found

Can a nanoparticle's physicochemical properties alone predict which blood proteins will coat it — the so-called protein corona? Models were trained on 144 observations covering 25 serum proteins across six formulations, using a dataset assembled synthetically from published corona compositions because raw proteomics data could not be obtained. The answer came back negative: tested on held-out nanoparticle types, random forest scored an R² of -0.158 and gradient boosting -0.679, both worse than a dummy model that simply predicts the mean (-0.013).

  • Prediction across formulations failed outright. Under leave-one-group-out validation over the six nanoparticle types, random forest reached R² = -0.158 and gradient boosting -0.679, against -0.0134 for a mean-predicting baseline, meaning both models were worse than useless for unseen particle types.
  • Protein properties outrank particle properties in the random forest's feature importances: protein length (0.235), protein molecular weight (0.170), ligand molecular weight (0.129) and protein hydropathy (0.115) all beat zeta potential (0.079) and every surface-chemistry indicator.
  • No individual physicochemical variable tracked protein abundance. The strongest correlation was ligand hydrophobicity at r = 0.0906 (p = 0.28), and a Kruskal-Wallis test across five surface chemistries found no effect on abundance (H = 4.7458, p = 0.31).
  • The intuitive charge-matching mechanism did not show up: nanoparticle zeta potential and protein isoelectric point correlated at r = 0.0043 (p = 0.96) across all pairs, and r = -0.0087 among proteins of above-median abundance.
  • Adding 3D structure helped only marginally. For the nine most abundant proteins whose predicted structures were retrieved, relative exposed hydrophobic surface area correlated with abundance at r = 0.267 (p = 0.49) — better than the sequence-based hydropathy score at r = -0.1015, but neither is significant.
  • The work still delivered a tool: a calculator that takes size, zeta potential, surface chemistry and core material and returns a predicted corona fingerprint plus a dysopsonin-to-opsonin stealth ratio. Its own output flags the weakness — one test case rated a cationic amine-coated particle as low immune risk, which the report notes contradicts established literature.

How it was done

Because raw corona proteomics could not be pulled from public repositories, a 144-row dataset was constructed synthetically from published compositions, covering 25 serum proteins across six formulations with gold, PLGA, polystyrene and silica cores, sizes from 20 to 100 nm, and zeta potentials from -42.3 to +32.7 mV. Thirteen predictors combined protein descriptors — length, molecular weight, isoelectric point and hydropathy — with particle size and charge plus RDKit-computed descriptors for each surface ligand. Dummy, random forest and gradient boosting regressors were then compared under leave-one-group-out cross-validation grouped by nanoparticle type, which deliberately tests generalisation to formulations the model has never seen. AlphaFold structures were retrieved for the nine most abundant proteins and solvent-accessible surface area computed with the Shrake-Rupley algorithm to test a structural hydrophobicity hypothesis, and the fitted model was wrapped in a command-line design calculator with rule-based safety warnings.

Data sources

  • Synthetic dataset of 144 observations — 25 serum proteins across 6 nanoparticle formulations, built from literature-reported corona compositions
  • Tenzer et al., ACS Nano 7(4):3253-3263 (2013) — size-dependent corona composition
  • Walkey et al., ACS Nano 8(3):2439-2455 (2014) — corona fingerprinting and cellular interaction
  • Monopoli et al., Journal of the American Chemical Society 133(8):2525-2534 (2011)
  • Ke et al., ACS Nano 11(11):11773-11784 (2017)
  • AlphaFold Protein Structure Database — 9 predicted structures retrieved and analysed

Limitations

The training data is synthetic rather than experimental, generated from literature-derived abundance patterns with simplified size, charge and surface-chemistry rules, so the negative result reflects that construction as much as real corona biology. The sample is also small — 144 observations over six formulations, and nine structures for the structural test — and the models treat the corona as a static snapshot, ignoring protein exchange kinetics and cooperative binding.

Figures from this analysis

How this research was produced

K-Dense Web planned and ran this nanomedicine investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.

Share:
Drug Discovery

Natural Products Chemical Space

Cluster 715K+ compounds from the COCONUT database to identify bacterial-derived drug candidates using ML dimensionality reduction.

Genomics

Longevity Gene Analysis

Map model organism longevity genes to human orthologs and evaluate their validation through GWAS studies.

Astronomy

JWST Exoplanet Target Prioritization

Rank 124 habitable zone exoplanets for JWST atmospheric characterization using TSM/ESM metrics and observability windows.

Run this kind of analysis on your own question

Try K-Dense Web free and see how an AI co-scientist accelerates your research.