Skip to main content
Chemistry· 15-page report· 1 figure

SAMPL6 First-Principles pKa Prediction

Predict SAMPL6 pKa values from first principles and correlate predictions against experimental benchmarks.

What this research found

How accurate can a fully automated, purely physics-based pKa prediction be if it uses the cheapest quantum chemistry available? The 24 kinase-inhibitor-like fragments of the SAMPL6 blind challenge were predicted end to end using semiempirical GFN2-xTB energies, implicit-water solvation and a thermodynamic cycle, with no empirical pKa predictor and no experimental pKa values entering the prediction chain. The raw values carry a near-constant offset of about 130 pKa units; a single global linear calibration brings them to a mean absolute error of 1.74 and a root-mean-square error of 2.18 pKa units.

  • Raw first-principles predictions all landed between −136.8 and −102.4 pKa units, giving a mean absolute error of 129.7 — an almost constant offset rather than random scatter, consistent with the known over-stabilization of anionic species by low-cost semiempirical methods combined with continuum ionic solvation.
  • A single two-parameter global calibration, calibrated pKa = 0.159 × raw pKa + 25.53, reduced the error by roughly two orders of magnitude to a mean absolute error of 1.74 and a root-mean-square error of 2.18 pKa units across the 30 matched transitions.
  • The underlying chemical signal is moderate: Pearson r = 0.54 (p = 1.9 × 10⁻³, n = 30) and R² = 0.30. Because correlation is invariant under the linear map, this is the honest measure of how much information the raw computation carries.
  • The calibration slope of 0.159 means a computed change of one pKa unit corresponds to only about 0.16 experimental units, so predictions compress toward the mean — the calibrated range is 3.8 to 9.3 against an experimental range of 2.2 to 11.7, and the largest residuals cluster at the high-pKa end (SM06 pKa2, error 5.63; SM18 pKa3, error 4.39).
  • Twenty of the 30 calibrated predictions fall within 2.0 pKa units of experiment and 12 within 1.0 unit. Within every molecule the ordering of transitions is physically correct, with higher-charge deprotonations assigned lower pKa values.
  • Enumeration produced 133 protonation microstates across the 24 molecules (2 to 11 per molecule, median 5.5) and 43 adjacent charge-state transitions. All 133 geometry optimizations and Hessians converged in both gas phase and implicit water, with solvation stabilization scaling quadratically with net charge, from about 6 kcal/mol for neutrals up to 194 kcal/mol for dications.

How it was done

SAMPL6 structures and experimental values were curated from the official challenge files and cross-checked against the primary measurement paper. Ionizable sites were identified by substructure matching and local neighbour logic, excluding groups that do not ionize between pH 2 and 12, and microstates were built as the Cartesian product of per-site protonation choices, de-duplicated by canonical tautomer form and limited to net charges from −2 to +2. The lowest-energy conformer of each microstate was optimized with GFN2-xTB in both gas phase and analytical linearized Poisson–Boltzmann implicit water, followed by a Hessian with modified rigid-rotor–harmonic-oscillator thermochemistry at 298.15 K. Aqueous free energies were combined into Boltzmann charge-state partition functions and converted to macroscopic pKa values using a literature aqueous proton free energy of −270.3 kcal/mol, the only externally supplied number. Predicted transitions were matched to experimental values within each molecule by optimal assignment before a single global least-squares calibration.

Data sources

  • SAMPL6 pKa blind challenge set — 24 molecules (SM01–SM24), 31 experimental macroscopic pKa values from 2.15 to 11.74
  • Işık et al., J. Comput.-Aided Mol. Des. 32(10):1117 (2018) — UV-metric spectrophotometric titrations, three replicates per molecule
  • Işık et al., J. Comput.-Aided Mol. Des. 35(2):131 (2021) — SAMPL6 pKa challenge overview
  • Tissandier et al. (1998) and Kelly, Cramer & Truhlar (2006) — aqueous proton free energy reference
  • GFN2-xTB (xtb 6.7.1) with ALPB implicit water; RDKit, SciPy and NumPy

Limitations

The accuracy ceiling is set by the level of theory, not the workflow: GFN2-xTB has limited accuracy for the relative stability of protonation-site isomers, only one conformer per microstate was used so conformational entropy is neglected, and the continuum solvation model omits specific solute–water hydrogen bonding. One of the 31 experimental values (SM16's second pKa) could not be matched because the enumerated microstates for that molecule provide only one adjacent charge-state pair.

How this research was produced

K-Dense Web planned and ran this chemistry investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.

Share:
Chemistry

METLIN Retention-Time QSAR Prediction

Predict METLIN chromatographic retention times with QSAR models against observed values.

Chemistry

nmrshiftdb2 NMR Structure Elucidation

Elucidate molecular structures from nmrshiftdb2 NMR shifts using a multi-stage candidate funnel.

Chemistry

QM9 HOMO-LUMO Gap Prediction

Predict QM9 HOMO–LUMO gaps with ML regression and validate predicted-versus-true parity plots.

Run this kind of analysis on your own question

Try K-Dense Web free and see how an AI co-scientist accelerates your research.