Skip to main content
Chemistry· 18-page report· 1 figure

Buchwald-Hartwig C-N Coupling Yield Prediction

Predict Buchwald–Hartwig C–N coupling yields and rank reaction-component importance.

What this research found

Palladium-catalyzed carbon–nitrogen coupling is one of the most-used reactions in pharmaceutical synthesis, and its yield is notoriously hard to predict in advance. Working from the 2018 high-throughput benchmark of 3,955 reactions described by 120 quantum-chemical descriptors, gradient-boosted trees predicted percent yield with an R² of 0.953 on a random hold-out but only 0.576 when three isoxazole additives were withheld from training entirely. That gap — a near-threefold jump in error — is the practical measure of what these models cost when applied to genuinely new chemistry.

  • On a random 70/30 split, the histogram gradient-boosting model predicted yield on 1,187 held-out reactions with RMSE 6.07 percentage points and R² 0.953, an error comparable to or below the experimental reproducibility of the underlying high-throughput measurements.
  • Withholding all reactions containing 3,5-dimethylisoxazole, 5-phenylisoxazole, and benzo[c]isoxazole — chosen to span the high, medium, and low ends of the additive yield range — cut performance on those 540 reactions to RMSE 17.60 and R² 0.576. The low-yielding benzo[c]isoxazole is systematically over-predicted, since the model does not anticipate its stronger catalyst inhibition.
  • Model class matters far more under extrapolation than interpolation. A ridge-regression baseline is adequate on the random split (R² 0.707) but produces physically absurd predictions on unseen additives (R² −10.30, RMSE 90.85), while tree ensembles, bounded by their training targets, degrade smoothly to R² near 0.58.
  • By summed permutation importance the aryl halide is the dominant driver at 44.9% of attributed weight, followed by the additive at 21.5%, the ligand at 16.9%, and the base at 16.8%. Per descriptor the ranking reorders: the base leads at 1.022 with the aryl halide effectively tied at 1.013, while the ligand trails at 0.161 because its signal spreads thinly across 64 largely collinear descriptors.
  • The single strongest feature is the aryl-halide C3 nuclear magnetic resonance shift, at a permutation importance of 12.15 — more than 50% above the next descriptor — and it correlates positively with yield (Pearson 0.350, Spearman 0.431). Nine of the top fifteen descriptors correlate positively with yield and six negatively.
  • Yields span the full range with mean 33.1%, median 28.8%, and a heavy mass of failed reactions near zero. The additive is the widest experimental lever, moving mean yield from 46.7% down to 9.0% across the 22 isoxazoles tested.

How it was done

The complete four-factor grid of 22 additives, 15 aryl or heteroaryl halides, 3 bases, and 4 Buchwald biaryl phosphine ligands was reconstructed from the released descriptor tables, giving 3,955 reactions once the five cells lacking a measured yield were dropped. Component identities were recovered by exact matching of descriptor sub-blocks against the named tables rather than imputed, and a 23-point integrity audit passed on counts and combination uniqueness. Each reaction was represented purely by its 120 quantum-chemical descriptors — chemical shifts, atomic charges, frontier-orbital energies, vibrational modes, and size metrics — with no categorical component identity, which is what makes the additive hold-out a true extrapolation test. Ridge regression, a random forest, and histogram gradient boosting were tuned by 5-fold cross-validation on training data only, with the final model chosen by lowest cross-validated training error, never by test performance. Drivers were attributed by permutation importance over 20 repeats, cross-checked against univariate correlations.

Data sources

  • Ahneman, Estrada, Lin, Dreher & Doyle, Science (2018) — high-throughput Buchwald–Hartwig yield dataset, 3,955 reactions with 120 DFT descriptors
  • Public doylelab/rxnpredict and rxn4chemistry/rxn_yields repositories — mirrored descriptor tables and reaction matrix

Limitations

The extrapolation estimate rests on three withheld additives; leave-one-additive-out validation across all 22 would give a distribution of errors rather than a single number. Permutation importance is also blunted by heavy collinearity — 403 descriptor pairs exceed 0.90 in absolute correlation — so it identifies informative blocks and representative descriptors rather than uniquely causal ones, and all conclusions pertain to this specific chemical space.

How this research was produced

K-Dense Web planned and ran this chemistry investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.

Share:
Chemistry

METLIN Retention-Time QSAR Prediction

Predict METLIN chromatographic retention times with QSAR models against observed values.

Chemistry

nmrshiftdb2 NMR Structure Elucidation

Elucidate molecular structures from nmrshiftdb2 NMR shifts using a multi-stage candidate funnel.

Chemistry

SAMPL6 First-Principles pKa Prediction

Predict SAMPL6 pKa values from first principles and correlate predictions against experimental benchmarks.

Run this kind of analysis on your own question

Try K-Dense Web free and see how an AI co-scientist accelerates your research.