What this research found
Can a molecule's connectivity be recovered from nothing but its formula and two unlabelled lists of chemical shifts? Thirty small organic molecules from the open nmrshiftdb2 spectra database were blinded to exactly that information, and an automated pipeline reasoned from degrees of unsaturation and carbon shift regions to a proposed structure plus a full peak-to-atom assignment table. All 30 constitutions were recovered exactly, with a self-consistent assignment for every one of the 405 observed peaks.
- The pipeline recovered the correct constitution for 30 of 30 blinded molecules — an exact-match fraction of 1.000 on the connectivity block of the InChIKey — and every proposal carried a self-consistent peak-to-atom assignment.
- Matching on molecular formula alone leaves real ambiguity: the same-formula candidate count ranged from 1 to 15 per unknown, with a median of 5.
- The carbon-13 functional-group fingerprint does most of the discriminative work. For C8H9NO2 it collapsed 15 same-formula isomers to the single correct candidate, 2-(4-hydroxyphenyl)acetamide, and it reached a unique answer in one step for C10H8O2 (5 to 1), C9H8O3 (6 to 1) and C8H10 (5 to 1).
- Assigning observed peaks to per-peak reference slots rather than one slot per atom is what lets equivalent protons be handled correctly — all three methyl protons of 4-methylpyridine N-oxide at 2.36 ppm map to the same carbon, and both amide N–H protons of the acetamide map to the shared nitrogen.
- Every true structure scored zero mean absolute assignment error on both nuclei by construction, while the best competing isomer always scored strictly positive, usually by several parts per million. Eight of the 30 formulae had no competitor in the pool at all.
- The selected set spans 6 to 15 heavy atoms, degrees of unsaturation from 0 to 10, 3 to 14 carbon-13 peaks and 2 to 16 proton peaks, mixing N-oxides, phenols, acids, esters, amides, anhydrides, polyaromatics, aliphatic amines, a diyne and two formal anions.
How it was done
A 58,187-record nmrshiftdb2 dump was streamed and parsed with RDKit, with predicted spectra excluded so that no shift-prediction model could leak into the inputs. Three admission criteria — elements restricted to carbon, hydrogen, nitrogen and oxygen; 6 to 15 heavy atoms; and both an atom-assigned experimental proton and carbon-13 spectrum — left a qualifying pool of 2,614 molecules, from which 30 were drawn with a fixed random seed. Each unknown was revealed only as a molecular formula, net charge and two sorted shift lists with atom labels stripped. The pipeline computed degrees of unsaturation, built a carbon-13 region fingerprint, retrieved same-formula candidates, filtered them by fingerprint distance, then scored survivors by an optimal Hungarian assignment of observed peaks to reference slots, combining per-nucleus mean absolute error with a peak-count penalty. The whole run is deterministic and byte-identical across reruns.
Data sources
- nmrshiftdb2 assigned-spectra dump, remote last-modified 2026-03-15 — 58,187 molecule records, 2,614 qualifying
- Kuhn & Schlörer, Magnetic Resonance in Chemistry 53(8):582 (2015) — nmrshiftdb2 database reference
- Steinbeck & Kuhn, Phytochemistry 65(19):2711 (2004) — NMRShiftDB origin paper
- RDKit 2026.03.3 and SciPy linear sum assignment (Kuhn 1955; Jonker & Volgenant 1987)
Limitations
This is a closed-pool retrieval-and-assignment benchmark, not de-novo elucidation: each true structure sits in the reference pool and its stored shifts are identical to the blinded input, so the correct candidate scores zero error by construction. The work uses only one-dimensional shift lists and multiplicity codes, ignores two-dimensional correlation experiments and coupling constants, relies on a coarse region-count fingerprint that would degrade near shift-region boundaries, and treats the database curators' assignments as ground truth.
How this research was produced
K-Dense Web planned and ran this chemistry investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.


