Skip to main content

Evidence

Benchmarks

Claims about AI research agents are cheap. These are the measurements behind ours — including the ones that did not go our way.

On grading our own homework

Most of these benchmarks were run by us, on our own platform. That is a real conflict of interest and we do not pretend otherwise. What we can do is publish the prompts, the scoring, the reproducibility artifacts, and the results where a competitor beat us — so you can check the work rather than trust the summary. Every card below states what its result does not establish.

Platform

How K-Dense Web itself scores on external and head-to-head benchmarks.

50 verified questions

BixBench-Verified-50

K-Dense Web on Phylo's expert-cleaned subset of the BixBench biology-agent benchmark.

Result

45/50 correct — 90.0% accuracy.

What it doesn’t show

This is the cleaned subset, not raw BixBench. The cleaning was done by Phylo, not us, but a verified subset is an easier target than the original by construction.

Read the full benchmark
20 equally weighted tasks

K-Dense Web vs Claude Science

Head-to-head on interdisciplinary research tasks, scored on scientific quality and research execution.

Result

K-Dense Web took the higher mean composite score (67.8 vs 61.8) and led 14 of 20 prompts, with research execution the decisive gap: 14.7/25 vs 6.4/25.

What it doesn’t show

Claude Science earned the higher mean scientific-quality score (55.3/75 vs 53.1/75) and produced the single highest-scoring result in the comparison. We ran this benchmark on our own platform, and platform identity was visible to the evaluator.

Read the full benchmark

Agent skills

Whether our open-source agent skills measurably improve agent behaviour.

32 drugs for pKa, ~60 cloud workflows total

Rowan skill vs RDKit and experiment

pKa, logD, tautomer populations, ADME, and docking run by an agent with the Rowan skill, measured against RDKit and literature values.

Result

Rowan's pKa model reached 0.23 MAE (R² 0.986) against experiment, versus 1.15 MAE for an RDKit functional-group heuristic — error at the level of experimental reproducibility.

What it doesn’t show

RDKit ships no pKa predictor at all, so the baseline here is a hand-built heuristic rather than a fair tool-to-tool comparison. The skill was contributed by Rowan Scientific.

Read the full benchmark
250 agent runs

pyOpenMS skill for mass spectrometry

Ten real mass-spectrometry tasks solved with and without the pyopenms skill, across two model tiers.

Result

On a strong model the skill lifted task success from 96% to 100%, cut pyOpenMS API errors per run by 92%, and reduced cost per correct result by 10%.

What it doesn’t show

The gains come from documenting API changes in pyOpenMS 3.5.0 that models get wrong from memory. On a weaker model the skill was transformative only when the agent actually invoked it.

Read the full benchmark
10 skills, three model tiers, plus a four-stage screening workflow

NVIDIA BioNeMo NIM skills

Whether an agent skill improves reliability when calling NVIDIA NIM biomolecular microservices, across three model tiers.

Result

On hard single-call cases, Haiku with the skill beat Opus baseline at roughly one-fifth the cost. On a four-stage virtual-screening workflow it completed every run, where the baseline completed one in five.

What it doesn’t show

Skills did not make the underlying models better at science — ProteinMPNN recovery stayed around 40% across every arm. The gain is delivery reliability, not scientific accuracy, and skills did not beat pasted-in documentation on single calls when all the needed context was present.

Read the full benchmark

Frontier models

What the newest frontier models can and cannot do on scientific work.

240 images

Scientific image models

Four image models generating scientific figures, scored on figure quality and latency.

Result

Nano Banana 2 Lite was fastest by a wide margin — 3.8 s median versus 49.0 s — with all 60 of its generations completing successfully.

What it doesn’t show

It is not the best model in the study. GPT Image 2 leads on raw scientific figure quality, and scientific accuracy remains the hardest part of the problem for every model tested.

Read the full benchmark
35 cases

Google Omni Flash for scientific video

Omni Flash generating scientific video, scored on visual quality, text rendering, physical dynamics, and factual correctness.

Result

Visual quality is genuinely strong: on-prompt, polished clips in under a minute, useful for exploring scientific visuals under human supervision.

What it doesn’t show

Not trustworthy unattended. Text rendering, physical dynamics, and factual consistency are all unreliable, and the visual polish makes the factual errors harder to catch. The right posture is expert-vetted drafting, not autonomous explanation.

Read the full benchmark

Benchmark it yourself

The most useful benchmark is the one on your own work. Run K‑Dense on a deliverable you already know the cost and quality of, and compare.