Skip to main content

Evidence

Benchmarks

Claims about AI research agents are cheap. These are the measurements behind ours — including the ones that did not go our way.

On grading our own homework

Most of these benchmarks were run by us, on our own platform. That is a real conflict of interest and we do not pretend otherwise. What we can do is publish the prompts, the scoring, the reproducibility artifacts, and the results where a competitor beat us — so you can check the work rather than trust the summary. Every card below states what its result does not establish.

Platform

How K-Dense Web itself scores on external and head-to-head benchmarks.

50 verified questions

BixBench-Verified-50

K-Dense Web on Phylo's expert-cleaned subset of the BixBench biology-agent benchmark.

Result

45/50 correct — 90.0% accuracy.

What it doesn’t show

This is the cleaned subset, not raw BixBench. The cleaning was done by Phylo, not us, but a verified subset is an easier target than the original by construction.

20 equally weighted tasks

K-Dense Web vs Claude Science

Head-to-head on interdisciplinary research tasks, scored on scientific quality and research execution.

Result

K-Dense Web took the higher mean composite score (67.8 vs 61.8) and led 14 of 20 prompts, with research execution the decisive gap: 14.7/25 vs 6.4/25.

What it doesn’t show

Claude Science earned the higher mean scientific-quality score (55.3/75 vs 53.1/75) and produced the single highest-scoring result in the comparison. We ran this benchmark on our own platform, and platform identity was visible to the evaluator.

Agent skills

Whether our open-source agent skills measurably improve agent behaviour.

32 drugs for pKa, ~60 cloud workflows total

Rowan skill vs RDKit and experiment

pKa, logD, tautomer populations, ADME, and docking run by an agent with the Rowan skill, measured against RDKit and literature values.

Result

Rowan's pKa model reached 0.23 MAE (R² 0.986) against experiment, versus 1.15 MAE for an RDKit functional-group heuristic — error at the level of experimental reproducibility.

What it doesn’t show

RDKit ships no pKa predictor at all, so the baseline here is a hand-built heuristic rather than a fair tool-to-tool comparison. The skill was contributed by Rowan Scientific.

250 agent runs

pyOpenMS skill for mass spectrometry

Ten real mass-spectrometry tasks solved with and without the pyopenms skill, across two model tiers.

Result

On a strong model the skill lifted task success from 96% to 100%, cut pyOpenMS API errors per run by 92%, and reduced cost per correct result by 10%.

What it doesn’t show

The gains come from documenting API changes in pyOpenMS 3.5.0 that models get wrong from memory. On a weaker model the skill was transformative only when the agent actually invoked it.

10 skills, three model tiers, plus a four-stage screening workflow

NVIDIA BioNeMo NIM skills

Whether an agent skill improves reliability when calling NVIDIA NIM biomolecular microservices, across three model tiers.

Result

On hard single-call cases, Haiku with the skill beat Opus baseline at roughly one-fifth the cost. On a four-stage virtual-screening workflow it completed every run, where the baseline completed one in five.

What it doesn’t show

Skills did not make the underlying models better at science — ProteinMPNN recovery stayed around 40% across every arm. The gain is delivery reliability, not scientific accuracy, and skills did not beat pasted-in documentation on single calls when all the needed context was present.

Frontier models

What the newest frontier models can and cannot do on scientific work.

100 questions, three models, one run each

CompBioBench v1 on K-Dense BYOK

DeepSeek V4 Pro, GPT-5.6 SOL, and Gemini 3.8 Flash on Genentech's CompBioBench v1, run through K-Dense BYOK with the same prompt, sandbox, skills, and tools.

Result

Gemini 3.8 Flash scored 94% and GPT-5.6 SOL 92%, a statistical tie. DeepSeek V4 Pro scored 81%, with up to 3 points of that gap attributable to the harness rather than the model.

What it doesn’t show

One run per model. The harness changed between runs, five r3 answers were operator-selected re-runs, and the leaderboard returns an aggregate score only, so we do not know which questions any run got wrong. Costs are not comparable across models with differing cache behavior.

178 tasks, 9 models, 1,602 runs

K-Bench 01

Nine frontier models running real scientific requests from K-Dense Web users end to end, bare in a stock harness, scored blind by three LLM judges against an 8-dimension rubric.

Result

Only one model averages above the 8/10 bar where a scientist would accept the work, and honesty-related failures such as overclaiming touch 40% of all runs. Full methodology and results are published as a pre-print on arXiv.

What it doesn’t show

The top-scoring model is also one of the judges and rates its own runs +0.83 points above where its peers put them; by the only neutral judge the winner changes. The tasks stay internal to avoid contamination, so the result is not externally reproducible, and there is no expert human baseline.

240 images

Scientific image models

Four image models generating scientific figures, scored on figure quality and latency.

Result

Nano Banana 2 Lite was fastest by a wide margin — 3.8 s median versus 49.0 s — with all 60 of its generations completing successfully.

What it doesn’t show

It is not the best model in the study. GPT Image 2 leads on raw scientific figure quality, and scientific accuracy remains the hardest part of the problem for every model tested.

35 cases

Google Omni Flash for scientific video

Omni Flash generating scientific video, scored on visual quality, text rendering, physical dynamics, and factual correctness.

Result

Visual quality is genuinely strong: on-prompt, polished clips in under a minute, useful for exploring scientific visuals under human supervision.

What it doesn’t show

Not trustworthy unattended. Text rendering, physical dynamics, and factual consistency are all unreliable, and the visual polish makes the factual errors harder to catch. The right posture is expert-vetted drafting, not autonomous explanation.

Benchmark it yourself

The most useful benchmark is the one on your own work. Run K‑Dense on a deliverable you already know the cost and quality of, and compare.