---
title: "Introducing Citation Telephone: Trace a Cited Claim Back to Its Source"
description: "Citation Telephone follows a cited claim back through the papers it cites, reads each one, and shows where the claim came from and what changed on the way."
updatedAt: "2026-10-05"
author: "Timothy Kassis, PhD"
authorTwitter: "timothykassis"
authorLinkedIn: "https://www.linkedin.com/in/timothykassis/"
tags: ["Product", "AI", "Research", "Scientific Integrity"]
canonical: "https://www.k-dense.ai/blog/introducing-citation-telephone"
---
In January 1980, the New England Journal of Medicine printed a five-sentence letter from Jane Porter and Hershel Jick. Among 11,882 hospitalized patients who had received at least one narcotic, they found only four well-documented cases of addiction in people with no history of it. The letter said nothing about patients taking opioids at home for months, yet it went on to be cited hundreds of times as evidence that addiction to prescribed opioids is rare. When [Leung and colleagues traced its citations in 2017](https://doi.org/10.1056/NEJMc1700150), they found 608 citing articles; 439 of them (72%) cited the letter as evidence that addiction was rare, and most never mentioned that its patients were in hospital.

Few of those authors set out to mislead anyone. Each cited the paper in front of them a little more confidently than the last, and the claim drifted one hop at a time, like a message in the children's game of telephone. The pattern is common. A [meta-analysis of 28 studies of quotation accuracy](https://doi.org/10.7717/peerj.1364) in medical journals estimated that about a quarter of citations (25.4%) don't support the statement they're attached to, and about one in eight are major errors. A [study of one citation network in biomedicine](https://doi.org/10.1136/bmj.b2680) showed how a belief can gain authority through citation alone, including "the conversion of hypothesis into fact".

Checking a citation means opening the cited paper and finding the passage it was cited for. Checking where a claim really came from means doing that again for every paper it passes through. So today we're releasing **[Citation Telephone](https://www.k-dense.ai/tools/citation-telephone)**, a free tool that does the walking for you. Pick a cited sentence, and it follows the claim back through the papers it cites, reads each one, and shows you where the claim came from and what changed on the way.

## One sentence, followed back

Here is a sentence from a 2024 paper in the Journal of Medicinal Chemistry: "More importantly, nuclear PD-L1 can promote cell proliferation in an immune-dependent manner without interacting with PD-1." It cites a 2024 review in Advanced Science. Citation Telephone opened the review's full text and found the passage it was cited for. The review says nuclear PD-L1 "promotes immune-independent cell proliferation without binding to PD-1", the opposite of what the citing sentence says. The review in turn cites a [2022 study in Cancer Letters](https://doi.org/10.1016/j.canlet.2021.12.017), which reports the finding from its own experiments: nuclear PD-L1 "promotes cell cycle progression in an immune-independent manner in BRAF V600E-mutated CRC."

![Citation Telephone trail for the nuclear PD-L1 claim: the 2024 citing sentence says immune-dependent and is marked Changed against the 2024 review it cites, which says immune-independent and passes the claim on to a 2022 Cancer Letters study, the origin; the review is selected, with its passage highlighted](https://www.k-dense.ai/blog/introducing-citation-telephone/pd-l1-trail.png)

The trail shows what a quick check would miss. The first hop is marked **Changed**: one word ("dependent" for "independent") reverses the claim. The second hop carries no label, because the model's readings were split (partly matches 39%, matches 34%, goes further 25%) and none reached the bar for showing one. That split is itself a hint: the review states as a general property of nuclear PD-L1 what the original study showed in one kind of colorectal cancer. At the origin, the tool also compares your starting sentence with what the origin says, and that comparison comes back Changed (81%). None of this needs you to trust the tool's reading: each step opens the passage it rests on, highlighted inside its paragraph, and the readings behind every label are one click away.

## What each hop tells you

Every line on the trail compares the citing sentence with the passage it found in the cited paper:

- **Matches**: the paper says what it's cited for, allowing different wording and rounding.
- **Partly matches**: the paper backs part of what it's cited for, but not all of it.
- **Goes further**: the citing sentence claims more than the paper shows. It might be more certain, causal where the paper reports an association, or general where the paper studied one group.
- **More cautious**: the citing sentence claims less than the paper.
- **Changed**: the paper addresses the same question but says something else, such as a different number, population, cell type or direction.
- **Not there**: no passage of the paper states what it's cited for.

Every paper also gets a result that decides whether the trail goes on. An **origin** reports the finding as its own result, from its own experiments, data or pooled analysis, and the trail ends there. A paper that **passes it on** states the claim as background and cites earlier work for it, and the trace follows those references, choosing only the ones cited for this claim when a sentence cites several. A paper can also state the claim with **no source given**, or turn out to **not say it** at all. If a paper's full text isn't open access, it is marked **abstract only**, and that never counts as "doesn't say it", because a claim missing from an abstract may still be in the paper.

## Two ways to start

The usual way in is through a paper. Paste the DOI, PubMed ID or PMC ID of any open-access paper in Europe PMC, and Citation Telephone lists every sentence in it that cites something, grouped by section, with a filter to find the one you mean. Choose a sentence and the references it cites appear underneath; pick up to three to trace.

![The From a paper panel with a 2026 Genes & Development review loaded, its cited sentences listed by section, the sentence about AlphaFold and protein-ligand complexes selected, and its two references, Jumper et al. 2021 and Abramson et al. 2024, checked as sources to trace](https://www.k-dense.ai/blog/introducing-citation-telephone/picker.png)

The other way in is a claim from anywhere: a grant review, a news story, a slide. Paste the sentence and the DOI or PubMed ID of the paper it cites, and the trace starts from that paper. Either way, the page fills in as each paper is read, and when the trail ends you can pick any reference that wasn't followed and trace that one too.

## More trails

Machine learning results travel fast, and so do their citations. A 2026 review in Genes & Development states that "AlphaFold and related models now predict protein and protein-ligand complexes at near-experimental accuracy," citing both the [2021 AlphaFold 2 paper](https://doi.org/10.1038/s41586-021-03819-2) and the [2024 AlphaFold 3 paper](https://doi.org/10.1038/s41586-024-07487-w). Citation Telephone reads both in full and marks each as an origin that **partly matches**. The AlphaFold 2 passage reports "accuracy competitive with experimental structures in a majority of cases" at CASP14, a result about single protein chains, not complexes or ligands. The AlphaFold 3 passage does cover protein-ligand complexes, reporting that the model "greatly outperforms classical docking tools such as Vina" without using structural inputs, which is a strong result but a different one. Each paper backs part of the sentence, and neither passage says that complexes are predicted at near-experimental accuracy.

![The AlphaFold trail: the 2026 claim about protein and protein-ligand complexes at near-experimental accuracy, with two origins, the 2021 AlphaFold 2 paper and the 2024 AlphaFold 3 paper, each marked Partly matches, and the AlphaFold 2 abstract sentence on accuracy competitive with experimental structures highlighted](https://www.k-dense.ai/blog/introducing-citation-telephone/alphafold-trail.png)

Sometimes the cited paper says the opposite. A 2024 paper writes that "it is well-known that confidence intervals provide evidence of post hoc statistical power," citing a [2019 paper in General Psychiatry](https://doi.org/10.1136/gpsych-2019-100069). That paper's own simulations show that post hoc power analyses "do not indicate true power for detecting statistical significance." The trail ends at an origin, and the hop is marked **Changed**.

![The post hoc power result: the 2019 paper is the origin, the hop is marked Changed, and the highlighted passage reports simulation results showing that post hoc power analyses do not indicate true power](https://www.k-dense.ai/blog/introducing-citation-telephone/post-hoc-power.png)

Most trails are less dramatic, and that is useful too. A 2024 paper says sleep is "a crucial period for brain memory consolidation," citing a 2022 systematic review of sleep loss and physical performance. The review mentions memory in one sentence that cites five papers for five different roles of sleep, and the trace follows only the one cited for memory: a [2019 meta-analysis](https://doi.org/10.1016/j.smrv.2019.05.006) that found a small to medium benefit of sleep on prospective memory. The trail runs two hops to an origin that reports its own pooled result, and no hop goes further than its source or changes it; the first is only a partial match, because the citing sentence says far more about sleep than the review does.

## How it works

A trace repeats the same steps at every paper, and only one of them needs a model. Finding papers, reading their structure and following references are plain code.

1. **Read the citations.** Open-access papers are read from their [Europe PMC](https://europepmc.org/) full text, in the JATS XML format publishers deposit. In that format every in-text citation is linked to its entry in the reference list, so the tool knows exactly which reference each sentence cites, whether it is written as [12], as a superscript, as "(Smith et al., 2010)" or as a range like [3–6]. For the rare paper with no citation markup at all, citations are matched in the text by number, or by first author and year.
2. **Find the cited paper.** Each reference is matched to its Europe PMC record by PubMed ID, PMC ID or DOI. References without identifiers are matched by title, or through Crossref, and a match only counts when nearly every word of the title agrees and so does the year.
3. **Read it.** A paper in the open-access subset is read in full, including tables and figure legends. For any other paper, only the abstract can be read, and the tool says so.
4. **Find the passage.** The model reads the paper paragraph by paragraph, then sentence by sentence, and picks the passage closest to what the citing sentence cites it for. It is told how the sentence cites this paper ("[14]") so that, in a sentence citing several papers for different points, it looks only for this paper's part. It picks two candidates: the closest statement, and the paper's own result on the claim.
5. **Judge it.** For each candidate, the model gives a probability that it states what the paper is cited for, and compares it with the citing sentence. Separately, with the claim out of view, it decides whether the passage is the paper's own result or a description of earlier work. If the sentence cites other papers, it also scores which of them are cited for this claim.
6. **Follow or stop.** Fixed rules turn those readings into a result. Origins end a branch. Papers that pass the claim on are followed into the references cited for it, up to three per sentence, six hops and 24 papers per trace. Further down the trail, the tool also compares the claim you started with against each passage, so drift that builds up over several hops shows even when each single hop looks small.

The model is [Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev), the System One decision model from TypeSafe that also powers [Rigor Scan](https://www.k-dense.ai/blog/introducing-rigor-scan) and [Claimscape](https://www.k-dense.ai/blog/introducing-claimscape). Jev doesn't write prose. It answers typed questions with probabilities, and its fit for this job is the same as for the other two tools. When it picks a passage, the only options it is offered are the paper's own paragraph and sentence IDs, so the evidence on screen is always text the paper contains. When it judges a passage, the result is a set of probabilities, so the rule for "origin" or "goes further" can be written down and applied the same way to every paper. And TypeSafe quotes response times of 70 to 500 milliseconds, which is what lets a multi-hop trail fill in while you watch. Every question, threshold and rule is published word for word under [How a trace is made](https://www.k-dense.ai/tools/citation-telephone#method), and you can copy the whole method as JSON.

## How we tested it

Each hop rests on one judgment: given a sentence and a paper it cites, does the paper state what the sentence cites it for, and if so, where? We measured that judgment directly, on real citations, and we report the results with their uncertainty, how the labels were made, and what the tests can't tell you.

### The test sets

We built three sets of citation pairs from open-access papers in Europe PMC. For each citing paper, we picked sentences at random from its introduction or discussion (roughly 60 to 450 characters long, citing a single reference in set 1 and one to three in sets 2 and 3), chose one of the references each cites, and looked it up in Europe PMC. Each citing paper contributed up to two pairs whose cited paper has open-access full text and, in sets 2 and 3, one whose cited paper is closed, so only its abstract can be read. Each also contributed a control: the same sentence paired with another open-access paper from its reference list that isn't cited anywhere in that paragraph. Controls are deliberately hard, because a paper from the same reference list is usually on the same topic.

| Set | Citing papers | Pairs | Topics | Labels | Role |
|---|---|---|---|---|---|
| 1 | 30 | 85 | 15 | One annotator, not blind | Development |
| 2 | 26 | 90 | 26 | Blind reviewers | Tuning v0.2 |
| 3 | 23 | 82 | 23 | Blind reviewers | Measurement only |

The 257 pairs come from 79 citing papers on 44 topics, from cancer immunotherapy and the gut microbiome to dental caries, coral reefs and protein design. No citing paper appears in more than one set, so set 3 shares no sentences with the sets used to build and tune the tool.

All labels were made by AI, not by human experts, and you should weigh the results with that in mind. Set 1 was labeled by a single annotator, an AI model working with us, after it had seen an earlier version's output, so it is not blind. Sets 2 and 3 were labeled by separate AI reviewers (instances of Claude, a different model family from Jev) that saw only the citing sentence, how it cites the paper, and the cited paper's text, never the tool's output. For each pair they gave a verdict (supports, partly supports, goes further, contradicts, not stated, or cannot tell from an abstract), quoted the best passage, and said whether that passage is the paper's own result and whether it cites other work. Five pairs were dropped because a reviewer's run was interrupted, three of them on papers about pathogens and toxins, and set 3 avoided those topics.

### What we measured

- **Found rate:** among pairs where the cited paper's full text was read and the reviewer judged that it addresses what it is cited for (supports, partly supports, goes further or contradicts), the share where the tool also found a passage that addresses it.
- **False-find rate:** among full-text pairs where the reviewer judged that the paper does not state what it is cited for, the share where the tool reported a passage anyway.
- **Closed papers:** how often a paper read only from its abstract was wrongly reported as not saying something.
- **Comparison and result type:** how often the label on a hop (matches, partly matches, goes further, changed) and the result (own result, cites earlier work, no source given) agree with the reviewer.

### Results

![Accuracy of method v0.2 by test set, with Wilson 95% confidence intervals: the supporting passage was found in 41 of 44 pairs in set 1, 36 of 42 in set 2, 38 of 39 in set 3 and 115 of 125 overall; support was reported anyway in 2 of 27, 3 of 24, 0 of 21 and 5 of 72](https://www.k-dense.ai/blog/introducing-citation-telephone/accuracy.png)

| Set | Supporting passage found | Support reported anyway |
|---|---|---|
| 1 | 41 of 44 (93%; 95% CI 82 to 98%) | 2 of 27 (7%; 2 to 23%) |
| 2 | 36 of 42 (86%; 72 to 93%) | 3 of 24 (13%; 4 to 31%) |
| 3, held out | 38 of 39 (97%; 87 to 99.5%) | 0 of 21 (0%; 0 to 15%) |
| All three | 115 of 125 (92%; 86 to 96%) | 5 of 72 (7%; 3 to 15%) |

Jev's answers vary slightly between runs, so we ran the whole evaluation three times. Across runs the found counts moved by at most one pair per set (40 to 41 in set 1, 36 in every run of set 2, 38 in every run of set 3), and the false finds by at most one (2 to 3 in set 2, 0 to 1 in set 3). The table shows the first run. Set 2 shaped the final thresholds, so its numbers are optimistic. Set 3 is the honest estimate, and with 39 and 21 pairs its intervals are wide.

Closed papers behaved as designed. In all three runs, none of the 46 pairs whose cited paper could only be read from its abstract was reported as not saying something. Where the abstract states the finding as the paper's own result, the paper is still marked as an origin.

The labels on each hop are the softest part. Where the tool found support, its comparison agreed with the reviewer's in 47 to 55% of pairs, depending on the set and run. Nearly all the disagreements are the tool saying "partly matches" where the reviewer saw a full match. That is the direction we chose to err in: "goes further" and "changed" are shown only when the model gives them at least 50%, and in every run, one citation that reviewers judged sound got one of those warnings (in set 3), while none of the three problems the reviewers found was labeled a match. Three problems is far too few to estimate how reliably the tool catches them, so treat the warnings as leads, not verdicts. The tool's call on whether a passage is the paper's own result or background agreed with the reviewers in 68 to 81% of pairs.

The tests also measured the problem the tool exists for. Of the 130 real citations the blind reviewers read, 12 (9%; 95% CI 5 to 15%) cited a paper that didn't state the claim or said something else. That is close to the rate of major quotation errors in the [meta-analysis](https://doi.org/10.7717/peerj.1364) above.

### What changed, and what didn't help

The first public version (v0.1) found the supporting passage in 31 of 42 pairs of set 2 (74%; 59 to 85%) when we first ran it, before any tuning. Version 0.2 made three changes to the judging and three to the reader:

- **A lower bar for support.** Many misses had the right sentence highlighted but a support reading just under 0.5, so a passage now counts from 0.4.
- **Warnings need confidence.** "Goes further" and "changed" are shown only at 50% or more; below that the passage gets its best other reading. This cut warnings on sound citations across sets 2 and 3 from seven to one or two, depending on the run.
- **A clearer "matches".** Parts of a sentence that are about other papers, or about the citing paper's own work, no longer count against the passage.
- **The reader** now treats any link into the reference list as a citation (one publisher marks them differently, which had hidden every citation in one of the test papers), reads citations from the text in papers that don't mark them up, and keeps an author-year citation that opens a sentence with that sentence.

Three ideas didn't help, and we left them out. Giving the model the sentence before the claim, to resolve words like "these" or "such", removed one or two false finds per 45 controls but cost about five points of found rate. Asking the support question a second way and averaging the two answers added nothing beyond what the lower threshold already gave. A stricter definition of "partly matches" moved errors from "partly matches" into "goes further", which is worse for a reader.

### Errors we still see

The misses cluster in a few kinds of sentence. Partial matches sit near the threshold: a sentence that cites a study of bats' lungs for a finding about their livers is labeled partial by the reviewer and was missed by the tool in two of the three runs. Citations of a resource rather than a finding, such as "We conducted experiments on the CXR dataset of the 2021 SIIM-FISABIO-RSNA Machine Learning COVID-19 Challenge [44]", get the right passage highlighted but a low support reading. Sentences that lean on the one before, such as "Only a few of them explored…", give it little to look for. And it misses subtle swaps of a single term: one citing paper turned "fibroblastic reticular cells" into "fibroblastic reticulocytes", and the hop came back as a match.

### Robustness

Separately from accuracy, we ran the paper reader over 175 open-access papers from 133 journals, published between 2008 and 2026. None failed or stalled, and all 14,684 citation links in them resolved to a reference. We also ran 24 full traces end to end, several of them more than one hop deep, and all finished without an error.

### Limits of this evaluation

- **The labels are AI judgments.** The reviewers are a different model family from the one the tool uses, which limits shared blind spots but doesn't remove them, and no human expert checked the labels.
- **The samples are small.** The intervals above are wide, especially for set 3 and for the false-find rate.
- **The selection is narrow.** All pairs come from open-access papers in Europe PMC, mostly biomedical, and from introduction and discussion sentences citing at most three references. Methods citations, longer citation lists and other fields may behave differently.
- **One hop at a time.** The labels judge single hops. A trail several hops long multiplies the chances of an error, so check each hop on a long trail.
- **Some topics are missing.** Pairs about pathogens and toxins were dropped from labeling.

The labeled pairs and the tool's output from all three runs are [available as a JSON file](https://www.k-dense.ai/blog/introducing-citation-telephone/citation-telephone-eval-v0.2.json), so you can check any number on this page. The questions, thresholds and rules that produced them are published word for word on the tool page under [How a trace is made](https://www.k-dense.ai/tools/citation-telephone#method).

## What it can't do

Citation Telephone can only read what is openly available. Full text comes from the Europe PMC open-access subset; papers that are free to read but not openly licensed, and everything behind a paywall, are read from their abstracts. Abstracts rarely say where a claim came from, so many trails end at a closed paper. A good example is the widely repeated claim that the gut holds ten times more bacterial cells than the body has human cells. Many papers cite a 2005 Science review for it, which can only be read as an abstract, while a [2016 re-estimate](https://doi.org/10.1371/journal.pbio.1002533) found the two numbers to be of the same order.

Books, reports, guidelines, websites and papers outside Europe PMC can't be read at all. Claims made only in a figure can be missed, though tables and figure legends are read. The limits of three references per sentence, six hops and 24 papers keep a trace fast; you can continue from any paper where a trail stopped. Above all, the tool checks citations, not truth. An origin can itself be wrong, and a claim can be true and still badly cited.

## Fix it with K-Dense Web

A broken trail usually leads to one of a few next steps: find a source that supports the claim as written, rewrite the sentence to match its sources, read the papers that were closed, or check the rest of a paper's citations the same way. The bar under every trail turns your choice into a brief for [K-Dense Web](https://www.k-dense.ai/products/k-dense-web), our AI co-scientist, which can retrieve full texts, search the wider literature and draft the corrected sentence with the right citations.

![The Fix it with K-Dense Web bar under a trail, opened to show the next steps: find a source that supports the claim as written, rewrite the sentence to match its sources, check the rest of the starting paper's citations and check the claim against the wider literature](https://www.k-dense.ai/blog/introducing-citation-telephone/handoff.png)

The brief lists the claim, every paper on the trail with its identifiers and result, and the passage each result rests on. It is built from fixed templates, so nothing in it is generated, and K-Dense Web fetches the papers itself.

## Private by default

Your claim is sent with the cited papers' public text to be read, and K-Dense doesn't store or log it. The trail lives in your browser tab. A link copied from the page reruns the trace from the same sentence, so sharing a result shares where to find it, not a stored copy.

## Try it

[Citation Telephone](https://www.k-dense.ai/tools/citation-telephone) is free, with no account needed. The page opens on the nuclear PD-L1 trail, and more examples, including the AlphaFold, post hoc power and sleep trails above, are one click away. It is most useful on a sentence you're about to cite, a claim in a manuscript you're reviewing, or a number that sounds a little too clean. If it finds something interesting, or gets something wrong, we'd like to hear about it at [contact@k-dense.ai](mailto:contact@k-dense.ai).
