Skip to main content

Introducing Citation Telephone: Trace a Cited Claim Back to Its Source

Citation Telephone follows a cited claim back through the papers it cites, reads each one, and shows where the claim came from and what changed on the way.

12 min read
Share:

In January 1980, the New England Journal of Medicine printed a five-sentence letter from Jane Porter and Hershel Jick. Among 11,882 hospitalized patients who had received at least one narcotic, they found only four well-documented cases of addiction in people with no history of it. The letter said nothing about patients taking opioids at home for months, yet it went on to be cited hundreds of times as evidence that addiction to prescribed opioids is rare. When Leung and colleagues traced its citations in 2017, they found 608 citing articles; 439 of them (72%) cited the letter as evidence that addiction was rare, and most never mentioned that its patients were in hospital.

Few of those authors set out to mislead anyone. Each cited the paper in front of them a little more confidently than the last, and the claim drifted one hop at a time, like a message in the children's game of telephone. The pattern is common. A meta-analysis of 28 studies of quotation accuracy in medical journals estimated that about a quarter of citations (25.4%) don't support the statement they're attached to, and about one in eight are major errors. A study of one citation network in biomedicine showed how a belief can gain authority through citation alone, including "the conversion of hypothesis into fact".

Checking a citation means opening the cited paper and finding the passage it was cited for. Checking where a claim really came from means doing that again for every paper it passes through. So today we're releasing Citation Telephone, a free tool that does the walking for you. Pick a cited sentence, and it follows the claim back through the papers it cites, reads each one, and shows you where the claim came from and what changed on the way.

One sentence, followed back

Here is a sentence from a 2024 paper in the Journal of Medicinal Chemistry: "More importantly, nuclear PD-L1 can promote cell proliferation in an immune-dependent manner without interacting with PD-1." It cites a 2024 review in Advanced Science. Citation Telephone opened the review's full text and found the passage it was cited for. The review says nuclear PD-L1 "promotes immune-independent cell proliferation without binding to PD-1", the opposite of what the citing sentence says. The review in turn cites a 2022 study in Cancer Letters, which reports the finding from its own experiments: nuclear PD-L1 "promotes cell cycle progression in an immune-independent manner in BRAF V600E-mutated CRC."

Citation Telephone trail for the nuclear PD-L1 claim: the 2024 citing sentence says immune-dependent and is marked Changed against the 2024 review it cites, which says immune-independent and passes the claim on to a 2022 Cancer Letters study, the origin, marked Partly matches

The trail shows two things a quick check would miss. The first hop is marked Changed: one word ("dependent" for "independent") reverses the claim. The second hop is marked Partly matches, because the review states as a general property of nuclear PD-L1 what the original study showed in one kind of colorectal cancer. At the origin, the tool also compares your starting sentence with what the origin says, and that comparison comes back Changed as well. None of this needs you to trust the tool's reading: each step opens the passage it rests on, highlighted inside its paragraph.

What each hop tells you

Every line on the trail compares the citing sentence with the passage it found in the cited paper:

  • Matches: the paper says what it's cited for, allowing different wording and rounding.
  • Partly matches: the paper backs part of what it's cited for, but not all of it.
  • Goes further: the citing sentence claims more than the paper shows. It might be more certain, causal where the paper reports an association, or general where the paper studied one group.
  • More cautious: the citing sentence claims less than the paper.
  • Changed: the paper addresses the same question but says something else, such as a different number, population, cell type or direction.
  • Not there: no passage of the paper states what it's cited for.

Every paper also gets a result that decides whether the trail goes on. An origin reports the finding as its own result, from its own experiments, data or pooled analysis, and the trail ends there. A paper that passes it on states the claim as background and cites earlier work for it, and the trace follows those references, choosing only the ones cited for this claim when a sentence cites several. A paper can also state the claim with no source given, or turn out to not say it at all. If a paper's full text isn't open access, it is marked abstract only, and that never counts as "doesn't say it", because a claim missing from an abstract may still be in the paper.

Two ways to start

The usual way in is through a paper. Paste the DOI, PubMed ID or PMC ID of any open-access paper in Europe PMC, and Citation Telephone lists every sentence in it that cites something, grouped by section, with a filter to find the one you mean. Choose a sentence and the references it cites appear underneath; pick up to three to trace.

The From a paper panel with a 2026 Genes & Development review loaded, its cited sentences listed by section, the sentence about AlphaFold and protein-ligand complexes selected, and its two references, Jumper et al. 2021 and Abramson et al. 2024, checked as sources to trace

The other way in is a claim from anywhere: a grant review, a news story, a slide. Paste the sentence and the DOI or PubMed ID of the paper it cites, and the trace starts from that paper. Either way, the page fills in as each paper is read, and when the trail ends you can pick any reference that wasn't followed and trace that one too.

More trails

Machine learning results travel fast, and so do their citations. A 2026 review in Genes & Development states that "AlphaFold and related models now predict protein and protein-ligand complexes at near-experimental accuracy," citing both the 2021 AlphaFold 2 paper and the 2024 AlphaFold 3 paper. Citation Telephone reads both in full and marks each as an origin that partly matches. The AlphaFold 2 passage reports "accuracy competitive with experimental structures in a majority of cases" at CASP14, a result about single protein chains, not complexes or ligands. The AlphaFold 3 passage does cover protein-ligand complexes, reporting that the model "greatly outperforms classical docking tools such as Vina" without using structural inputs, which is a strong result but a different one. Each paper backs part of the sentence, and neither passage says that complexes are predicted at near-experimental accuracy.

The AlphaFold trail: the 2026 claim about protein and protein-ligand complexes at near-experimental accuracy, with two origins, the 2021 AlphaFold 2 paper and the 2024 AlphaFold 3 paper, each marked Partly matches, and the AlphaFold 2 abstract sentence on accuracy competitive with experimental structures highlighted

Sometimes the cited paper says the opposite. A 2024 paper writes that "it is well-known that confidence intervals provide evidence of post hoc statistical power," citing a 2019 paper in General Psychiatry. That paper's own simulations show that post hoc power analyses "do not indicate true power for detecting statistical significance." The trail ends at an origin, and the hop is marked Changed.

The post hoc power result: the 2019 paper is the origin, the hop is marked Changed, and the highlighted passage reports simulation results showing that post hoc power analyses do not indicate true power

Most trails are less dramatic, and that is useful too. A 2024 paper says sleep is "a crucial period for brain memory consolidation," citing a 2022 systematic review of sleep loss and physical performance. The review mentions memory in one sentence that cites five papers for five different roles of sleep, and the trace follows only the one cited for memory: a 2019 meta-analysis that found a small to medium benefit of sleep on prospective memory. The trail runs two hops to an origin that reports its own pooled result, and no hop goes further than its source or changes it; the first is only a partial match, because the citing sentence says far more about sleep than the review does.

How it works

A trace repeats the same steps at every paper, and only one of them needs a model. Finding papers, reading their structure and following references are plain code.

  1. Read the citations. Open-access papers are read from their Europe PMC full text, in the JATS XML format publishers deposit. In that format every in-text citation is linked to its entry in the reference list, so the tool knows exactly which reference each sentence cites, whether it is written as [12], as a superscript, as "(Smith et al., 2010)" or as a range like [3–6]. In a test on 175 papers from 133 journals, all 14,684 of their citation links resolved to a reference. For the rare paper with no citation markup at all, citations are matched in the text by number, or by first author and year.
  2. Find the cited paper. Each reference is matched to its Europe PMC record by PubMed ID, PMC ID or DOI. References without identifiers are matched by title, or through Crossref, and a match only counts when nearly every word of the title agrees and so does the year.
  3. Read it. A paper in the open-access subset is read in full, including tables and figure legends. For any other paper, only the abstract can be read, and the tool says so.
  4. Find the passage. The model reads the paper paragraph by paragraph, then sentence by sentence, and picks the passage closest to what the citing sentence cites it for. It is told how the sentence cites this paper ("[14]") so that, in a sentence citing several papers for different points, it looks only for this paper's part. It picks two candidates: the closest statement, and the paper's own result on the claim.
  5. Judge it. For each candidate, the model gives a probability that it states what the paper is cited for, compares it with the citing sentence, and decides whether it is the paper's own result or a description of earlier work. If the sentence cites other papers, it also scores which of them are cited for this claim.
  6. Follow or stop. Fixed rules turn those readings into a result. Origins end a branch. Papers that pass the claim on are followed into the references cited for it, up to three per sentence, six hops and 24 papers per trace. Further down the trail, the tool also compares the claim you started with against each passage, so drift that builds up over several hops shows even when each single hop looks small.

The model is Jev, the System One decision model from TypeSafe that also powers Rigor Scan and Claimscape. Jev doesn't write prose. It answers typed questions with probabilities, and its fit for this job is the same as for the other two tools. When it picks a passage, the only options it is offered are the paper's own paragraph and sentence IDs, so the evidence on screen is always text the paper contains. When it judges a passage, the result is a set of probabilities, so the rule for "origin" or "goes further" can be written down and applied the same way to every paper. And TypeSafe quotes response times of 70 to 500 milliseconds, which is what lets a multi-hop trail fill in while you watch. Every question, threshold and rule is published word for word under How a trace is made, and you can copy the whole method as JSON.

How accurate is it

We tested Citation Telephone on three sets of real citations: 257 pairs drawn from 79 open-access papers on 44 topics, from cancer immunotherapy and the gut microbiome to dental caries, coral reefs and protein design. Each pair is a sentence and a paper it cites. The sets include sentences that cite several papers at once, citations of papers whose full text is closed, and controls: the same sentence paired with a different paper from its reference list. We labeled the first set ourselves by reading the cited papers. The second and third were labeled blind by independent reviewers, who read each cited paper without seeing the tool's output. We used the second set to tune the tool and kept the third purely for measurement.

Across all three sets, when the cited paper supports the sentence at least in part, Citation Telephone found the supporting passage in 114 of 125 cases (91%). When it doesn't, the tool reported support in 6 of 72 (8%), and most of those were related papers that it flagged as going further, changed or only partly matching. On the third set, which played no part in tuning, the figures were 38 of 39 (97%) and 1 of 21 (5%). A closed paper read from its abstract was never reported as not saying something.

The labels on each hop are the softest part. When the tool finds support, its comparison agrees with the reviewers' about half the time, mostly because it says "partly matches" where a reviewer saw a full match. We tuned it to err in that direction: "goes further" and "changed" are shown only when the model is confident, so across the two blind sets just two citations that reviewers judged sound got one of those warnings, while every problem the reviewers found was flagged. Its call on whether a passage is the paper's own result or background agreed with the reviewers about three times in four.

The tests also showed the errors to expect. The tool misses subtle swaps of a single term: one citing paper turned "fibroblastic reticular cells" into "fibroblastic reticulocytes", and the hop came back as a match. Vague citing sentences, such as "this correlation has also been found in…" or "detailed discussions of these techniques can be found elsewhere", give it little to look for, and different runs can settle on different passages. The tests also made the problem it measures concrete. Of the 130 real citations the reviewers read, 12 (about one in ten) cited a paper that didn't state the claim or said something else, close to the major-error rate in the research above. So when the tool says a cited paper doesn't support a sentence, check the passage it shows you, and expect that it is often right.

What it can't do

Citation Telephone can only read what is openly available. Full text comes from the Europe PMC open-access subset; papers that are free to read but not openly licensed, and everything behind a paywall, are read from their abstracts. Abstracts rarely say where a claim came from, so many trails end at a closed paper. A good example is the widely repeated claim that the gut holds ten times more bacterial cells than the body has human cells. Many papers cite a 2005 Science review for it, which can only be read as an abstract, while a 2016 re-estimate found the two numbers to be of the same order.

Books, reports, guidelines, websites and papers outside Europe PMC can't be read at all. Claims made only in a figure can be missed, though tables and figure legends are read. The limits of three references per sentence, six hops and 24 papers keep a trace fast; you can continue from any paper where a trail stopped. Above all, the tool checks citations, not truth. An origin can itself be wrong, and a claim can be true and still badly cited.

Fix it with K-Dense Web

A broken trail usually leads to one of a few next steps: find a source that supports the claim as written, rewrite the sentence to match its sources, read the papers that were closed, or check the rest of a paper's citations the same way. The panel next to every trail turns your choice into a brief for K-Dense Web, our AI co-scientist, which can retrieve full texts, search the wider literature and draft the corrected sentence with the right citations.

The Fix it with K-Dense Web panel with next steps to find a source that supports the claim as written, rewrite the sentence to match its sources, read the papers that were closed, check the rest of the starting paper's citations and check the claim against the wider literature

The brief lists the claim, every paper on the trail with its identifiers and result, and the passage each result rests on. It is built from fixed templates, so nothing in it is generated, and K-Dense Web fetches the papers itself.

Private by default

Your claim is sent with the cited papers' public text to be read, and K-Dense doesn't store or log it. The trail lives in your browser tab. A link copied from the page reruns the trace from the same sentence, so sharing a result shares where to find it, not a stored copy.

Try it

Citation Telephone is free, with no account needed. The page opens on the nuclear PD-L1 trail, and more examples, including the AlphaFold, post hoc power and sleep trails above, are one click away. It is most useful on a sentence you're about to cite, a claim in a manuscript you're reviewing, or a number that sounds a little too clean. If it finds something interesting, or gets something wrong, we'd like to hear about it at contact@k-dense.ai.

Run this kind of analysis yourself

K‑Dense Web is an AI co-scientist that plans, runs, and writes up real research — from literature to code to figures.

Enjoyed this article? Share it with others!

Share:
Back to all posts