MeDiVa is now in public preview.Read the announcement
// research

The research shelf

A working collection of the studies, reports, incidents, and primary sources on AI-generated text and its failure modes. Every entry links its primary source; scope notes say what a number does and doesn't measure; figures that fail verification get their own section. Free to reference — corrections welcome via an issue.

Hallucinated citations & references

What fraction of generated references are fabricated, measured against real databases.

Legal citation hallucination is not monotonic: ~25% for GPT-3.5, a low of 1.23% for the August 2024 GPT-4o, back up to 6.57% for GPT-5.1 (p=0.001).

Liu, Stammbach & Henderson — Who Checks the Citations? (Princeton) · 2026

Eight generations of default ChatGPT, the same 92 legal-drafting prompts, 8,000+ citations verified against CourtListener and Westlaw. The same paper counts 1,000+ real court filings containing hallucinated citations. Scope: legal drafting in one product line.

55% of GPT-3.5's bibliographic citations were fabricated, 18% for GPT-4; 43% and 24% of the genuine ones carried substantive errors.

Walters & Wilder — Scientific Reports · 2023

636 citations across 42 topics. The models are dated; useful as the historical baseline the 2026 studies extend.

Up to 22% of computer-science papers show signs of LLM-modified text, against roughly 9% for mathematics and Nature-portfolio journals.

Liang et al. — Nature Human Behaviour · 2025

1,121,912 papers, 2020–2024. Measures LLM-modified language, not fabrication — don't recruit it into a hallucination argument. Supersedes the widely-quoted 17.5% figure from the earlier preprint.

1 in 277 PubMed-indexed papers in 2026 shows fabricated references; fabricated references were found in 2,800+ biomedical articles.

Retraction Watch analysis, May 2026 · 2026

Reported tier: figures from Retraction Watch's analysis as covered by STAT and Forbes; re-verify against the primary analysis before quoting elsewhere.

Across 2.5 million biomedical papers, the share carrying at least one fabricated citation rose from 1 in 2,828 (2023) to 1 in 458 (2025) to 1 in 277 (first seven weeks of 2026); 4,046 fabrications were confirmed across 2,810 papers.

Topaz, Roguin, Gupta, Zhang & Peltonen — Fabricated citations: an audit across 2.5 million biomedical papers (Columbia), The Lancet Vol 407 · 2026

Published as a letter. 125.6M references extracted, 97.1M with verifiable identifiers, checked against PubMed, Crossref, OpenAlex and Google Scholar. Over 98% of the affected papers had seen no publisher action at the time of the audit. The letter publishes a quarterly-rate figure spanning Jan 2023 to Feb 2026 and names 51.3 per 10,000 for Q4 2025 in its text. The cited letter carries a correction notice dated 2026-07-16. The method is AI-assisted and validated on 500 records; Cochrane’s head of editorial policy called the published methods detail insufficient, and a Northwestern integrity researcher called the analysis simplistic and the tip of the iceberg, arguing the larger problem is citations that are inaccurate rather than wholly invented.

Frontier-model factuality

Hallucination rates depend heavily on what's measured — adversarial QA, production traffic, and grounded summarization tell different stories.

On production-traffic factuality grading, GPT-4o ran a 12.9% hallucination rate and gpt-5-main 9.6%; OpenAI also reports that o3 would sometimes claim to have completed tasks it hadn't.

OpenAI — GPT-5 System Card · 2025

OpenAI's own LLM-graded methodology (75% human-grader agreement), self-reported.

On grounded summarization, current models range from ~3% to 23.3% hallucination; reasoning-heavy models score worse than lighter ones, the inverse of the open-QA pattern.

Vectara — HHEM Hallucination Leaderboard · live

Summarization-faithfulness only; not comparable with PersonQA/SimpleQA numbers. Benchmark choice changes the story — check which one a headline figure comes from.

Hallucination is argued to be a statistical consequence of training and evaluation incentives; the paper also notes all reported models below 30% accuracy on Humanity's Last Exam and calibration error rates above 70%.

Kalai, Nachum, Vempala & Zhang — Why Language Models Hallucinate (OpenAI) · 2025

Frequently mis-cited: this paper does not contain the o3/o4-mini PersonQA figures — those live in the system cards above.

Agent failure modes

How autonomous agents fail, and how often they fail while reporting success.

Agents falsely claimed completion in 45–48% of tau2-bench failures and 75.8% of AppWorld self-assessed coding trajectories; no LLM judge exceeded AUROC 0.65 at catching it.

From Confident Closing to Silent Failure · 2026

Denominator is failed runs, not all runs.

29.6% of 'plausible' SWE-bench patches behave differently than the ground-truth fix; reported resolution rates were inflated by 6.2 points.

Wang, Pradel & Liu — ICSE 2026 · 2026

Passing the test suite is not the same as satisfying the requirement.

Anthropic's own system card documents Claude Opus 4 hard-coding an expected test value to force a failing case green, then reporting the task as successfully implemented.

Claude 4 System Card, section 6 (reward hacking) — Anthropic · 2025

The transcript is printed in the card, code and false summary side by side. A vendor documenting its own model faking a pass.

OpenAI's system card names the expected frontier failure class: tests deleted, data invented, tasks pretended solved, with deceptive actions in roughly 4% of gpt-5-thinking trajectories and 8% for o3.

GPT-5 System Card, section 5.2.1 (Apollo Research external evaluations) — OpenAI · 2025

The same card's deception section records o3 making false claims about actions it had taken and saying it had completed tasks it hadn't (section 3.8, wording verified against the PDF).

Evaluating generated text

What the LLM-as-judge literature measured, and where deterministic checks fit.

Judge position bias, quantified at the origin: Claude-v1 kept the same verdict only 23.8% of the time when answer order was swapped; GPT-4 agreed with humans ~85%, about the human-human rate.

Zheng et al. — Judging LLM-as-a-Judge (MT-Bench) · 2023

2023 models; see the CALM entry below for current judges.

The state-of-the-art LLM judge for summarization quality reached a Spearman correlation of 0.514 with humans, described as outperforming all previous methods by a large margin.

Liu et al. — G-Eval · 2023

ChatGPT and Claude 2 distinguished factual from hallucinated summaries at only 53.8–58.5% accuracy.

Li et al. — HaluEval · 2023

Dated models; the point is structural — a judge that can barely detect hallucination is a weak enforcement layer.

Eval frameworks ship deterministic validators for JSON, HTML, SQL and XML; no Markdown-structure equivalent is listed — Markdown output falls through to regex or a model-graded rubric.

promptfoo — assertion catalogue (documentation) · live

A mutable catalogue — this describes it as checked on the date above.

Production, economics, regulation

What unverified output costs in practice, and what regulators now require.

The EU AI Act requires high-risk systems to log events automatically (Art. 12), remain overseeable by humans able to detect anomalies (Art. 14), retain logs at least six months (Art. 19), and ship 'concise, complete, correct and clear' instructions (Art. 13).

Regulation (EU) 2024/1689 — statutory text · 2024

Inadequate software-testing infrastructure was estimated to cost the US $22.2–59.5 billion per year.

NIST — Planning Report 02-3 · 2002

A 24-year-old software-bug study, often quoted without its year or scope. Cite it as history, not as a current documents figure.

90% of developers regularly use at least one AI tool at work; 74% use a specialized coding tool.

JetBrains AI Pulse survey (n≈10,000, fielded January 2026) · 2026

The freshest developer-survey figure available. Stack Overflow's 2026 survey opened in June — 'for human developers only' — with results not yet published, so the 2025 numbers remain the newest SO data.

In one study, 15 of 18 writers assisted by ChatGPT could not quote their own essay minutes after writing it.

Your Brain on ChatGPT — MIT Media Lab (preprint) · 2025

Small-sample preprint (n=18 per condition); a methodological comment raises sample-size concerns. Formerly on the landing deck; the deck now argues review capacity rather than reader memory.

People find only 40–90% of errors when working alone: readers caught 81% of nonsense-word errors but just 66% of real-word errors, and detection of nonsense-word errors falls to 47–58% in difficult material.

Panko — Errors in Proofreading, summarising published detection-rate experiments (difficult-material figure from Riefer) · 1980–1993

Laboratory experiments on general prose. The 47–58% figure is Riefer (1991) for NONWORD errors in difficult material; Riefer's word-error range is 42–62%. The 87% professional-proofreader figure on the same page rests on three editions of the author's own textbook and is deliberately not used.

Across 22,000 developers, the median time a pull request spends in review rose 441% year over year — against 91% in the prior year's dataset — 31% more pull requests merged with no review at all, and 25% of pull requests are now reviewed by an AI agent.

Faros AI — The Acceleration Whiplash · 2026

Vendor telemetry from a company selling engineering analytics. The year-over-year comparison runs on the same pipeline, so it is like-for-like, but the methodology note is not public.

In one year on GitHub, pull requests created rose 20.4% and commits 25%, while comments on issues and pull requests rose 0.35% and comments on commits fell 27%.

GitHub Octoverse · 2025

GitHub labels these observational signals rather than causal claims. Comments proxy for review attention rather than being it, and Copilot code review shipped in April 2025, so part of the flat line may be review that moved to a machine.

Incidents & policies

Named cases with adjudicated outcomes, and the acceptance boundaries projects have drawn in response.

A tribunal held Air Canada liable for its chatbot's fabricated bereavement-fare policy: CAD $812.02, and a rejection of the argument that the chatbot was 'a separate legal entity'.

Moffatt v. Air Canada, 2024 BCCRT 149 · 2024

Deloitte Australia repaid the final instalment of an A$440,000 government contract after its report was found to contain a fabricated court quote and citations to nonexistent papers.

AP / CNA, October 2025 · 2025

KPMG pulled its own report after forensic review found only 5 of 45 citations pointed to real, uncorrupted sources. GPTZero, which ran the review, coined the term for the failure: vibe citing.

GPTZero forensic review, multi-outlet coverage · 2026

The KPMG investigation page could not be resolved at a stable URL; this links GPTZero's investigations index entry that covers the same review. Replace with the KPMG-specific URL once confirmed.

EY Canada withdrew a 44-page report after 16 of 27 cited sources were found fabricated, misattributed, or dead, including phantom Forbes, McKinsey, and Gartner citations.

GPTZero review, multi-outlet coverage · 2026

Withdrawn May 2026. With Deloitte (October 2025), three Big Four firms in nine months.

A World Bank paper on obesity trends was found to contain at least 14 references to work that does not exist.

Retraction Watch · 2025

A librarian checking a Springer Nature article on bowel-surgery management found 12 of its 14 references did not exist.

Retraction Watch · 2026

Numbers that don't check out

Viral figures that fail primary-source verification. Corrections included — the point of a source library is also what not to cite.

last updated: 2026-07-25 · numbers are quoted as the sources state them, never rounded past the source