Claude Hallucination Rate: What the Published Data Actually Shows
TL;DR
- Three published numbers put Claude's hallucination rate at 9.8%, 60.8% and 6% higher than the last model - different tasks and different denominators, and Anthropic itself says its newest model invents facts more often than the one it replaced.
There is no single published Claude hallucination rate, and the three numbers that do exist are 9.8%, 60.8%, and "6% higher" than the previous model. They are all correct. They measure summarising a document you supplied, guessing instead of abstaining on closed-book questions, and a relative change between two versions — three different tasks on three different datasets. Anyone quoting one figure as the rate has not read what it measures.
Everything below comes from Anthropic's own system card and model documentation, the AA-Omniscience benchmark paper and its live leaderboard, and Vectara's hallucination leaderboard repository, all retrieved 2026-09-18. No vendor marketing is used as evidence.
Key Takeaways
- Anthropic does publish hallucination measurements, in section 6.5 of the Claude Opus 5 system card, but as a net score and relative deltas rather than a rate (Claude Opus 5 system card, retrieved 2026-09-18, SOURCED).
- The newest model is not the most factual. Anthropic states that Opus 5 "hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall" (same source, SOURCED).
- On summarising supplied documents, Claude models score 9.8% to 12.2% hallucination — and no Claude 5-generation model is on that leaderboard at all (Vectara hallucination leaderboard, last updated 11 May 2026, retrieved 2026-09-18, SOURCED).
- On closed-book questions, Claude Opus 5's "hallucination rate" is 60.8% — which means the share of times it guessed rather than abstained when it did not know, not the share of its answers that are wrong (AA-Omniscience, retrieved 2026-09-18, SOURCED).
- Anthropic publishes no error rate for long-document work, which is the task most business readers actually care about.
The short answer
Three organisations publish something you could call a Claude hallucination rate. Each measures a different failure on a different dataset, and the figures differ by a factor of six as a result.
| What was measured | Dataset | Latest Claude figure | Source | Retrieved |
|---|---|---|---|---|
| Faithfulness of a summary to a supplied document | 7,700+ articles, temperature 0 | 9.8% (Haiku 4.5) to 12.2% (Opus 4.6) | Vectara HHEM-2.3 leaderboard | 2026-09-18 |
| Guessing instead of abstaining on closed-book questions | AA-Omniscience, 6,000 questions | 60.8% (Opus 5, max effort) | Artificial Analysis | 2026-09-18 |
| Change in factual hallucination between two versions | AA-Omniscience public split | 6% higher than Opus 4.8 | Anthropic Opus 5 system card | 2026-09-18 |
These cannot be averaged, ranked against each other, or quoted interchangeably. The rest of this post is what each one means.
What Anthropic publishes about Claude Opus 5
The Claude Opus 5 system card has a section titled "Honesty and hallucinations". It is worth reading directly, because it is more candid than a vendor document has to be.
The measurement. Anthropic states: "We measured factual accuracy on the public split of AA-Omniscience, a 41-topic, closed-book benchmark drawn from various economic and academically relevant domains. The model is given no web search or knowledge-base access when answering, and must instead answer from its own knowledge. Each answer is graded as either correct, incorrect, or an abstention" (retrieved 2026-09-18, SOURCED).
The result. "Claude Opus 5 received a net score of 0.49, which places it in between Opus 4.8 and the two Mythos models." Broken down by grade, Anthropic reports that Opus 5's "accuracy is 11% higher than Opus 4.8, but its rate of hallucinations is also 6% higher" (same source, SOURCED).
That trade-off is stated plainly twice more in the card's own summary of findings:
- "The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall."
- "We found a surprising number of cases in which Opus 5 confidently stated an answer about which it was in fact unsure."
ANALYSIS: a vendor publishing that its newest flagship invents facts more often than the model it replaces is the single most useful line in this whole topic, and it is the opposite of what a procurement deck would tell you. It also explains why "we upgraded to the newest model" is not a hallucination control.
A second, separate kind of hallucination. The card treats "input hallucination" as its own dimension, defined as "Hallucinating or materially misrepresenting the contents of files, tool outputs, or past user turns". On that measure, Anthropic reports Opus 5 "misleads the user at rates similar to or lower than Opus 4.8, Mythos 5, and Sonnet 5. The one exception is input hallucination, where the mean rose slightly but within the range of expected noise" (retrieved 2026-09-18, SOURCED).
So there are two Claude hallucination measurements inside one document, pointing in slightly different directions, and neither is expressed as a percentage of outputs.
What the card does not give you. The grade breakdown behind the 0.49 net score is published as a chart, not a table, so the exact correct, incorrect and abstention percentages cannot be read out of the document text. Anthropic's figures for honesty under pressure use the public MASK split, reported as "the fraction of prompts where the model did not contradict its own elicited belief under pressure" with 95% confidence intervals and n=904 — again as a chart.
What the closed-book number actually means
The 60.8% figure is the one most likely to be misread, so it needs its definition in full.
AA-Omniscience is a 6,000-question benchmark covering "42 economically relevant topics within six different domains", scored on an Omniscience Index described as "a bounded metric (-100 to 100) measuring factual recall that jointly penalizes hallucinations and rewards abstention when uncertain, with 0 equating to a model that answers questions correctly as much as it does incorrectly" (AA-Omniscience, arXiv:2511.13029, retrieved 2026-09-18, SOURCED).
Its hallucination rate is "defined as the proportion of questions it attempted to answer when it was unable to get the answer correct" (same paper, SOURCED).
Read that twice. It is a denominator of questions the model got wrong or did not know, not a denominator of all answers. A hallucination rate of 60.8% means that when Claude Opus 5 did not have the answer, it produced one anyway about six times in ten instead of saying it did not know. It does not mean 60.8% of its answers are false.
The Claude rows on the live leaderboard, read on 2026-09-18:
| Model variant, as listed | Omniscience Index | Accuracy | Hallucination rate |
|---|---|---|---|
| Claude Fable 5.1 (max with fallback) | 43.45 | 67.2% | not in the data read |
| Claude Fable 5 (with fallback) | 43.30 | 65.3% | 63.6% |
| Claude Fable 5.1 (xhigh with fallback) | 42.38 | 66.2% | not in the data read |
| Claude Fable 5.1 (high with fallback) | 40.80 | 64.9% | 68.8% |
| Claude Opus 5 (max) | 37.07 | 60.9% | 60.8% |
Source: Artificial Analysis, AA-Omniscience, retrieved 2026-09-18, SOURCED. The leaderboard lists many effort-level and fallback variants per model; these are the Claude rows and figures we could read on that date.
Two things follow. The variant with the highest accuracy is not the one with the lowest hallucination rate, which is the calibration point the benchmark exists to make. And the benchmark's authors found this pattern across labs: at launch, only three frontier models scored above zero on the index, and "high hallucination is the dominant factor driving these low scores" (paper, SOURCED).
These figures move fast. The paper, submitted November 2025, reported Claude 4.1 Opus with the highest index of the models then evaluated, at 4.8. The Claude rows above sit between 37 and 43 on the same index. Ten months, an order of magnitude. Any post quoting an index score without a date is quoting nothing.
One precision note. Anthropic describes the public split it used as covering 41 topics; the benchmark paper describes the full dataset as 42 topics across six domains. We report both as each source states them and did not resolve the difference.
What the summarisation number actually means
Vectara's leaderboard measures one narrow thing well: "how often an LLM introduces hallucinations when summarizing a document" (leaderboard repository, last updated 11 May 2026, retrieved 2026-09-18, SOURCED). Models summarise over 7,700 articles at temperature 0, and the hallucination rate is the complement of the factual consistency rate.
The Claude rows, in full:
| Model, as listed | Hallucination rate | Answer rate |
|---|---|---|
| claude-haiku-4-5-20251001 | 9.8% | 99.5% |
| claude-sonnet-4-20250514 | 10.3% | 98.6% |
| claude-sonnet-4-6 | 10.6% | 99.9% |
| claude-opus-4-5-20251101 | 10.9% | 98.7% |
| claude-opus-4-1-20250805 | 11.8% | 92.4% |
| claude-opus-4-20250514 | 12.0% | 91.0% |
| claude-sonnet-4-5-20250929 | 12.0% | 95.6% |
| claude-opus-4-7 | 12.0% | 98.0% |
| claude-opus-4-6 | 12.2% | 99.8% |
Source: same repository, retrieved 2026-09-18, SOURCED.
The answer-rate column is not decoration. A model that declines more documents has fewer opportunities to introduce an unsupported claim, so a rate of 12.0% at a 91.0% answer rate and 12.0% at a 98.0% answer rate are not the same performance (ANALYSIS). Vectara filters non-answers out of the scoring and publishes the column precisely so the comparison stays honest.
The gap that matters most. The newest Claude models on this leaderboard are the 4.x generation. No Claude Opus 5, Sonnet 5 or Fable 5.1 row exists on it, while Anthropic's current model documentation lists exactly those as the current line-up (Anthropic models overview, retrieved 2026-09-18, SOURCED). For the model you are most likely to be using today, there is no independent summarisation figure at all.
Vectara also states the limit of the whole approach, and it is the sentence to remember: "Determining hallucinations is impossible to do for any ad hoc question as it's not known precisely what data every LLM is trained on" (same source, SOURCED). That is why the leaderboard measures faithfulness to a supplied document instead, and why its numbers do not transfer to open-ended questions.
Why these numbers cannot be combined
The same word covers three different failures, and each benchmark picks one:
| The failure | Where it shows up | Which number covers it |
|---|---|---|
| Inventing content a supplied document does not support | Summaries, RAG answers, document Q&A | Vectara, 9.8-12.2% |
| Guessing when it does not know instead of abstaining | Closed-book factual questions | AA-Omniscience, 60.8% for Opus 5 |
| Misrepresenting files, tool outputs or earlier turns | Agentic and long-session work | Anthropic's input-hallucination dimension, reported as relative change |
ANALYSIS: a single "hallucination rate" for a model is therefore not a conservative simplification, it is a category error. The question that has an answer is narrower: on my task, with my documents, how often is this model wrong, and does it tell me when it is unsure? The same reasoning applies to every accuracy figure a vendor quotes, which is the argument in what "99% accurate" actually means.
What is not published
Stated plainly, because these gaps are the finding:
- No error rate for long-document work. Anthropic publishes none, and that is the task most teams are buying the model for. The documented controls that do exist are in how to reduce errors in long documents.
- No figure for the current generation on an independent summarisation benchmark, as above.
- No rate for citation fabrication specifically, by Anthropic or by either leaderboard. The causes of that failure, and the prompting that reduces it, are in how to stop ChatGPT from making things up — the mechanism is the same whichever model is generating.
- No exact grade breakdown behind the system card's 0.49 net score, because it ships as a chart.
- Nothing measured on your documents. Every figure here is someone else's corpus. The only number that describes your workload is one you produce, and the method is in the 10-prompt AI accuracy test.
Anthropic's Transparency Hub does add qualitative comparisons for its newest releases — its automated audit "finds that Claude Mythos 5.1 is somewhat more honest and hallucinates less than previously released models", and that it "hallucinates inputs, meaning it invents or materially misrepresents the contents of files or earlier user messages, significantly less than previous models in our investigations" (retrieved 2026-09-18, SOURCED). Useful, directional, and still not a rate.
How to use this when someone quotes you a number
- Ask which task the figure measures. Summarisation faithfulness and closed-book recall are not the same claim about the same thing.
- Ask for the denominator. A 60.8% hallucination rate with "when it did not know" attached is a calibration statistic; without that clause it reads as a catastrophe.
- Ask for the date and the model string.
claude-opus-4-6and Claude Opus 5 are different rows with different numbers, and one of them is not on the leaderboard. - Ask what the abstention rate was. Accuracy without it tells you nothing about whether the model guesses.
- Then test it on your own material. The published figures narrow the choice; they do not answer it.
The vendor-side version of these questions, with what a deflection sounds like, is in 8 accuracy questions to ask any AI vendor. The same documentation-first method applied to another provider is in Gemini accuracy: what Google's own documentation says, and the general pattern of which error types reach published pages is the 14 AI content errors.
How This Guide Was Sourced
Written and maintained by the LogicBalls editorial team (logicballs.com). Disclosure: LogicBalls builds AI writing tools, and routes work to models including Claude-class models. Every limitation described here applies to output produced through our own product.
AI involvement. AI assisted with research and drafting. Every quotation, figure and table row above was matched against its source text on 2026-09-18 — the system card PDF, the benchmark paper PDF, the leaderboard repository README and the live leaderboard data — before use, and the two tables were checked row by row against the published figures.
Sources. The Claude Opus 5 system card, section 6.5 (Honesty and hallucinations) and section 6.4.3 (Misleading users); Anthropic's models overview and Transparency Hub; AA-Omniscience (arXiv:2511.13029) and its live leaderboard at Artificial Analysis; Vectara's hallucination leaderboard repository README, last updated 11 May 2026.
What could not be verified. The exact correct, incorrect and abstention percentages behind the system card's 0.49 net score are published only as charts and are not stated in the document's text. Two of the five Claude variants on the AA-Omniscience leaderboard had no hallucination-rate value in the data we could read, and the table says so. Anthropic's public split is described as 41 topics against the paper's 42, and we did not resolve which is current.
What is not claimed. No single Claude hallucination rate, no ranking of Claude against any other model as a purchase, and no figure of our own. We ran no test; every number here is external, dated and linked.
No LogicBalls telemetry is used in this guide.
Frequently Asked Questions
What is Claude's hallucination rate?
There is no single published figure. On summarising supplied documents, Claude models score 9.8% to 12.2% on Vectara's leaderboard. On closed-book questions, Claude Opus 5's AA-Omniscience hallucination rate is 60.8%, meaning how often it guessed rather than abstained when it did not know. Anthropic itself publishes a net factuality score and relative changes, not a rate.
Does Anthropic publish hallucination data at all?
Yes. Section 6.5 of the Claude Opus 5 system card reports factual accuracy on the public split of AA-Omniscience, a net score of 0.49, and honesty under pressure on the MASK split. It also treats input hallucination as a separate dimension in section 6.4.3.
Is the newest Claude model the most accurate?
Not on this measure, and Anthropic says so. The Opus 5 system card states the model "hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall". Higher accuracy and lower hallucination do not arrive together automatically.
Why is one number 10% and another 60%?
Different denominators. Vectara's 10% is the share of document summaries containing an unsupported claim. AA-Omniscience's 60% is the share of occasions the model answered anyway when it did not know the answer. Both are correctly reported and neither substitutes for the other.
Which figure should I use to choose a model?
Whichever matches your task, plus the abstention behaviour. For summarising and RAG, the faithfulness leaderboard is the closer analogue. For closed-book questions, calibration matters more than accuracy. For anything consequential, test on your own documents rather than adopting someone else's corpus.
Related reading
- AI Content Errors: The 14 Mistakes That Reach Published Pages
- How to Stop ChatGPT From Making Things Up
- Claude Accuracy: How to Reduce Errors in Long Documents
- Gemini Accuracy: What Google's Own Documentation Says
- AI Accuracy Rate: What "99% Accurate" Actually Means When You Buy AI
- Why AI Makes Up Statistics (And How to Get Real Ones Instead)