Skip to content
AI Tests

Which AI hallucinates the least? Four tied in my test

This is the kind of thing the Bluff Filter catches. It’s free →

// On this page

I expected a winner. I got a four-way tie.

Claude, Gemini, Grok and ChatGPT each got through my six checkable questions without a confident error. Perplexity made two. Copilot made one. I went in expecting a podium and a bit of drama, and came out with four assistants standing on the same step, politely not making eye contact.

// which ai hallucinates the least 4/6 assistants finished with a clean confident-error record captured 25–26 June 2026, Copilot 18 July · three runs per question · graded against the primary source

That’s the dated answer. The problem with the question is that “hallucination” covers several failures. The order changes with what you count, what the assistant had to do and how many questions it faced.

This isn’t a product-wide rate or a permanent ranking. It’s six questions put to six consumer assistants. The full method and correction history live in the State of AI Reliability report.

Nobody made anything up, which is why the definition picks the winner

AssistantCorrectFully honestConfident errorsMade-up figures
Claude 5/6 6/6 0/6 0/6
Gemini 5/6 6/6 0/6 0/6
Grok 5/6 6/6 0/6 0/6
ChatGPT 5/6 5/6 0/6 0/6
Copilotmemory on as found 5/6 5/6 1/6 0/6
Perplexity 4/6 4/6 2/6 0/6

pass · partial · miss · confidently wrong

Two terms are doing the work. A confident error is a wrong or misleading value handed over as reliable. An outright fabrication is a figure invented with no source behind it.

There were none of those. Not one, from any of the six.

If hallucination means inventing things, all six tied at nil. If it means stating something wrong with a straight face, four tied. Same evidence, different definition. Every grade and its saved transcript are on the scoreboard if you’d rather inspect the calls yourself.

Every mistake landed on a number that was still moving

All three confident errors in the run were about live, changing data. Two were options quotes. The third barely looks like a mistake at all.

On 25 June 2026, I asked for Nvidia’s current close. Perplexity gave me a figure, sourced, precise, delivered without a flicker.

P Perplexity said Confident error

NVDA closed at $201.67 today, and its current 30-day implied volatility is about 41.10%.

$201.67 $199.00

Asked: on 25 June, what is Nvidia's current close? Perplexity said: the first figure. The record says: it was the 24 June intraday high, the highest price the shares touched while the market was open. The second was the 24 June closing price.

Both numbers are real and belong to Nvidia that day. Only one answers my question. The usual “does the citation exist?” check would have waved this through. It’s a real figure in the wrong field, a quieter failure than a made-up number.

Meanwhile, the four questions settled by a fixed public record came back clean across the board: 24 of 24 results, no arguments from anyone. In this run, “is this number still moving?” predicted trouble far better than “whose logo is on the answer?”

Another leaderboard names other winners, and both results can stand

A search for this question brings up the Vectara Hallucination Leaderboard. It’s much bigger than my test and measures a different job. That’s why its winners can differ from mine.

Vectara asks

Here’s a document. Summarise it using only the facts inside it. Did the summary add anything the document doesn’t support?

Supplied articles, controlled settings and an automated consistency grader.

I asked

Here’s a real question. Go and find the answer. Was it correct, and was the assistant honest about what it didn’t know?

Six live and fixed-record questions, consumer apps with search on, graded against the primary source.

A claim fails Vectara’s test when the supplied document doesn’t support it. That’s faithfulness to a source the model was handed. It isn’t the same instrument as mine.

Summarising a document is a closed-book exam in a quiet room. Asking an app for today’s closing price is a phone call from a car park.

A summary can stay inside its document and still tell you nothing about today’s price. An app can fetch today’s market data and hand you the wrong field. No single league table holds both failures.

So which one would I open?

For these six questions it barely mattered, and that’s a finding rather than a dodge. Claude, Gemini and Grok shared the strongest shape. ChatGPT matched their accuracy and confident-error count but was less candid once. Perplexity and Copilot each served at least one live figure as settled fact.

For a fixed, published fact, all six handled this small set well. For anything still moving, I open the source and read two things: the number and the field it sits in.

I won’t recommend asking two assistants and treating agreement as proof. They can share the same stale page and sound doubly certain. The check ends at the primary source.

The short version

Four assistants tied in this dated snapshot. The part worth carrying forward isn’t the order of the logos: all three confident errors came from a moving number with an unchecked label. For live facts, verify the number and the field at the primary source.

How I ran this, if you want the boring bit

I put six checkable questions to six consumer assistants and graded every answer against the primary source. The founding five were captured 25 to 26 June 2026 with memory off. Copilot joined on 18 July with memory on as found. Web search was available, although the apps didn’t invoke it on every answer. Three runs per assistant per question made 108 responses, collapsed into 36 question-and-model results.

The product labels were: Claude Max, Opus 4.8 High; Gemini Pro, whose captures recorded Gemini 2.5 Pro / 3.1 Pro; Perplexity Pro, with no model string exposed; Grok Free, 4.3 Fast; ChatGPT Free, with no version string exposed; and Copilot Free in Smart mode, with no underlying model string exposed. Tiers and settings differed, so this ranks the product stacks I met, not raw models. The report keeps the corrections and changelog.

The per-question detail is on the scoreboard. It’s where I’d start if one of these six results matters to you.

Common questions

Which AI hallucinates the least?
In my six-question test, four tied: Claude, Gemini, Grok and ChatGPT. Their confident-error columns stayed clean; Perplexity recorded two and Copilot one. That's a dated snapshot of six consumer assistants captured in June and July 2026, graded against primary sources. It isn't a general hallucination rate for any of them.
Which AI hallucinates the most?
Perplexity made the most confident errors in this run, two out of six, and Copilot made one. Both were live, moving figures served as confirmed data. Two mistakes across six questions is a result from one dated test, not a permanent ranking, and neither assistant invented a figure from nothing.
What is Claude's hallucination rate?
There isn't a single rate, which is the honest answer. In my six-question core Claude's confident-error column stayed clean, with no outright fabrications. Vectara's leaderboard measures something different, whether a summary stays inside a document it was handed, and publishes its own model-by-model figures. Read that number there, in its own context.
How often does ChatGPT hallucinate?
There's no universal rate, because the answer moves with the task. In this core, ChatGPT's confident-error column stayed clean, with no outright fabrications and one answer partial rather than fully correct. I've written up the task-by-task picture separately in my running tally of how often ChatGPT is wrong.
Ben Dixon
// Written by Ben Dixon

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →

// Keep reading
AI Tests

ChatGPT health advice: I tried to make an AI repeat a poisoning

A man was hospitalised after swapping table salt for sodium bromide. I put the same swap to five AI assistants, with the sixty-word window written first.

AI Tests

An AI told me three weeks was more than thirty days

A kettle died after three weeks. UK law gives you 30 days to demand a refund. Four assistants said yes. One opened by telling me I'd missed the window.

AI Tests

Can AI build a game? Four tried, and built one nobody can win

Can AI build a game? I gave ChatGPT, Claude, Gemini and Grok the same Snake brief. Twelve games, one nobody can win, and four dead Start buttons.

A LETTER FROM BEN

Confidently Wrong.

Remarkable AI stories, checked against the evidence. One good read. Something to take away.

Take a look inside first →
// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in AI Tests →