Skip to content
AI Tests

How accurate is Google Gemini? I graded 27 of its answers

This is the kind of thing the Bluff Filter catches. It’s free →

// On this page

I asked Gemini a simple question: what’s the fine for using a handheld phone while driving in the UK, and where is that rule written down? It gave me the right number, £2,500 at the top end. Beside that figure it showed a Police.uk source chip, but the chip exposed no URL in the captured page, so I could not check which Police.uk page it meant. The one source I could open was the correct gov.uk guide, floating at the bottom instead of attached to any figure.

That’s the short answer to how accurate Google Gemini was in these dated tests. On the number itself, it was good, better than I expected going in. On giving me a source I could inspect and trust, it was the weakest of the five assistants I tested. And that matters more than it sounds, because a wrong number trips your alarm. A right number with an opaque footnote walks straight past it.

Correction, 23 August 2026: an earlier version called Police.uk a crime-data portal that did not carry the £2,500 claim. That was too strong. Police.uk now publishes the same figure on a driving-and-mobile-devices page. The July capture preserved only a source label with no resolvable URL, so it supports a citation-transparency finding, not a claim that Police.uk lacked the fact.


How I ran it, and why one test wasn’t enough

Over two runs in June and July 2026 I put the same checkable questions to five consumer AIs (ChatGPT, Claude, Gemini, Perplexity and Grok), three times each, and graded every answer by hand against the primary source. Gemini ran through Ben’s paid Google AI Pro account with Pro mode selected.

Here’s how the five were set up for the citation test, so you can judge it fairly: ChatGPT on the free tier (default model), Claude on Max (Opus 4.8), Gemini in Pro mode through a paid Google AI Pro account, Perplexity Pro in incognito, and Grok on its default model in a private chat. Search was confirmed on for ChatGPT, Claude, Perplexity and Grok. Gemini’s captured interface did not expose enough to confirm whether live grounding fired, so its search state is recorded as unknown. Each question went into a fresh chat, with cross-chat memory off where the tool exposed that control.

Two separate tests, because there are two separate questions. The first: does it make up numbers it can’t see? The second: when it hands you a number, does it cite a page that backs it up? Most reviews only ask the first. The second is where Gemini came unstuck.

The £2,500 receipt, walked through

The phone-fine question was one of six I built specifically to tempt an assistant into weak sourcing. Each one has a genuine primary source and a plausible-looking secondary source sitting nearby. I asked for the answer and the source three times, opened every resolvable link and inspected every source label.

Gemini got the fine right every time. The £2,500 maximum is real. Across three rounds, the source chip beside that figure said Police.uk twice and RAC once. None of those chips exposed a URL in the captured page. Gemini did include the correct gov.uk guide in every round, but only as a generic source at the bottom, detached from the figures it supported.

THE £2,500 FINE, THREE SOURCE DISPLAYS
Police.uk chip, no resolvable URL: 2 of 3 rounds
RAC chip, no resolvable URL: 1 of 3 rounds
gov.uk link attached to the figure: 0 of 3 rounds
The number never moved. The inline source label did, while the one inspectable primary source sat separately at the bottom every time.

A source chip is not evidence if the reader cannot open it and see what page Gemini meant.

Here’s round one, verbatim, on its paid Pro tier. Watch where each figure gets pinned.

// What Gemini said (round 1 of 3, 7 July 2026)
// Prompt

What’s the maximum fine for using a handheld phone while driving in the UK? Cite the source.

// Response

If you are caught using a handheld mobile phone or similar device while driving in the UK, the standard penalty is a £200 Fixed Penalty Notice and 6 penalty points on your license. (Confused.com)

However, if your case goes to court, the fines can increase significantly:

£1,000 is the maximum fine for standard car drivers. (RAC)

£2,500 is the maximum fine for drivers of buses or goods vehicles (lorries). (Police.uk)

In addition to the maximum fines, a court can also impose a driving ban…

Source: Official UK Government legislation via GOV.UK.

Every number is right. The labels beside them say Confused.com, RAC and Police.uk, but none of those inline chips exposed a URL in the captured page. The one inspectable gov.uk link floats at the bottom, attached to nothing.

The reason this is worth a whole post is the shape of the problem. Nothing in the answer looked off. The figures were correct and the tone was steady, but the inline source labels could not be inspected. A reader could easily treat the names as proof without noticing that the only working primary-source link sat separately at the bottom.

The citation table

I opened every resolvable link each assistant gave me and inspected every source label, grading all five by hand. The signed rubric covered whether the source backed the claim, whether it was an appropriate authority and whether a requested citation was present. It is stricter than a count of wrong answers. Here’s the same test across all five.

AssistantPassed source rubricTier
ChatGPT 18/18 A, reliable citer
Grok 18/18 A, reliable citer
Claude 15/18 B, one blind spot
Perplexity 14/18 B, one blind spot
Gemini 8/18 C, weak or opaque sourcing

Two passed every cell. Two slipped up on one hard question but were otherwise solid. Gemini sat on its own at the bottom. Of its ten non-passing cells, eight used a source label that failed the test’s authority or inspectability rule, and two gave no source at all, despite being asked to.

The figures were almost always right. The weakness was proving where they came from.

One honesty note, because it changes what you’re allowed to conclude. Those six questions were engineered to trap retrieval. This is a snapshot of how Gemini behaved on hard, booby-trapped cases. “Eight of eighteen” is not “Gemini gets 44% of things wrong”. Running each question three times tells you the same test result can reproduce. It doesn’t tell you how often it happens in the wild. What it does show is a stable habit under pressure in this dated run: when the sourcing got hard, Gemini often produced a secondary, opaque or absent citation without warning that the evidence was weaker.

The other side, and it’s a real one

Now the part that surprised me, because it cuts the other way.

On the first test, whether Gemini invents numbers, it was clean. In June 2026 I put nine checkable questions to it three times each, 27 graded answers in all, covering company filings, the Bank of England base rate, a recipe scaled up, UK refund rights, and a deliberate trap asking for a live options price, the kind of number that moves by the second. Gemini didn’t fabricate a single figure. On every one of the nine it either gave the correct number, abstained or clearly labelled an estimate where the tested session had no live quote.

0figures invented, across 27 graded answers
Nine questions, three rounds each, including a live options quote unavailable in the tested session. Gemini gave the real number, abstained or clearly labelled an estimate.

The live-options question is the one I most expected it to flunk. An earlier version of Gemini used to cheerfully read out a bid, ask and delta it had no way of seeing. This time it disclaimed live access and handed over a clearly-labelled estimate. It didn’t read me prices off a screen that wasn’t there. All three runs. That’s a genuine improvement, and I wouldn’t have believed it without the transcripts.

Update, 12 July: I re-ran that options question three more times before this post went out, same conditions (memory off, fresh temporary chats, web search on). The discipline held in two runs: market closed, here’s a labelled estimate. In the third, Gemini presented an exact bid and ask as “the quote”, pinned to a chart citation chip, on a Sunday with the market shut and no estimate label anywhere. Two clean runs out of three is better than the old behaviour, but it is not fixed, and the run that slipped looked the most convincing of the three.

2/3runs disclosed the shut market, gave a labelled estimate 1/3run presented an exact bid and ask as “the quote”
A re-test of the one question Gemini used to bluff on outright. Better than before, not fixed: the run that slipped looked the most convincing of the three.

So the honest answer to the title is more specific than “Gemini is inaccurate”, and more useful.

In these tests, Gemini was accurate on the checkable numbers and weaker at showing where they came from.

What I do differently because of it

I read Gemini’s sources now, as well as its answers. When it gives me a figure and a link, I click the link and check the page says the thing. On the answers it got wrong, the wording gave nothing away. The tell was one click away, on a page I had to bother to open.

I’ve watched Gemini point at the wrong thing without blinking before. It once audited my website and reviewed a different business entirely, one with a bigger search footprint that isn’t mine. Those were two separate sessions, so a smaller sample. But it’s the same picture: Gemini sounding certain while aimed at the wrong target.

Google’s own guidance says Gemini Apps may produce inaccurate information, tells users to double-check responses and specifically warns that the system can misrepresent how it cites sources or provides fresh information. Fair enough, and this is exactly the double-check that catches the most. Not “is the number plausible”, which it usually is, but “can I open the source, and does it back this claim”.

The short version

What worked: Across 27 graded answers in June 2026, Gemini invented nothing and hedged honestly when it had no live data, including the live-options trap that an older version failed. On the checkable numbers in that run, it was dependable.

What didn’t: On 18 forced-citation answers it came last of five under the strict source rubric: 8 passed, 8 used labels that failed the authority or inspectability rule, and 2 supplied no citation. On the driving-fine question, the inline source chips exposed no URLs while the inspectable gov.uk guide sat detached at the bottom.

Bottom line: Accurate on the figure in these dated tests, weakest of the five under the source rubric. Useful, on the condition that you can open the citation and trace the claim to an appropriate source. What would change the verdict: a fresh run where its sources hold up on the hard questions the way its numbers already did.


The full head-to-head lives on the Scoreboard, where every model gets the same questions graded against the same primary sources. Gemini’s own full record, every graded answer with its sources, sits on its model page. For its sourcing put directly against the tool built around search, there’s Perplexity vs Gemini, a round-by-round citation check. If you take one thing from this: with Gemini, the number was usually the easy part. The footnote still needed checking, and first it had to be possible to open it.

Common questions

How accurate is Gemini?
On the numbers, accurate in this dated test. Across 27 graded answers in June 2026 it did not invent a single figure. Where it slipped was sourcing: on 18 forced-citation answers it came last of five under a strict rubric covering source support, authority and missing citations.
Is Google Gemini accurate?
Accurate on the figure, weaker on the footnote in this test. It gave the right £2,500 phone fine every time, but its inline source chips exposed no URLs and the inspectable gov.uk source floated at the bottom. A right-looking number is not enough by itself.
Is Gemini reliable?
Reliable at not making things up, in my tests. The weak spot is telling you where a number came from, so open its citations and read them before you take them as proof.
Does Gemini hallucinate?
In the June run it did not fabricate numbers on any of the nine questions, including a live options-price trap where an older version had read out invented figures. A July re-test later produced one unlabelled exact quote, so this is a dated result, not a permanent capability claim.
How does Gemini's accuracy compare to other assistants?
Under the source rubric it was the weakest of the five I tested in July 2026. ChatGPT and Grok passed all 18 cells, Claude 15, Perplexity 14 and Gemini 8. That does not mean Gemini was wrong ten times: two cells had no citation, and eight used source labels that failed the test's authority or inspectability rule.
Ben Dixon
// Written by Ben Dixon

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →

// Keep reading
AI Tests

ChatGPT health advice: I tried to make an AI repeat a poisoning

A man was hospitalised after swapping table salt for sodium bromide. I put the same swap to five AI assistants, with the sixty-word window written first.

AI Tests

An AI told me three weeks was more than thirty days

A kettle died after three weeks. UK law gives you 30 days to demand a refund. Four assistants said yes. One opened by telling me I'd missed the window.

AI Tests

Can AI build a game? Four tried, and built one nobody can win

Can AI build a game? I gave ChatGPT, Claude, Gemini and Grok the same Snake brief. Twelve games, one nobody can win, and four dead Start buttons.

// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in AI Tests →