Skip to content
AI Tests

AI cites the wrong source: I put 6 UK questions to 5 assistants and checked every inspectable source

This is the kind of thing the Bluff Filter catches. It’s free →

// On this page

The £2,500 figure was right. Gemini told me that using a handheld phone while driving a lorry or bus in the UK can cost you a £2,500 fine, and it can. Beside the figure it displayed a Police.uk source label.

The important detail is what the capture does and does not preserve. The label exposed no destination URL, so I could not inspect the page Gemini meant. The only resolvable receipt was the correct GOV.UK guide, floating separately at the bottom. Police.uk now carries the same figure on its own driving-and-mobile-devices page. The answer was numerically right; its inline provenance was opaque.

That is the whole test, and the whole worry. A dead link gives itself away. A plausible label can look just as reassuring while giving you nothing to inspect. And when a real link does open, a trusted domain can still hide a page that does not carry the claim. Both failures demand the same response: require the exact page, then find the claim on it.

ONE ANSWER, GRADED TWICE
The number: £2,500 is the real maximum
The inline Police.uk label exposed no URL
The inspectable GOV.UK guide sat detached below
Grade the answer and it passes. Grade the inline provenance under it and it doesn’t. Nothing on the screen tells you which one you’re looking at. Gemini (Pro), 7 July 2026.

What I did

I took six everyday UK questions and put each one to five AI assistants on 7 July 2026, each in its default model with memory off and a fresh chat each time. Web search was confirmed on for ChatGPT, Claude, Perplexity and Grok; Gemini’s captured search state is unknown. The tiers weren’t identical, and that matters, so here they are: ChatGPT on the free tier, Claude on Max (Opus 4.8), Gemini on Pro, Perplexity on Pro, Grok on its default fast model. The board that follows is one run per assistant per question; I then put the whole set through twice more the same day, three rounds in all, which I come back to near the end. I opened every resolvable citation against a source list verified before the run. A source label with no recoverable URL could not pass that audit.

I engineered each of the six questions to tempt a lazy retrieval toward a plausible-but-wrong page, the sort of everyday thing a person actually asks: online returns and your rights, stamp duty on a £300,000 home, free childcare hours for a 9-month-old, the State Pension age, the speed limit in Wales, the maximum fine for using your phone at the wheel. The trap in each case is a genuine, authoritative-looking page that sits right next to the correct one. Change-of-mind returns are governed by one law, but a more famous law about faulty goods has a real page a model can grab. Stamp duty is a different tax in Scotland and Wales, but the England page is the one everyone links. Childcare rules changed last year, and the old page is still up.

That’s one run each, which makes this a snapshot of a single day, and I’ll come back to what that does and doesn’t let me say. Worth flagging now: ChatGPT was on the free tier; Grok used its default Fast model, but the saved run record does not preserve its account tier. So this is not evidence that paid tools beat free ones.

The board

Five assistants, six questions each: here’s who held and who slipped.

AssistantHeldSlippedWhat stood out
ChatGPT 6 0 Flagged the Scotland and Wales tax split unprompted
Grok 6 0 Cleanest sourcing of the lot, one correct page, no clutter
Claude 5 1 Sourced the fine to a page that doesn’t carry it
Perplexity 4 2 Led a £1,000 court fine with a solicitor’s page over gov.uk
Gemini 3 2 The most opaque provenance of the group

A word on how I’m counting a “slip”, because three different faults sit in that column. One is an inspectable page that does not support the specific claim. Another is opaque provenance: a source label with no recoverable page. The third is a correct page that is incomplete: right figure, right official source, but it leaves out something a reader needs. Gemini has the second and third. I’ll be clear below about which is which. (Gemini’s sixth answer, on the returns question, sits in neither column: it gave the right answer but attached no source at all when I asked for one, a fourth kind of miss I’ve left uncounted and haven’t forced into “held” or “slipped”.)

Two things to say straight away, because the headline “AI cites the wrong source” can slide into “all AI is broken”, and that isn’t what I found. ChatGPT and Grok cited a page that held the claim on all six questions. This is a real split. The failures below sit in the answer-source pair: sometimes the page is weak, sometimes the provenance is opaque, and sometimes the answer leaves out where its source applies. The answer is often right, which turns out to be the more unsettling half.

// the board 3/5 failed a strict source rubric: unsupported page, opaque provenance or incomplete jurisdiction

The three that slipped, with the record

Gemini made the inline sources impossible to inspect. On the handheld-phone question, Gemini gave the right penalties, £200 fixed penalty and six points, £1,000 in court, £2,500 for a lorry or bus. It put bare Confused.com, RAC and Police.uk labels beside those figures, none with a recoverable destination URL in the capture. Here is the relevant stretch, exactly as it came back:

Gm Gemini said, 7 July 2026 Opaque provenance

£1,000 is the maximum fine for standard car drivers. RAC

£2,500 is the maximum fine for drivers of buses or goods vehicles (lorries). Police.uk

The one genuine GOV.UK link sat at the very bottom as a generic source, attached to nothing in particular. That page carries all three figures. The inline labels were bare text, so the historical pages behind them cannot be checked from the saved record. A current Police.uk page carries the £2,500 figure too, which is why the honest finding is opacity and detached provenance, not that Police.uk lacked the fact.

Perplexity gave a childcare figure that its own source contradicts. I asked how many hours of free childcare a working parent of a 9-month-old in England gets right now. Here is what it opened with, word for word:

Px Perplexity said, 7 July 2026 Confidently wrong

A working parent of a 9-month-old in England is currently entitled to 15 hours of free childcare per week (for 38 weeks of the year, or up to 570 hours annually), with the full 30 hours per week rollout starting from September 2025.

It said 15 hours, and described the 30-hour offer as “starting from September 2025”, as if it were still ahead of us. September 2025 was about ten months in the past. The current answer is 30 hours. What makes this the sharpest one is that Perplexity’s own sources contradicted it. It had cited a stack of real gov.uk pages, and at least two of them, read properly, say the thing it was denying, one of them describing the 30-hour tier for exactly this age group. The links were real. The model read them as if the calendar hadn’t moved.

The links were real. The model read them as if the calendar hadn't moved.

The childcare figure was one of Perplexity’s two slips on the board. The other came on the phone-fine question, where it led the £1,000 court fine with a solicitors’-firm marketing page and left gov.uk trailing beneath it. I come back to that one in the re-runs, because unlike the childcare figure it held.

Claude backed the right fine with a link that goes nowhere useful. Claude gave the correct court fines on the phone question, but sourced the £1,000 and £2,500 to a solicitor firm’s page, and cited gov.uk separately for the £200 fixed penalty. Same shape as Gemini’s miss, milder: right numbers, but the specific figure a reader most wants, the maximum fine, was pinned to a commercial blog when the primary page was open in the same answer. Click that link to check the £2,500 and it redirects you to the firm’s homepage, which carries none of the figures at all. Right number, and a citation that quietly leads nowhere.

I’m counting Gemini’s stamp-duty answer as its second slip, and it’s the gentlest of them, worth being precise about. Gemini gave the correct £5,000 for England and Northern Ireland, on the correct gov.uk page. Nothing wrong on the page it cited. What it did was present that as the answer to a question about “a £300,000 home” without mentioning that Scotland and Wales run different taxes at different rates, so a reader in Edinburgh or Cardiff would take an England figure off a real England page and never learn it didn’t apply to them. ChatGPT and Grok both caught that trap out loud, unprompted. The citation was correct. It just wasn’t complete.

The part that should worry you

0hedges, across every non-passing source answer
Not one “I’m not certain”, not one lower-confidence phrasing. The citations that failed arrived in the same calm tone as the ones that held.

After I’d graded the board, I went back through every non-passing source answer looking for a single hedge, any “I’m not certain”, any lower-confidence wording, and found nothing. Every one of them arrived as plainly as the correct answers did. That’s the finding underneath the finding: no visible seam between the citations that held and the ones that didn’t. The opaque Police.uk label looked as reassuring as an inspectable citation. The stale childcare figure was stated as plainly as a true one.

That’s what makes weak provenance more dangerous than a broken link. The broken link is honest about being broken. An opaque label or unsupported page comes dressed identically to a real source, in the same calm finished tone the assistant uses when it’s right, and nothing in the answer tells you which is which. The weakness is invisible until you require the exact page and inspect it.

For contrast, look at what the clean answers did. ChatGPT opened its stamp-duty answer with “Assuming you mean England or Northern Ireland” and named the Scotland and Wales divergence off its own bat. Grok cited one correct gov.uk page six times out of six and did the same. When these tools handle a source well, they often show their working, they flag the assumption, they name the boundary. The absence of that signal doesn’t prove a problem. Its presence is a decent sign of a good one.

I re-ran them, because one answer is not a rate

One run is only a story, so I put every question back through all five twice more, the same day, three rounds in total. A few things came out of it, and every one is the point of the exercise.

The weak-provenance habit held, and more plainly than I’d expected. Asked again for the phone-driving fine, Gemini put an opaque Police.uk label beside the £2,500 lorry-and-bus figure in the second round, then an opaque RAC label beside it in the third. The correct GOV.UK guide appeared in every round, but only as a detached closing source. The specific inline destinations are unrecoverable, so this proves a repeated auditability failure, not three proved wrong pages.

The phone fine wasn’t the only place that reflex showed. Across all three rounds, Gemini put a Notting Hill Genesis label beside the State Pension transition bands. In two of three, it put an Office for Statistics Regulation label beside the Welsh 20mph effective-date or definition claim. Those organisations do not set pension or road law, but the captures preserve labels rather than recoverable destination URLs. The evidence therefore shows repeated opaque, wrong-remit provenance, not exact pages I opened. The board caught two Gemini slips on the first pass; the closer look found the pattern was wider than that.

Perplexity ran both ways at once. Its court-fine sourcing held: across all three rounds it led the £1,000 court fine with a solicitors’-firm marketing page over gov.uk, citing that page four times in one round. The stale childcare figure did the opposite. On both re-runs, Perplexity corrected itself and gave the right 30 hours, from the same pages it had misread the first time. That self-correction is the more unsettling half, honestly. The same question, on the same day, handed me the wrong answer once and the right one the next, in the identical confident voice both times. You cannot lean on it being wrong any more than you can lean on it being right.

FIRST RUN
15 hours, with the 30-hour offer described as still to come “from September 2025”. That date was ten months in the past.
BOTH RE-RUNS, SAME DAY
30 hours, which is right, read off the same gov.uk pages it had misread the first time.
Nothing changed but the run. The wrong answer and the right one arrived in the same steady voice, which is why catching this by ear is not a plan.

So I am not claiming a percentage, and you should be wary of anyone who does off a handful of runs. What I have is a dated record of what these tools did, and two same-day re-runs of whether it held, the run that sets the sourcing tier in the State of AI Reliability report. The sourcing misses mostly held; the stale figure went the other way the second time. Both point at the same move: open the source yourself.

Update, 26 July 2026; wording corrected 23 August: the result above remains the dated 7 July record. On 23 August I narrowed the wording around Gemini’s opaque labels without changing any count or preserved receipt. Two days after Anthropic made Opus 5 the default on Claude’s Max tier, I re-ran all six questions on Claude alone, three fresh chats each, memory off. Claude’s slip above did not repeat: all three runs pinned the £1,000 and £2,500 court figures to gov.uk itself, and one run named the “unlimited fine” claim some commercial pages carry, traced it to the 2015 change in magistrates’ fine limits, and sided with gov.uk. The counterweight, because there is always one: a new one-run slip on the Welsh 20mph question, a precise-looking SI number that belongs to a different law entirely. Claude only, three runs per question, so read it as a dated re-check of one assistant, not a fresh board. Both entries are in the evidence log.

What this changed for me

I already open links on answers that matter. This tightened the rule: for anything I’m about to act on, a fine, a rule, a rate, a date, I first require an exact page, then check whether it actually says the thing. A source label with no destination fails at step one; a real page without the claim fails at step two.

THE TEN SECONDS, IN ORDER
Does the link open? Every one I could click did.
Does the label sound official? Police.uk does.
Is there an exact page URL you can inspect?
Is the page even about the thing you asked?
Is the exact figure on it, in words you can point at?
Does it apply to your nation and circumstances?
The top two are the checks people usually run, and neither caught a miss here. The exact-page check catches opaque labels; the final three test whether an inspectable source supports the claim and applies to the reader.

It’s the same ten-second habit I’ve written about when checking whether an AI’s sources are real at all: open the page, find the specific claim, ask whether the page is even about the thing you asked. This test just widened it from one tool to five and made the traps harder, and the habit held up as the only thing that reliably worked. If you’re picking a tool to lean on for this kind of everyday research, it’s worth knowing that free and paid tiers can behave differently on exactly these questions, which is part of why I keep a running audit of the free AI tools for this sort of research, since one blanket recommendation can’t be trusted across them. These five slips join the rest of the record in the running log of AI mistakes I catch.

The short version

What worked: ChatGPT and Grok cited a source that backed the claim on all six questions, and both flagged a jurisdiction trap without being asked. When these tools source well, they tend to show the seam, naming the assumption or the boundary.

What didn’t: Three assistants failed the strict source rubric, and all of them did it with zero hedge, in the same tone as their correct answers. Gemini put opaque labels beside the driving-fine figures while the inspectable GOV.UK guide sat detached below; Perplexity served a stale childcare figure contradicted by its own other sources.

Bottom line: Conditional, and the condition is you. On these six trap-laden questions, the weak receipt was invisible from the answer and only fell apart when I required an exact page and checked it, which took about ten seconds. What would change the verdict: if these tools reliably linked to the exact page that holds the claim, that ten seconds would stop earning its place. On 7 July 2026 it very much still did.

One more time on the honest limit, because it’s important. The board was one run per question, six questions, one day, with two more rounds the same day to see what held. It tells you who slipped and who held on the traps I set, with the dated evidence, and nothing more. It is not “Gemini gets sources wrong a third of the time”, and I’d ask you not to read it that way. A single non-passing receipt is enough to matter, though, because you weren’t going to question the one that looked fine. Gemini’s driving-fine answer is the cleanest proof of that interface trap: right number, plausible labels, no inspectable inline page. The neighbouring failure is worse and looks identical from the outside, because there the number itself was never real: a share price ChatGPT quoted as live that the stock never traded at, citation and all.

The check itself takes ten seconds per source, and I’ve written it up as a reusable method: how to check if a ChatGPT citation is real or fake.

Common questions

Do AI assistants cite the wrong source?
Sometimes, and the dangerous form is subtle. On six everyday UK questions put to five assistants on 7 July 2026, three failed the strict source rubric. Some linked a real page that did not back the claim; Gemini also placed unresolvable source labels beside a correct driving fine. A dead link is easy to catch. A plausible page or opaque label is not.
Which AI assistants got the sources right?
In this snapshot, ChatGPT and Grok cited a correct source on all six questions and flagged a jurisdiction trap unprompted. Gemini slipped the most, Perplexity on two, Claude on one. The board is one run per question; I re-ran them all twice more the same day (the source misses mostly held across all three rounds, one figure corrected itself), so this is a dated snapshot of who held on the day, not a reliability rate.
Why is a working AI link more dangerous than a broken one?
A broken link gives itself away, so you go and check. A working link that opens a real, official-looking page passes your sniff test, so you stop checking, and that is exactly where a citation that doesn't back the claim survives.
How do you check whether an AI's cited source actually backs the claim?
First require an exact page you can inspect. Then open it and look for the specific claim, not just whether the page loads. Ask whether it covers the thing you asked, carries the exact figure and applies to your situation. It takes about ten seconds.
Ben Dixon
// Written by Ben Dixon

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →

// Keep reading
AI Tests

An AI told me three weeks was more than thirty days

A kettle died after three weeks. UK law gives you 30 days to demand a refund. Four assistants said yes. One opened by telling me I'd missed the window.

AI Tests

Can AI build a game? Four tried, and built one nobody can win

Can AI build a game? I gave ChatGPT, Claude, Gemini and Grok the same Snake brief. Twelve games, one nobody can win, and four dead Start buttons.

AI Tests

Do AI model upgrades fix mistakes? It fixed mine, then made a worse one

Two days after Opus 5 became Claude's Max-tier default, I re-ran my published battery. The documented mistake vanished. A new one appeared, better dressed.

// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in AI Tests →