The Bluff Filter is this kind of check, on one page. Take it with you →
// On this page
On 20 June 2026 I asked ChatGPT a plain money question and asked it to show me the source. An ISA is the UK’s tax-free savings and investment account, and I wanted to know whether you can move part of this year’s money to a new provider or have to move the whole lot. It answered, roughly right, and handed me a gov.uk link.
The link worked. Real government page, current, live. It was about what happens to your ISA if you move abroad or die. Useful, if the question had been either of those.
I put the same question to Perplexity the same afternoon. It cited a different gov.uk page and quoted the sentence that carries the rule. Two links, both real, both gov.uk. Only one held the answer. So the thing that went wrong was never the link. It was the page underneath it.
The two ways a source goes wrong
There are two separate failures here, and most writing on the subject squashes them into one.
Fabrication is the famous one. The model invents a reference: a study nobody wrote, a link that goes nowhere. A dead link gives itself away on the first click, which makes it the easy one. A 2023 study in Scientific Reports went through 636 citations ChatGPT produced across 84 papers it had been asked to write, and found 55% of the older model’s references fabricated, and 18% of GPT-4’s. Those runs were April 2023, though, and neither model has been available in ChatGPT for over a year.
Misattribution is the other one, and it has no reputation at all. The source exists. The page loads, the domain is often as official as it gets, and the claim still isn’t on it. That is the ISA case above.
Turning on web search takes away the failure you can catch and leaves you the one you can’t, which is one of the stranger things about search-grounded answers.
If you’d rather catch this before you act on an answer, the Bluff Filter is the one-page set of instructions I paste in so a model marks which parts of its answer it looked up and which parts it is guessing at.
Why ChatGPT makes up sources
Three things are going on, and they stack.
The model writes the citation. With no search running, it isn’t fetching a reference and copying it out. It is producing the most likely next words, and a source is just more words. It knows the shape of a citation for a claim like yours: a plausible journal, a plausible year, a URL that looks like one that would exist. The shape is all it sets out to produce. Nothing in that process goes and checks that the source is there.
With search on, the links are real. The join is still written. The search step fetches pages about your topic, so the links come back genuine. The model then decides which page gets attached to which sentence, and it makes that decision the same way it made the sentence. Nothing checks afterwards that the page it picked contains the claim it pinned on it.
That gap is the whole of misattribution, and it is the part none of the explanations I read covered.
And it guesses rather than saying it doesn’t know. OpenAI’s own explanation, published in September 2025, is blunter than the usual one: “language models hallucinate because standard training and evaluation procedures reward guessing over acknowledging uncertainty.” A model that says “I’m not sure” scores worse on the tests it’s measured on than one that has a go. So it has a go. Anyone who has sat a multiple-choice paper with no penalty for a wrong answer knows the logic.
What I caught, and what I didn’t
Here is the ISA answer in ChatGPT’s own words, and Perplexity’s. The whole run is in the ChatGPT sources test.
Under the section “What amounts can be transferred”, HMRC says: “transfer some or all of the current year’s ISA subscriptions, the investments bought with those subscriptions, and any income arising on those investments” […]
the only gov.uk address anywhere in that answer was /individual-savings-accounts/if-you-move-abroad-or-die, which carries none of this
The clearest rule from the government is: “You can transfer all or part of the savings in your Individual Savings Account (ISA) from one provider to another at any time.”
ChatGPT quoted HMRC accurately, then hung the quote on a page that carries nothing of the kind. The transfer rule lives on /transferring-your-isa, which is the page Perplexity opened and quoted back. When I checked again in late July, gov.uk had split that part of the guide into two separate pages, and the address ChatGPT gave me redirected to the ISA overview. Even the wrong page didn’t stay put. Catches: a real page, on the right domain, standing in for the page that holds the rule.
Earlier in the same session I’d asked ChatGPT the annual fee on a popular global index fund, the Vanguard FTSE All-World ETF. It said 0.19% a year, cited four links on Vanguard’s own domains, and noted the fee used to be 0.22% before Vanguard cut it. I opened every one. Real Vanguard pages, right number, right fund. Catches: nothing, and that is the point. Same tool, same session, sourcing exactly right.
The shape isn’t ChatGPT’s alone, either. On 7 July 2026 Gemini gave me the correct £2,500 fine for using a handheld phone as a lorry driver and sourced it to Police.uk, the crime-map site.
What the usual advice misses
Search the title. Resolve the DOI (an academic paper’s permanent link). Check the publisher. Confirm the page loads. Every fix I found when I went looking asks whether the source exists, and a live gov.uk page passes all of them without breaking stride.
The research has the same shape. The Scientific Reports study graded every citation on its bibliographic parts, authors, title, journal, volume, pages, and a second study, on 115 medical references, did the same. Neither asked whether a real source supported the claim it had been attached to, and the second says as much in its own limitations. OpenAI’s help page lists “fabricated quotes, studies, citations or references to non-existent sources” among its hallucination examples, and stops there.
None of that is a broken promise, which is worth saying plainly. ChatGPT search offers “fast, timely answers with links to relevant web sources”.
Relevant is what's promised. The claim being on the page is what you need. The ISA answer sat in the gap between those two words.
The check, in one line
Open the page, and look for the specific claim on it.
That’s the whole thing. It’s what I did on 20 June, and it’s the only reason this post exists rather than a wrongly-sourced ISA answer sitting in my notes with a government pill next to it, looking checked. The full version, four steps and about ten seconds, is how to check if a ChatGPT citation is real or fake. Card 01 on the guardrails hub is the paste-in that gets a model to say whether it retrieved a source or recalled it. And when an answer arrives with a stack of citations, open them all before you read any of them. The dead and the invented fall out in one pass; whether a real page actually backs a specific claim stays yours.
The short version
What worked: Opening the cited page caught it in seconds. On the day, the gov.uk link behind the ISA answer opened a real government page about a different subject, and nothing in the wording gave that away.
What didn’t: No existence check would have found it. Search the title, follow the link, look at the domain: a working gov.uk page on the wrong subject clears all three.
Bottom line: With web search on, the links ChatGPT hands you are real. Nothing then checks whether the page it picked says the thing it said, because the join between a claim and its link is written the same way the sentence was. The model has no more idea than you do. What would change that: a sourced answer that reliably points at the page carrying the claim. On what I’ve run, we aren’t there.
One limit worth stating. This is two days of testing: two questions on the first, six on the second. The ISA test was 20 June 2026, one run per tool, ChatGPT on the free tier with web search on, in a temporary chat with memory off, which is how a first-time user meets it. The five-assistant board on 7 July was one run per question, then the whole set twice more the same day, three rounds in all, and Gemini put that phone fine on a source other than gov.uk in every round. That’s enough to show the failure is real and what it looks like. It is nowhere near enough to say how often it happens, and any rate quoted off a few runs is doing more work than it can carry. Which assistants hold up on sourcing over a longer run of questions is what the Scoreboard is for. The ones I’ve caught so far are in the record of AI mistakes.
Common questions
- Why does ChatGPT make up sources?
- Two reasons. Without web search it writes what a source for a claim like that usually looks like, rather than fetching a real one. With search on the links are real, but the model still decides which page goes with which sentence, and nothing in the pipeline checks that pairing.
- How do I stop ChatGPT making up sources?
- You can reduce it by asking the model to name its source and say whether it retrieved the page or recalled it, but you can't remove it. The check stays with you: open the cited page and find the exact claim on it. That takes about ten seconds.
- What's the difference between a fabricated source and a misattributed one?
- A fabricated source doesn't exist: the paper was never written, or the link is dead, and one click catches it. A misattributed source is real and loads, often on an authoritative site, and simply doesn't contain the claim it was cited for. The second survives every existence check.

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →
The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.