Skip to content
AI Tests

Does a prompt to stop AI hallucinations work? I graded mine, and one answer got worse

I graded my own anti-bluffing prompt, three runs a cell. It defused every flat bluff on the citation and maths traps, then made one citation answer worse.

This is the kind of thing the Bluff Filter catches. It’s free →

// On this page

Asked which study found that people check their phones 150 times a day, Perplexity handed me a citation: Wilcockson, Ellis and Shaw, 2018, in Cyberpsychology, Behavior, and Social Networking. Three real researchers. A real journal. A real paper, which has never contained that figure.

The interesting part is which half of the test that came from. It came from the run carrying my own prompt to stop AI hallucinations, and it arrived with a label attached. It happened twice: one of those runs called the claim “inferred”, the other gave it a “medium confidence” verdict.

The bluff had not gone anywhere. It had filled in a form.

Minutes either side of it, in the same session, the same question asked plain got it right. No journal study. An industry report. That same 2018 paper mentioned only as adjacent research, which is all it is.

The prompt is the Bluff Filter, the free paste-in instruction set this site gives away, so grading it against a primary source the way I grade the assistants was overdue. Three runs a cell, both ways, one day.

Scope, before any of the numbers

I ran it on three assistants: ChatGPT on the free tier, Gemini Flash, and Perplexity on a logged-in free plan showing a Pro preview banner. Four questions, all traps I picked because they have a checkable right answer: two citations that do not exist (the study behind “people check their phones 150 times a day”, and the one behind “it takes 21 days to form a habit”), a maths question that looks like a multiplication, and a two-turn ask for the exact source URL behind a statistic. Every question ran both ways, plain and filtered, three times each. Within a round the two arms ran back to back, minutes apart, so neither ever got a better day than the other. The rounds themselves were not all one sitting: round one was captured on 24 July 2026 and rounds two and three the following night, each in a fresh thread. Every answer was graded against a primary source.

The board is uneven, and the table below shows it. All three assistants ran the URL trap and the habit-study trap, three times each way, which is the six in their rows. The phone-checking citation and the brownie maths went to Perplexity alone, after an earlier pass had found the other two clean on both, which is why its row reads out of twelve.

These are counts on a small trap set, not a rate. Every result that separated the two arms came from Perplexity. ChatGPT and Gemini did not serve a single flat wrong answer in either arm, which puts a ceiling on what this test can show.

AssistantWrong, stated flat: plainWrong, stated flat: filtered
ChatGPTfree 0/6 0/6
GeminiFlash, temporary chat 0/6 0/6
Perplexityfree plan, Pro preview banner 4/12 2/12

Both filtered survivors are the same class: the source URL. Look only at the traps where the bluff lives inside the answer itself, the citation and the maths, and the filter is clean. Six plain Perplexity runs produced three flat wrong answers. Six filtered ones produced none.

  • ChatGPTNothing wrong stated flat either way. In one filtered run it corrected its own source attribution mid-answer, unasked.
  • GeminiNothing wrong stated flat either way, and the most obedient of the three about the four-stage structure. Its links were the weak spot: one “Direct PDF Link” was a Google search wrapper.
  • PerplexityThe only assistant whose two arms differed at all. Halved its flat wrong answers with the filter on, and produced the one result that got worse.

The answer that got worse

The plain run is the one I would keep. Perplexity said the 150-checks-a-day figure is commonly attributed to Kleiner Perkins’ Internet Trends report, not to a peer-reviewed journal article, and that it could not verify a journal citation for the statistic. It then offered two real smartphone-usage papers, Wilcockson 2018 and Andrews 2015, as adjacent research. That is the correct answer on the point that mattered, and it holds up: the number comes from industry, not from any journal. One caveat I owe you, since this post is about attribution: “commonly attributed to Kleiner Perkins” is where the figure gets cited, not where it started. The trail runs back further, to a number Nokia put out around 2010 and a consultant who popularised it. Perplexity was right that it is industry rather than academia, and loose about the rest.

Perplexity, plain prompt, saying the 150 times a day figure is attributed to a KPCB Internet Trends report and not a journal study, with the closing line boxed in orange.
Perplexity, no filter, 24 July 2026. The boxed line is the one the filtered runs lost.

With the filter attached, that same paper moved. In two of three runs it stopped being nearby research and became the answer, the thing I had asked for, in one of them sitting under a heading that read Exact citation.

Perplexity, running the Bluff Filter, giving Wilcockson, Ellis and Shaw 2018 under an Exact citation heading, boxed in orange.
Perplexity, running the Bluff Filter, 25 July 2026. Round two of three. The same question asked plain, five minutes earlier in the same sitting, got it right.

Underneath it, the filter’s own labelling did its work exactly as instructed:

Px Perplexity said Labelled, and wrong

Sourced: the journal, year, and authors come from the paper reference shown in the source.

Inferred: this is the best match for the “150 times a day” statistic because the figure is linked to that study in later coverage.

Risk: if you need the original source for the exact 150-check claim, the secondary reporting may not be enough; the paper itself is the thing to verify against.

Nothing there breaks the rule I wrote. The filter says to separate what the model knows from what it is inferring, and to label every claim the answer depends on, and that is precisely what happened. The stage it walked past is the one before: SCOPE, which tells the model it is allowed, and required, to stop and refuse rather than guess. The same file names the failure to watch for as “stating a fabricated source, citation, figure, or quote in the same calm tone as a real one”. A guess is labellable. So the guess got stated.

The first run went the other way entirely, and it is the reason I ran the question three times at all. On that pass it was the plain arm that bluffed, naming Oulasvirta 2011 flat as the study, no hedge, when no journal paper reports that figure at all. The filtered arm on the same pass got it right, down to saying in as many words that it could not verify a peer-reviewed study with that number. So on this one question the filter caught the first bluff and then induced the next two.

Label your guesses turns out to be a door as well as a gate.

It is the difference between a stranger telling you they don’t know where the station is, and a stranger pointing confidently down the wrong road while mentioning they’re only about 70% sure. The second one is more informative and gets you more lost.

Where it does nothing at all

The URL trap is the failure that survived the filter, and I will not pretend to be surprised by it. A block of text pasted at the top of a chat cannot open a web page. Asked for the direct link behind a UK AI-usage figure, Perplexity wrote out a line beginning “Direct URL:” above an Office for National Statistics address that returns a 404, while its own citation chip in the same answer carried the working address for the same report. A dead ONS link served up as the direct page happened in two of the three filtered runs, each time with the filter’s own scope-and-risk scaffolding sitting on the page above it.

The honest read: the filter changes how an answer talks about its own certainty, and it has no reach into whether a link resolves. A 404 with a confidence level attached to it is still a 404. Checking whether a cited source exists stays a job you do yourself, in another tab.

Labelling is not correctness

The brownie question is the plainest illustration, and it is the class our own Bluff Filter page uses as its worked example. Scale a recipe from an 8-inch tin to a 16-inch one at the same batter depth, and the bake time barely moves, because heat travels through the depth and the depth has not changed. Perplexity got it wrong in all three filtered runs, and in two of the three plain ones. Between 43 and 50 minutes across the three, medium confidence every time, with an instruction to start checking at 40 or 45. The real answer is nearer 30 to 35, the figure I fixed in writing before the first capture and the one piece of ground truth here with no fetchable source behind it.

So the filter did not save the brownie. It stopped the number arriving as a fact. The worst plain run said 60 minutes, flat, no range, on the reasoning that four times the area roughly doubles the time. The filtered runs gave a wrong estimate, a confidence level, and an instruction to put a skewer in early, which is the behaviour that gets you a decent brownie out of a bad calculation.

That is the whole promise, and it is a smaller promise than most people reading the words "refuses to fabricate" will assume.

How I graded it

Both arms of every question ran on the same assistant, in the same session, minutes apart, in a fresh chat with memory off, so neither arm got a better day than the other. The filtered arm used the shipped Bluff Filter text, not a paraphrase: pasting it in unchanged was the pre-committed procedure, and the capture notes record the filter by name rather than storing its text. Ground truth was fixed in writing before the first capture and checked at grading: every cited paper looked up, every URL fetched for a status code and read for the figure it was supposed to carry. Thirty-eight new captures plus the pilot’s twenty-four, all screenshotted, all kept. The three-verdict scale is the same one the Scoreboard uses.

One caveat I cannot grade away: Perplexity’s tier. The first batch, on the afternoon of the 24th, logged the account as Pro; the batch the following night as “Free plan” with a Pro preview banner, so treat the tier as ambiguous. Both arms of every question ran in the same session, so the within-question comparison holds regardless of which model was serving.

The short version

What worked: On the traps where the bluff lives inside the answer, the citation and the maths, the filter went three flat wrong answers in six plain runs to none in six filtered runs, and it turned a flat 60-minute bake time into a hedged estimate with a check-it instruction.

What didn’t: It has no effect on fabricated source URLs, twice writing out a dead link as the direct source page. And on the citation trap it twice promoted a real, topically adjacent paper into the answer slot with a hedge attached, where two of the three plain runs had correctly said no such study exists.

Bottom line: Useful, with a narrower promise than the words suggest. It makes an unverified claim visible rather than making it go away, and on one trap the labelling gave a guess somewhere to live. What would change the verdict: more questions per trap class, and a rewritten SCOPE stage that makes “no such source exists” an answer the model has to reach for, rather than one it can walk past the moment it finds something plausible.

I would still paste it in. The alternative, on this evidence, is a model that gets the same things wrong without telling you which parts it made up. But the honesty pass belongs on the product page as much as in the results, and I would rather write that sentence myself than have a reader find it. The test’s key failures and catches are logged in the evidence register.

Common questions

Does a prompt stop AI hallucinations?
It reduces the confident kind and does not remove the underlying invention. On this trap set, a paste-in instruction to label every guess took flat wrong answers from three of six plain Perplexity runs to none of six filtered. The same instruction also produced two wrong citations, labelled as inferences rather than withheld.
Does telling an AI to flag its guesses make it more accurate?
No. It makes the guess visible, which is a different thing. In this test the filtered arm got the brownie question wrong all three times, and the plain arm got it wrong twice. The only change was that the wrong number arrived with a confidence level and an instruction to check the tin early.
Ben Dixon
// Written by Ben Dixon

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →

// Keep reading
AI Tests

Does ChatGPT make up stock prices? Yes, and I caught two that never traded

Does ChatGPT make up stock prices? Yes. Asked for a live price, it gave me two NVDA figures that never traded that day. Here's the 15-second check first.

AI Tests

Best AI for math: which is most reliable with numbers?

What's the best AI for math? On everyday sums the assistants are level. The real test is the number that looks like a sum and isn't, and who catches it.

AI Tests

Is ChatGPT good at maths? I graded four AIs on 120 real answers

I put four AI assistants through 120 graded everyday sums. Every final answer was right. The mistakes were sitting above the working.

// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in AI Tests →