The Bluff Filter is this kind of check, on one page. Take it with you →
// On this page
Copilot got the bones question right. Asked how many bones are in a typical adult human skeleton, it said 206, added that babies start with about 270 to 300, and split the adult count into 80 and 126. Six tidy lines, most with a grey chip beside them.
Those chips lead to three pages, listed in the Sources panel under Citations. Below them sits a second heading, More, with twenty entries tied to no sentence at all. Three Pinterest pins. A Shutterstock search for skeleton stock photos. Britannica on Supination. And, second from the top:
I can’t think of a skeleton that needs the colour of the sun. The entry isn’t wrong, exactly. It just isn’t pinned to anything, and a source pinned to nothing can’t back a sentence. So the check this pilot points to is small: look for the sentence first, then the source.
This is a small pilot, and here’s how small
I settled the shape of this pilot on 19 September: six ordinary questions, six assistants, one run of each. The questions ran from bones and the Eiffel Tower’s height to why Einstein failed maths at school. He didn’t, which was the point of that one.
The runs happened on 25 and 26 September 2026, each in the assistant’s own temporary, incognito or private chat. ChatGPT, Claude and Perplexity were on signed-in Pro accounts. Gemini, Grok and Copilot were signed in. Grok and Copilot showed no paid plan, and Gemini’s plan wasn’t checked. AI agents did the asking: Codex captured 31 runs, and Builder, the Claude agent that builds this site, captured four of Grok’s after I’d signed in to Grok for it. Grok’s sixth answer never arrived after two retries, so there are 35 answers.
Most of what they showed was pinned to nothing
Twenty of the 35 answers showed a source of any kind. Four of those sources were chips with no link to follow, all of them pinned to a sentence.
// sources shown, per assistant, alphabetical 156 of 212 sources sat only in a list, pinned to no sentence.
pinned to a sentence only in a list
25–26 Sep 2026 · bars drawn to scale · each source counted once per answer. Counted as entries on screen it’s 281 of 378, largely because Grok’s panel lists many pages twice.
A listed page may be excellent. Nothing tells you which sentence it’s meant to back, so there’s no claim to check it against.
ChatGPT and Claude went the other way. Each showed sources in one answer out of six, both times about the Eiffel Tower. No prompt asked for links, though one asked where the euro sign is defined. That’s behaviour on six questions, not a flaw.
The pinned ones mostly held
A pinned source is one you can check, so AI graders held each one against the exact sentence beside it.
A “partly” could be a number the page put a little differently, like Copilot’s newborns. The answers held up too: of the 34 we could settle, 30 were fully right, 4 partly, and none was wrong. This isn’t a story about invented facts.
Two that missed
Two misses are worth keeping, because the pin tells you something.
What Grok said, Eiffel Tower question, 26 September 2026
A new digital radio antenna was installed around 2012, adding roughly 6 metres (nearly 20 feet) to the original 312 m structure at that time…
The chip beside that sentence opens the tower operator’s own article, dated 31 March 2022, saying the tower had just gained 6 metres. Grok cited the right page and still came out a decade early.
The screen, with the source chip at the end
Grok, Fast, Private Chat, signed-in account. Captured 26 September 2026.
The same answer says its figure is “confirmed directly on the official Eiffel Tower website”. For the 330 metres, it is. The date needed one more look.
Perplexity’s miss is quieter. Its line “It opened on November 4, 1861” carried one chip holding two pages. The first, the University of Washington’s About page, says only “Since our founding in 1861”, and the graders marked that pair unsupported. The second, a surgery department history page, says “founded on November 4, 1861”: the same day, but founding isn’t opening, so that pair came out partly. The date is right. Neither page quite says it.
Three checks before you trust a source
-
Find the pin. Which sentence is this source attached to? A list at the bottom is pinned to nothing.
-
Find the words. Open it and search the page for the number, name or date. Right topic isn’t the same as saying it.
-
Check it’s the same fact. Match the date and the detail. Grok’s own source said 2022, not 2012.
These steps come from the citation check we grade with, and this pilot didn’t test whether they help readers catch more. What they will do is stop you trusting a page for a sentence it never mentions. If the link itself looks shaky, check it’s real first, and the Source Ladder helps once you know what the page is.
What this pilot can’t tell you
A pilot, 25–26 September 2026. Not a ranking.
Claude and OpenAI’s GPT-6 Sol graded without seeing each other’s work, and the 43 of 234 pairs they still split on after a second round are out of every count above. Both tripwires the protocol set in advance stayed quiet.
How the grading worked
Claude was Grader A and GPT-6 Sol was Grader B. Both checked 185 of the 191 settled pairs. By design Grader B reviewed 30 of the 35 answers, so the other six settled pairs rest on Grader A alone.
The first tripwire: on 29 pairs from a random draw of five answers, the graders split on support twice, under the one-in-five limit. The second: 17 of the 191 settled pairs couldn’t be verified, under the limit of a quarter.
Several sources in one answer can share one cause, like a single search, so they aren’t independent votes.
For each factual claim in your answer, put its source straight after that sentence, not in a list at the end. Quote the exact words on that page that say it, with the date if there is one. If no page you retrieved says it, mark the claim “unsourced” rather than attaching a nearby link.
This pilot didn’t test the prompt. It asks for pins you can check, and you still have to open one.
The short version
What worked: Pinned sources mostly backed the sentence beside them.
What didn’t: Most sources on show were pinned to nothing, and a few pinned pages didn’t say what their sentence said.
Bottom line: A pilot, not a ranking. Check the pins. The pile underneath has nothing to be checked against.
Next time an answer arrives with twenty sources, count the pins. The chips beside the sentences are the sources doing the work. The rest is reading material, and some of it is about the sun.
Ben tests AI answers and follows the sources behind AI stories. His work includes original experiments, reported stories and practical checks, with the evidence and limits made clear. About Ben →
Enjoyed this? Get the next Confidently Wrong letter →
The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.
