Skip to content
DIXON.AI The AI Reliability Tracker
Menu
Guardrails

Do AI Citations Support the Answer? A Pilot Found Most Were Pinned to Nothing

A pilot: six assistants, six everyday questions. Most sources they showed sat in a list pinned to no sentence. The pinned ones mostly held up.

The Bluff Filter is this kind of check, on one page. Take it with you →

// On this page

Copilot got the bones question right. Asked how many bones are in a typical adult human skeleton, it said 206, added that babies start with about 270 to 300, and split the adult count into 80 and 126. Six tidy lines, most with a grey chip beside them.

Those chips lead to three pages, listed in the Sources panel under Citations. Below them sits a second heading, More, with twenty entries tied to no sentence at all. Three Pinterest pins. A Shutterstock search for skeleton stock photos. Britannica on Supination. And, second from the top:

Copilot's Sources panel for a question about how many bones are in an adult skeleton: three pinned pages under Citations, then a More list whose second entry, boxed in orange, is titled What Color Is The Sun.
Copilot, Temporary chat, 25 Sep 2026. Three pinned pages, then the More list. The boxed row is not about bones, and it is pinned to no sentence.

I can’t think of a skeleton that needs the colour of the sun. The entry isn’t wrong, exactly. It just isn’t pinned to anything, and a source pinned to nothing can’t back a sentence. So the check this pilot points to is small: look for the sentence first, then the source.

This is a small pilot, and here’s how small

I settled the shape of this pilot on 19 September: six ordinary questions, six assistants, one run of each. The questions ran from bones and the Eiffel Tower’s height to why Einstein failed maths at school. He didn’t, which was the point of that one.

The runs happened on 25 and 26 September 2026, each in the assistant’s own temporary, incognito or private chat. ChatGPT, Claude and Perplexity were on signed-in Pro accounts. Gemini, Grok and Copilot were signed in. Grok and Copilot showed no paid plan, and Gemini’s plan wasn’t checked. AI agents did the asking: Codex captured 31 runs, and Builder, the Claude agent that builds this site, captured four of Grok’s after I’d signed in to Grok for it. Grok’s sixth answer never arrived after two retries, so there are 35 answers.

Most of what they showed was pinned to nothing

Twenty of the 35 answers showed a source of any kind. Four of those sources were chips with no link to follow, all of them pinned to a sentence.

A listed page may be excellent. Nothing tells you which sentence it’s meant to back, so there’s no claim to check it against.

ChatGPT and Claude went the other way. Each showed sources in one answer out of six, both times about the Eiffel Tower. No prompt asked for links, though one asked where the euro sign is defined. That’s behaviour on six questions, not a flaw.

Claude's answer giving the Moon's average distance as about 384,400 km, with the row under it boxed in orange and labelled no links: copy, read-aloud and retry buttons, and no source chip.
Claude, incognito, 25 Sep 2026. The Moon question: right answer, nothing to click.

The pinned ones mostly held

A pinned source is one you can check, so AI graders held each one against the exact sentence beside it.

Supported131
Partly30
Unsupported11
Contradicted2
Can’t verify17
Pinned source against its sentence, out of 191 settled pairs, drawn to scale. Disputed pairs are left out.

A “partly” could be a number the page put a little differently, like Copilot’s newborns. The answers held up too: of the 34 we could settle, 30 were fully right, 4 partly, and none was wrong. This isn’t a story about invented facts.

Two that missed

Two misses are worth keeping, because the pin tells you something.

What Grok said, Eiffel Tower question, 26 September 2026

A new digital radio antenna was installed around 2012, adding roughly 6 metres (nearly 20 feet) to the original 312 m structure at that time…

2012Grok’s date
2022its own source

The chip beside that sentence opens the tower operator’s own article, dated 31 March 2022, saying the tower had just gained 6 metres. Grok cited the right page and still came out a decade early.

The screen, with the source chip at the end

Grok's answer paragraph with 'A new digital radio antenna was installed around 2012' boxed in orange, ending in a Toureiffel source chip.

Grok, Fast, Private Chat, signed-in account. Captured 26 September 2026.

The same answer says its figure is “confirmed directly on the official Eiffel Tower website”. For the 330 metres, it is. The date needed one more look.

Perplexity’s miss is quieter. Its line “It opened on November 4, 1861” carried one chip holding two pages. The first, the University of Washington’s About page, says only “Since our founding in 1861”, and the graders marked that pair unsupported. The second, a surgery department history page, says “founded on November 4, 1861”: the same day, but founding isn’t opening, so that pair came out partly. The date is right. Neither page quite says it.

Perplexity's answer with 'It opened on November 4, 1861' boxed in orange. The open source chip, marked Trusted, shows the About the UW page, which says only 'Since our founding in 1861'.
Perplexity, incognito, 25 Sep 2026. The first page behind the chip gives the year, not the date.

Three checks before you trust a source

  1. Find the pin. Which sentence is this source attached to? A list at the bottom is pinned to nothing.
  2. Find the words. Open it and search the page for the number, name or date. Right topic isn’t the same as saying it.
  3. Check it’s the same fact. Match the date and the detail. Grok’s own source said 2022, not 2012.

These steps come from the citation check we grade with, and this pilot didn’t test whether they help readers catch more. What they will do is stop you trusting a page for a sentence it never mentions. If the link itself looks shaky, check it’s real first, and the Source Ladder helps once you know what the page is.

What this pilot can’t tell you

What one run of six questions can and can’t show
✓Where each source sat ✓Whether pinned pages said it ✗Which assistant is best ✗Whether listed pages were good ✗How it looks next month

A pilot, 25–26 September 2026. Not a ranking.

Claude and OpenAI’s GPT-6 Sol graded without seeing each other’s work, and the 43 of 234 pairs they still split on after a second round are out of every count above. Both tripwires the protocol set in advance stayed quiet.

How the grading worked

Claude was Grader A and GPT-6 Sol was Grader B. Both checked 185 of the 191 settled pairs. By design Grader B reviewed 30 of the 35 answers, so the other six settled pairs rest on Grader A alone.

The first tripwire: on 29 pairs from a random draw of five answers, the graders split on support twice, under the one-in-five limit. The second: 17 of the 191 settled pairs couldn’t be verified, under the limit of a quarter.

Several sources in one answer can share one cause, like a single search, so they aren’t independent votes.

// The three checks, as one prompt

For each factual claim in your answer, put its source straight after that sentence, not in a list at the end. Quote the exact words on that page that say it, with the date if there is one. If no page you retrieved says it, mark the claim “unsourced” rather than attaching a nearby link.

This pilot didn’t test the prompt. It asks for pins you can check, and you still have to open one.

The 49-second version. Synthetic voice.

The short version

What worked: Pinned sources mostly backed the sentence beside them.

What didn’t: Most sources on show were pinned to nothing, and a few pinned pages didn’t say what their sentence said.

Bottom line: A pilot, not a ranking. Check the pins. The pile underneath has nothing to be checked against.

Next time an answer arrives with twenty sources, count the pins. The chips beside the sentences are the sources doing the work. The rest is reading material, and some of it is about the sun.

Ben Dixon
// Written by Ben Dixon

Ben tests AI answers and follows the sources behind AI stories. His work includes original experiments, reported stories and practical checks, with the evidence and limits made clear. About Ben →

// Keep reading
Guardrails

How to check if ChatGPT cites your site

Normal analytics do not show what ChatGPT says about your site. Here's my frozen question set and the round where it described another company.

Guardrails

Why does ChatGPT make up sources? Two gov.uk links, only one held the rule

With web search on, ChatGPT's links are real. The failure is a working link to a page that doesn't hold the claim. Here's why it happens.

Guardrails

Does ChatGPT just agree with you? Mostly no, but watch the numbers

Does ChatGPT just agree with you? Mostly no. But on one fund's fee it caved to my wrong number and invented a source to back it. The 30-second check.

A LETTER FROM BEN

Confidently Wrong.

One striking AI story, the evidence behind it, and a check you can use.

Read a real edition first ↗

Enjoyed this? Get the next Confidently Wrong letter →

// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in Guardrails →