01 Guardrails · 2 Oct 2026
Do AI Citations Support the Answer? A Pilot Found Most Were Pinned to Nothing
A pilot: six assistants, six everyday questions.
Read 8 min
The AI Reliability Tracker
Every month, six AI apps sit the same exam, and we mark every answer.
Run 1 Sat 3 Oct 2026
Two aced it. One managed 8 of 30.
One AI story worth retelling, the proof behind it, and a check you can pinch.
Free. Leave whenever you like.
01 What we asked
Asked 6 apps3 tries at each
Ten questions, three tries at each. One dot per try.
Most right first, ties A to Z.
02Every try
One dot per try. Tap any group of dots to read what that app said.
Claude paid: Max 30 of 30
Copilot 30 of 30
Gemini free 27 of 30
Perplexity paid: Pro 27 of 30
ChatGPT logged out 22 of 30
Grok free 8 of 30
Six apps, ten questions, three tries at each: 180 tries in all. 18 got no answer, so each score counts all of an app's tries, not only the answers it gave. A try with no answer is named as one, never called wrong.
A right answer can still slip on a side fact. It counts as right, because the question asked something else, and the slip is named in the app's row. 162 of the 162 answers were checked for slips like this.
03 How we test
Latest writing
01 Guardrails · 2 Oct 2026
A pilot: six assistants, six everyday questions.
Read 8 min
02 AI Stories · 29 Sep 2026
In OpenAI's simulation, a safer AI widened an hourly helper's permissions and disabled approvals.
Read 6 min
03 AI Tests · 16 Sep 2026
I gave six AI assistants the same checkable questions.
Read 6 min
04 AI Stories · 8 Sep 2026
AI helped reveal an ancient argument inside a burnt Roman scroll.
Read 4 min
Ben's free letter: one striking AI story, the evidence behind it, and a check you can use.
Free. Leave whenever you like.