Skip to content
// AI reliability

AI reliability: how AI answers fail, and how to check one.

Reliability changes with the kind of question you ask, which is why a single accuracy figure for an AI assistant tells you almost nothing you can act on. On facts already fixed in the public record, the six assistants graded here came back clean on 24 of 24 results. On live, moving data, the same six produced 3 confident errors: a wrong value served as reliable, sourced, formatted, and carrying no hedge at all. The habit that helps is asking what kind of question you have in front of you, and what would catch this particular answer if it were wrong.

Assistants graded
Six (ChatGPT, Claude, Gemini, Perplexity, Grok and Copilot)
Method
Each question asked 3 times, memory off, graded against a primary source fixed before the run
Latest run
18 July 2026
Grader of record
Ben Dixon
// What reliability means for a consumer assistant

A single accuracy score hides the split that actually decides whether you get burned.

Ask an assistant something whose answer was written down once and stayed written down, a published figure, a statutory right, the menu path in a piece of software, and it is close to dependable. Ask it something that moved this morning and the same model will very often answer anyway, in exactly the same confident register, because nothing about the question told it that this one needed a live feed.

Fixed public record 100%
Everyday questions 94%
Live, moving data 42%
Clean results by kind of question. 24 of 24 on facts fixed in the public record · 15 of 16 on the everyday set (scaling a recipe, a menu path in a spreadsheet, a refund right) · 5 of 12 on live, moving data. Six assistants, each question asked 3 times, graded against the primary source. Latest run 18 July 2026.

That is the whole shape of the problem, and it is why the question “how reliable is AI” has no honest answer until you say which kind of question you mean. Every question in that battery, with what each assistant said and the source it was graded against, is on the Scoreboard.

// How answers actually go wrong

6 kinds of wrong, taken from a log rather than from a list.

Each of these is a shape that kept recurring in real sessions, which is how it earned a name. The register currently holds 55 logged failures and 50 logged catches, every row dated, screenshotted where a screenshot exists, and with the assistant named.

Counts derive from the register at build time, so this page cannot drift from it. The register itself is one flat, filterable log of every answer we have checked: the evidence register, or straight to the failures only.

// How we grade

Three outcomes, because right-or-wrong cannot see the dangerous one.

Correct Right against the primary source it was graded on, on every run.
Partial Wrong or incomplete, and openly flagged as uncertain. Imperfect and honest, which the rubric treats as good behaviour rather than a failure.
Confidently wrong Wrong on a checkable fact, served as reliable, with no hedge and no named uncertainty. The only outcome graded hard, because it is the one that costs you.

With only right and wrong to mark with, an honest “I cannot pull that live” scores the same as a confident invention, so the bluff becomes the better play. The middle outcome exists to stop that: a Partial is earned only by an answer that is wrong and says so openly. The full reasoning, with the research behind it, is on how we grade.

// Where the assistants stand

The graded record as of 18 July 2026.

6 checkable questions, each put to six assistants 3 times over, in fresh chats with memory off. Listed with the fewest confident errors first. Several assistants are level, and the order inside a level is not a ranking.

Assistant Correct Confidently wrong
Claude Max · paid 5/6 none
Gemini Pro · paid 5/6 none
Grok Free · Grok 4.3 Fast 5/6 none
ChatGPT Free 5/6 none
Copilot Free · Smart 5/6 1 /6
Perplexity Pro · paid 4/6 2 /6

Reading it: teal is a clean full mark · pink is a wrong answer served as certain. High is good for Correct; none is the win for Confidently wrong.

Counted over the 6 core questions only, 36 results in total, where one result is one assistant’s verdict on one question across 3 runs. Every tier is labelled on the board: an assistant graded on a free tier is not being compared like for like with a paid one, and saying so is cheaper than pretending otherwise. Outright fabrications in this run, meaning a figure invented with no source at all: 0.

This is a snapshot of one dated battery. The versioned report, with a changelog recording every re-grade and correction we have made to our own numbers, is The State of AI Reliability. The board itself, with a page per question and the receipt behind each verdict, is the Scoreboard.

Free when you subscribe

The Bluff Filter

A paste-in prompt that makes any AI flag what it’s guessing before you act on it. Learn it once, use it on every answer.

Plus one email a fortnight, showing where an AI went wrong.