AI reliability: how AI answers fail, and how to check one.
Reliability changes with the kind of question you ask, which is why a single accuracy figure for an AI assistant tells you almost nothing you can act on. On facts already fixed in the public record, the six assistants graded here came back clean on 24 of 24 results. On live, moving data, the same six produced 3 confident errors: a wrong value served as reliable, sourced, formatted, and carrying no hedge at all. The habit that helps is asking what kind of question you have in front of you, and what would catch this particular answer if it were wrong.
A single accuracy score hides the split that actually decides whether you get burned.
Ask an assistant something whose answer was written down once and stayed written down, a published figure, a statutory right, the menu path in a piece of software, and it is close to dependable. Ask it something that moved this morning and the same model will very often answer anyway, in exactly the same confident register, because nothing about the question told it that this one needed a live feed.
That is the whole shape of the problem, and it is why the question “how reliable is AI” has no honest answer until you say which kind of question you mean. Every question in that battery, with what each assistant said and the source it was graded against, is on the Scoreboard.
6 kinds of wrong, taken from a log rather than from a list.
Each of these is a shape that kept recurring in real sessions, which is how it earned a name. The register currently holds 55 logged failures and 50 logged catches, every row dated, screenshotted where a screenshot exists, and with the assistant named.
Counts derive from the register at build time, so this page cannot drift from it. The register itself is one flat, filterable log of every wrong answer and good catch from our published tests: the evidence register, or straight to the failures only.
- Nine kinds of AI failure, each named from a dated test, with the check that catches it
- A running tally of how often one assistant got things wrong, kept across dated tests
- What turning web search on fixes, and what it quietly makes worse
- A right answer arriving with the wrong source bolted on
- The thirty second version of the checking habit
Three outcomes, because right-or-wrong cannot see the dangerous one.
With only right and wrong to mark with, an honest “I cannot pull that live” scores the same as a confident invention, so the bluff becomes the better play. The middle outcome exists to stop that: a Partial is earned only by an answer that is wrong and says so openly. The full reasoning, with the research behind it, is on how we grade.
The graded record as of 18 July 2026.
6 checkable questions, each put to six assistants 3 times over, in fresh chats with memory off. Listed with the fewest confident errors first. Several assistants are level, and the order inside a level is not a ranking.
| Assistant | Correct | Confidently wrong |
|---|---|---|
| Claude Max · paid | 5/6 | none |
| Gemini Pro · paid | 5/6 | none |
| Grok Free · Grok 4.3 Fast | 5/6 | none |
| ChatGPT Free | 5/6 | none |
| Copilot Free · Smart | 5/6 | 1 /6 |
| Perplexity Pro · paid | 4/6 | 2 /6 |
Reading it: teal is a clean full mark · pink is a wrong answer served as certain. High is good for Correct; none is the win for Confidently wrong.
Counted over the 6 core questions only, 36 results in total, where one result is one assistant’s verdict on one question across 3 runs. Every tier is labelled on the board: an assistant graded on a free tier is not being compared like for like with a paid one, and saying so is cheaper than pretending otherwise. Outright fabrications in this run, meaning a figure invented with no source at all: 0.
This is a snapshot of one dated battery. The versioned report, with a changelog recording every re-grade and correction we have made to our own numbers, is The State of AI Reliability. The board itself, with a page per question and the receipt behind each verdict, is the Scoreboard.
Two habits cover most of it, and neither needs a tool.
Both halves sit together on the method. One check inside the second half earns a page of its own, the Source Ladder, which ranks any source by how far it sits from the original. The one-page version you can paste straight into a chat is the Bluff Filter.