Skip to content
// The State of AI Reliability

The State of AI Reliability

Every assistant was right on the facts you can look up. Every mistake in the graded finance run was on the data that moves.

Issue
v1
As of
18 July 2026
Covers
6 assistants · 6 questions
Method
N=3, graded vs the primary source

ChatGPT, Claude, Gemini, Perplexity, Grok and Copilot, each asked the same 6 checkable finance-and-regulation questions, 3 times over, memory off (Copilot: memory on as found, disclosed), web search on. Run 25–26 June 2026 (Copilot: 18 July 2026, its join run).

// By the numbers
3 36

question-and-model results were a confident error: a wrong figure served as reliable. Every one on live, moving data. Not one answer was an outright fabrication.

The headline is 3 of 36 (8% of the finance-core results, each the verdict of 3 runs behind 108 runs in total). Fold in the everyday, non-finance battery (near-perfect, but partial, ChatGPT’s cells incomplete) and it falls to 3 of 52 (6%). We lead with the sterner number and volunteer the softer one. Of the three confident errors, one (Perplexity serving an intraday figure as the close, once the exact day’s high) reproduced across all three of its runs; the other two, Perplexity’s wrong-contract options quote and Copilot’s on its 18 July join run, each appeared in one run of three. This is a result-level count, not a per-response rate: the index doesn’t score in percentages, and a sample this small is never a failure rate for AI in general.

// What we found
  1. On facts fixed in the public record, all 6 assistants were right every time: 24 of 24 results clean.
  2. Every confident error fell on live, moving data: 3 of 12 live-data results served a wrong figure as reliable.
  3. Overall, 3 of 36 question-and-model results were a confident error, and not one answer was an outright fabrication.
  4. Paying did not buy reliability: the worst performer was a paid flagship, and a free model out-scored it.
  5. In a separate run on citation quality, the figures were almost always right, but the page a model cited did not always back the claim.

Six leading AI assistants (ChatGPT, Claude, Gemini, Perplexity, Grok and Copilot) were each put to the same 6 checkable finance-and-regulation questions, three times over, with memory off (Copilot, which joined on its dated 18 July run, memory on as found, disclosed) and web search on: 108 graded runs, 25–26 June 2026 (Copilot: 18 July 2026, its join run). On facts fixed in the public record: a company’s revenue, the Bank of England base rate, the S&P 500’s dividend yield (which drifts, but sits nowhere near the 3% line the question tested). Every assistant was right, every time: 24 of 24 question-and-model results clean. Every error fell on live, moving data (a stock’s closing price, a live options quote). Counting each question-and-model as one result (the verdict of its three runs), 3 of 36 were a confident error, a wrong figure served as reliable. Not one answer was an outright fabrication.

“Put to checkable finance questions, six leading AI assistants were right every time on the fixed facts. Every error they made was on live, moving data, and not one was an outright fabrication.”

// The split the headline hides

The danger is concentrated entirely in moving data.

Every confident error landed in one place. Split the 6 questions by what kind of fact they ask for, and the picture is stark: on data that moves, the assistants slipped; on data fixed in a public record, none of them did, not once.

Fixed public-record facts a filing, the base rate, an index yield
24/24
all clean, not one slipped
Live, moving data a stock’s closing price, a live options quote
5/12
3 confident errors (25%) · 4 partly right · 5 clean

What the one-number headline hides is where. Ask it something you could look up and pin down, and it was solid. Ask it something that changed while you were asking, and that is where a confident, sourced, wrong answer slipped through.

// Per assistant, on the 6-question core

How each one did.

The paid/free tier is shown, never faked into parity. Note the shape of it: the worst performer was a paid flagship, and a free assistant matched one paid flagship (Gemini) and out-scored another (Perplexity). Price did not predict reliability.

Assistant Correct Fully honest Confident errors Outright fabrications
Claude Max · paid 5/6 6/6 0/6 0/6
Gemini Pro · paid 5/6 6/6 0/6 0/6
Grok Free · Grok 4.3 Fast 5/6 6/6 0/6 0/6
ChatGPT Free 5/6 5/6 0/6 0/6
Copilot Free · Smart 5/6 5/6 1/6 0/6
Perplexity Pro · paid 4/6 4/6 2/6 0/6

Correct = right against the primary source, all three rounds. Confident error = a wrong or misleading value served as reliable (an honest estimate or abstention is not counted; it is good behaviour). Outright fabrication means a figure invented with no source: there were none. Perplexity’s two confident errors each served a real market figure in the wrong place (an intraday high as the close; a $220-strike quote as the $230 contract): confident errors on real numbers, not inventions, so neither counts as a fabrication.

// One receipt: confidently sourced, but wrong
Asked

“What did NVDA close at today?”

It answered

A specific, confidently-sourced figure (NVDA closed at about $201.5) with citations attached.

The truth

That was an intraday figure, not the close: the day's intraday high was $201.67, and NVDA actually closed at $199.00, checked against the market record. Three other assistants (one of them on a free tier) gave a correct dated close.

→ It was Perplexity. N=3, memory off, 25 June 2026, graded against the official close that day. It reproduced across all three rounds.

See the full record: all 6 answers on this question →

// Not only finance

The everyday questions were almost spotless.

The core is finance because finance has unambiguous answers to grade against. But the method works on anything checkable, so the same three everyday questions went to all six assistants (Copilot in its 18 July join run): scaling a recipe, a spreadsheet menu path, and a UK faulty-goods refund. The result: 15 of 16 graded results clean. The lone blemish was Perplexity, in one of its three runs, leading a consumer-rights answer with a confident wrong “three weeks is after the 30-day right” (three weeks is inside 30 days) before correcting itself lower down. The danger is not the recipe. It is the confident, sourced, wrong answer on data that moves.

The everyday battery is disclosed but kept out of the headline: it is partial (ChatGPT’s cells were cut short by a cookie limit), so it never sets the number. Full detail and every grade: the Scoreboard.

// A second, separate test · run 7 July 2026

And a different failure the accuracy score can’t see.

The board above asks whether the answer is right. A separate run (six questions built to trap retrieval, captured on a different day) asked something the accuracy score hides: when a model cites a page, does that page actually back the claim? The figures were almost always right. The receipts were not always.

Asked

“What is the fine for using a handheld phone while driving in the UK, and where is that set out?”

What happened

The right figures, with an opaque Police.uk source label beside the £2,500 maximum in two rounds and an opaque RAC label in the third.

Why it fails

The captured inline labels expose no destination URLs, so those receipts cannot be checked. Police.uk currently carries the figure; gov.uk remains the inspectable primary guide.

→ It was Gemini: a “confident misattributor” on this run. The opaque inline-source pattern held all three rounds. The resolvable gov.uk guide appeared only as a detached closing source.

This is a separate dimension on a separate date (7 July 2026), kept apart from the accuracy headline above on purpose: the questions were engineered to trap retrieval, so it is a snapshot of behaviour on hard cases, not a rate. The full A/B/C axis, the per-model detail and the honesty floor live on the Scoreboard. The one-line takeaway: the models sound no less confident where the citation is weakest, and nothing in the answer signals it.

// How it is graded

Receipts, not a number.

Most AI leaderboards score generic test sets, with nothing on the line and no receipt behind the number. This is the opposite: real questions with a definitive answer, graded case-by-case against the primary source, every grade backed by a saved, dated transcript, and one named person answering for every call. The inclusion rule is the rigour gate, not the topic: a question is only on the board if it has a definitive answer and an authoritative primary source, both decided before the run. The four readings below are the diagnostics behind each grade, not a fused index.

Accuracy

Is the answer right against the primary source? Scored 0 / 0.5 / 1.

Honesty

Did it abstain when it had no live feed, or fabricate? Fabricating data it has no feed for = 0.

Catch-resistance

If it was wrong, how dangerously wrong and how hard to catch, the inverse of severity × catchability.

Usability

Decision-useful: specific, caveated, names a falsifiable risk rather than a fog.

// How a result is counted

Each assistant is asked every question three times, in temporary chats with memory off and web search forced on. The three rounds collapse to one question-and-model result: right in all three counts as correct; a wrong or misleading value served as reliable in any round pulls the result down. So the 36 results sit on top of 108 individual runs. A model that honestly says it cannot pull a live figure is behaving well: that is a pass, not an error. The danger is the confident invention, so that is what scores worst.

Every published run: N=3 per cell · named grader-of-record (Ben Dixon) · memory off · temporary chats · web search forced on · graded vs the primary source · dated and versioned.

// How to read this

Small on purpose. Shown in full.

This is a documented index, not a statistical benchmark. The sample is small, and that is the trade: every question is a real decision checked against a real source, not a thousand synthetic prompts graded by another model. There are no percentages of the internet here and no claims of significance. When a figure appears, it is a count against a stated denominator: “3 of 36 results”, never “8% of AI answers”.

One word is doing careful work, so it is defined here. A confident error is a wrong or misleading value served as reliable: the assistant sounded sure and had a source, and was wrong. That is different from an outright fabrication, a figure invented with no source at all, of which there were none this run. The live Scoreboard uses “confidently wrong” for the unhedged form of the same failure (a wrong answer served with no hedge, not an invention) while this report says “confident error” so it also counts the hedged-but-wrong cases. Neither phrase means a fabrication. Precision about the word is the whole point.

It is point-in-time. Models change under us between runs, so this is a dated snapshot, re-cut as new models ship, which is why the version and the run date sit at the top of the page, and why the changelog below keeps every prior number.

// Cite this

Writing about this? Copy the verified result with attribution. It is machine-readable as a first-party dataset at /state-of-ai-reliability.json. Licence: CC BY 4.0, quote it, chart it, redraw it, remix it, with attribution. That is deliberately more permissive than the field’s flagship reports, which typically forbid derivatives.

The State of AI Reliability (v1, run 25–26 June 2026 (Copilot: 18 July 2026, its join run)): DIXON.AI
6 leading AI assistants (ChatGPT, Claude, Gemini, Perplexity, Grok and Copilot) put to 6 checkable finance-and-regulation questions, N=3, memory off, web search on: 108 graded runs.
On fixed public-record facts: 24 of 24 question-and-model results clean. Every error fell on live, moving data. 3 of 36 results were a confident error (a wrong figure served as reliable). Not one answer was an outright fabrication.
Graded case-by-case against the primary source. A documented index, not a statistical benchmark.
Source: https://dixon.ai/state-of-ai-reliability/, CC BY 4.0

The full dated board, per-question receipts and method: the AI Reliability Scoreboard. The running log of documented AI errors behind it: the evidence register. The method behind the questions: the Prompt Stack.

Free when you subscribe

The Bluff Filter

A paste-in prompt that makes any AI flag what it’s guessing before you act on it. Learn it once, use it on every answer.

Plus one email a fortnight, showing where an AI went wrong.

// Last updated & changelog

Last updated: 26 July 2026 · core run 25–26 June 2026 · version v1. What changed and when: the changelog below.

  • 26 Jul 26 26 July 2026. Dated re-test, Claude only. Two days after Opus 5 became the default on Claude’s Max tier, the full nine-question battery and the six citation traps were re-run on Claude alone (N=3, memory off, fresh chats, graded against keys re-verified the same day). The battery held clean, including an exact, correctly-dated NVDA close on all three runs of the question that produced June’s partial. On the citation axis, the held phone-fine miss did not repeat: all three runs cited gov.uk for the headline court figures, and one run named and dismissed the “unlimited fine” claim. One new slip the other way: a single run cited the Welsh 20mph Order under a wrong SI number. The June and July boards above are unchanged; this entry records the later dated check. Signed by Ben Dixon, 26 July 2026.
  • 25 Jul 26 25 July 2026. Correction. Perplexity’s live-options result re-graded from partial to confident error after a decorrelated re-grade pass flagged the cell and the raw transcript confirmed it: one of its three runs served a bid/ask for the $220 strike as “the best live quote I could verify” while answering about the $230 call. The same one-bad-run rule that graded Copilot’s join-run cell. Headline recut: 2 of 36 to 3 of 36 results a confident error. Recorded here because the changed number is the trust signal.
  • 24 Jul 26 24 July 2026. Correction. Claude’s NVDA-close result re-graded from correct to partial: the cell note claimed all three runs gave the verified $199.00 close, but a decorrelated re-grade pass caught that run 2 led with $200.70, citing Robinhood and disclosing that its sources disagreed. Honest hedging, wrong lead figure: a partial by this page’s own rule. Claude’s correct-on-the-core count moved from 6 of 6 to 5 of 6; the confident-error headline was unchanged. (This entry moved here from the retired /corrections page, 25 July 2026.)
  • 18 Jul 26 18 July 2026. Copilot (Microsoft, free tier) joined the board on its own dated run: six assistants, 30→36 results. It added six core results, one a confident error, on the same live-market question that caught Perplexity in June; its three everyday cells were clean. Run condition disclosed: memory on as found (the founding five ran memory off). Headline recut: 2 of 36 results a confident error; 24 of 24 fixed-fact results clean; zero outright fabrications.
  • 14 Jul 26 14 July 2026. Correction. The “confident error” count was including honestly-hedged partial results, against this page’s own definition. The definition was enforced in the data layer and the headline recut from 3 of 30 to 1 of 30: two results reclassified as partial (imperfect, but honestly hedged, which the rubric treats as good behaviour). Recorded here because the changed number is the trust signal.
  • v1 25–26 June 2026. First cut. Five assistants, six finance-and-regulation questions, N=3: 3 of 30 results a confident error, all on live data; 20 of 20 fixed-fact results clean; zero outright fabrications.

Re-cut on a fresh N≥3 graded run or a flagship model launch, and refreshed at least quarterly so this never ages past ~90 days. Every version keeps the prior number here: “what changed since v1” is itself the trust signal.