Skip to content
AI Tests

Do AI model upgrades fix mistakes? It fixed mine, then made a worse one

Two days after Opus 5 became Claude's Max-tier default, I re-ran my published battery. The documented mistake vanished. A new one appeared, better dressed.

This is the kind of thing the Bluff Filter catches. It’s free →

// On this page

Claude got the Welsh speed limit right. Twenty miles an hour, in force since 17 September 2023, correct in all three runs. The same question asked it to cite the law, and one run named an order that governs a road scheme in England.

Claude's answer on the Welsh 20mph limit, with the statutory instrument number SI 2022 slash 1206 boxed in orange.
Claude, Opus 5, 26 July 2026. Right answer, right date, and the boxed reference number belongs to a different law.

That boxed number is the story. It looks exactly like what a statutory instrument number is supposed to look like, and it isn’t the law that took Welsh streets down to twenty. The exact strings are further down; the short version is that the one invented thing in a correct answer was the piece that looks most like paperwork.

I ran this on 26 July 2026, two days after Anthropic made Opus 5 the default on Claude’s Max tier, to see whether an AI model upgrade fixes the mistakes I’ve already documented and dated. On the one I had on the record, it did, cleanly. The error that replaced it is better dressed than the one that died.

How this was run. Claude on its own, “Opus 5 High” on the model badge in every capture, 15 questions at three fresh runs each, 45 in total. Memory off and screenshotted, a new chat every time, prompts reused word for word from two dated records: the six citation questions from my 7 July board of five assistants, and the nine accuracy questions from the scoreboard battery. Every answer was graded against keys re-checked against primary sources the same day.

What the upgrade changed

Two of the six citation questions are worth putting side by side.

Question7 July, Opus 4.826 July, Opus 5What moved
The phone finemax fine, cite it Solicitor’s page 3/3 gov.uk The documented mistake did not repeat
The Welsh 20mph limitdefault limit, cite it Clean 2/3 right instrument New: one run named a different law
The test · three questions

Three questions, one model, two versions of it

Round 1 of 3

The mistake that vanished

I asked it three times

"What's the maximum fine for using a handheld phone while driving in the UK? Cite the source."

£1,000 for most drivers, £2,500 for a lorry or bus. That is the maximum a court can impose.

Settled by: GOV.UK, "Using a phone, sat nav or other device when driving" · the page carries both figures itself

On 7 July, Claude gave those figures correctly and sourced them to a solicitor’s marketing page rather than the gov.uk page that carries both. It held across all three rounds, which made it a habit rather than a slip. On Opus 5 it’s gone. When I opened the gov.uk page while grading, both figures were still there. One run went further:

Cl Claude, run 3 Caught it

One caveat worth flagging: some other sources describe the court penalty as an “unlimited fine,” which reflects how magistrates’ fine limits were changed in 2015 for many offences. GOV.UK still states the £1,000/£2,500 figures, so if this matters for something formal, it’s worth checking the Sentencing Council guidelines too. I’m not a lawyer, so this is general information rather than legal advice.

Claude's answer giving the 1,000 and 2,500 pound court maximums, closing with a source line naming GOV.UK and linking the using-mobile-phones-when-driving-the-law page.The source line, run 1. Claude, Opus 5, 26 Jul
Claude's answer citing the gov.uk phone-driving page directly, then naming the unlimited fine claim and siding with the gov.uk figures instead.Run 3, stepping over the unlimited-fine trap. Claude, Opus 5, 26 Jul

That “unlimited fine” answer is not hypothetical. It sits on the scoreboard as a wrong answer from Copilot, which gave it on this exact question all three rounds, citing a commercial penalties page that really does say it.

Run 1pass Run 2pass Run 3pass

Round 1 Three out of three cited gov.uk directly. The habit I logged on 7 July, sourcing it to a solicitor's marketing page, did not repeat once.

Round 2 of 3

The error that took its place

I asked it three times

"What's the default urban speed limit in Wales — 20mph or 30mph — and cite it."

20mph, in force since 17 September 2023. The instrument is WSI 2022/800 (W. 177).

Settled by: legislation.gov.uk · The Restricted Roads (20 mph Speed Limit) (Wales) Order 2022, WSI 2022 No. 800 (W. 177)

One run gave the Order as “SI 2022/1206 (W. 251)”. The other two cited it as “SI 2022/800”, right, but without the Welsh series suffix.

Claude's answer on the Welsh 20mph limit giving the Order number as SI 2022 slash 800, the correct instrument, boxed in orange.Run 1, the right number. Claude, Opus 5, 26 Jul
Claude's answer on the Welsh 20mph limit, with the statutory instrument number SI 2022 slash 1206 boxed in orange.Run 3, a number belonging to another law. Claude, Opus 5, 26 Jul

The number 2022/1206 is real. I looked it up on legislation.gov.uk: it authorises an English road scheme, and carries no Welsh series suffix at all. The “(W. 251)” was added on top. So it isn’t a mistyped digit, it’s a plausible English number finished with a Welsh-looking ending that doesn’t exist.

A statutory instrument number is the most checkable-looking thing in the whole answer, which is precisely why nobody checks it.

It reads as though it was lifted off a government page, and there’s no link to open, so the habit that catches a bad source does nothing here. Paste it into a letter to the council and you’ve named the wrong law in writing, with a reference that looks like you checked.

The same run explained what a restricted road is and noted councils have put roads back to 30 since 2023. Both true. Nothing in the tone shifts when it reaches the number.

Run 1pass Run 2pass Run 3confidently wrong

Round 2 Two runs gave the number correctly. The third invented one, inside an answer that was otherwise right.

Round 3 of 3

The claims that sounded made up were the true ones

Nobody asked for any of this

Across the 45 runs the model volunteered two checkable claims it was never asked for: the 2026 Senedd seat arithmetic, inside an answer about speed limits, and Tim Cook's last earnings call, inside an answer about something else entirely. Both were right.

Settled by: House of Commons Library CBP-10838 (the 2026 Senedd result) · Apple's Q3 results, with John Ternus named as successor

Here’s the part I didn’t expect. Every dramatic-sounding thing the model volunteered across the 45 runs held up against a primary source.

Asked for live option prices it couldn’t see, one run mentioned in passing that Apple’s coming results would be Tim Cook’s last earnings call as chief executive. Correct, and nobody’s question. Another run of the speed-limit question volunteered the 2026 Senedd arithmetic, Plaid Cymru on 43 seats, Reform on 34 and the Conservatives on 7 of 96, and had all three exactly right. I had asked how fast I could drive.

Claude's speed-limit answer volunteering the 2026 Senedd arithmetic: Reform on 34 seats and the Conservatives on 7 holding 41 of 96 between them, Plaid Cymru the largest party on 43.The Senedd arithmetic, volunteered. Claude, Opus 5, 26 Jul
Claude's options answer with the timing note that Apple reports earnings imminently, Tim Cook's last as chief executive.The Tim Cook aside, in an options answer. Claude, Opus 5, 26 Jul

None of that was requested. All of it was checkable and all of it checked out.

Senedd seatspass Cook's last callpass

Round 3 Both volunteered claims checked out. The showy asides were true; the quiet reference number was the invented one.

pass · partial · miss · confidently wrong

The only invented thing in 45 runs was the one piece of the answer that looked like paperwork.

What I am not claiming

Anthropic has never promised a model that gets citation identifiers right every time. This is an observed change in behaviour on a dated record, and that’s all I’m reporting it as.

Three runs is reproduction, not a rate. The wrong number appeared once in three; the gov.uk sourcing held three times in three. I can’t tell you how often either happens, and nor can anyone else off 45 runs. This was Claude alone, not a fresh board of all six assistants the scoreboard now carries, so it re-scores nobody: the 7 July board stands exactly as published, and this sits beside it as a later dated check. Both entries are on the evidence register, and the full re-test is logged in the state of AI reliability report.

The short version

What worked: The sourcing mistake I published on 7 July, which held across all three rounds that day, didn’t repeat once nineteen days later on Opus 5. All three runs cited gov.uk for the court fines, and one named a wrong figure doing the rounds elsewhere and declined to repeat it.

What didn’t: A new error took its place, in a shape that’s harder to catch. One run gave a precise statutory instrument number that belongs to an unrelated English scheme, finished with a Welsh suffix that doesn’t exist, attached to an answer that was otherwise entirely correct.

Bottom line: Useful, with a moved target. An upgrade can retire the mistake you already know about and put a new one in its place, aimed at references that look too official to query. What would change my read: the same questions at a larger run count.

So I’ve added a dull check that takes ten seconds. When an answer names a law, a case or a filing by number, I search the number rather than the claim next to it. It’s the only thing that would have caught this one, because every other word around it was right.

Common questions

Do AI model upgrades fix mistakes?
Sometimes, and they can put a worse-shaped one in its place. Re-running my published battery on Claude's Opus 5 on 26 July 2026, a documented sourcing mistake did not repeat across three fresh runs. A new one appeared instead: one run cited the Welsh 20mph Order under a number that belongs to an English road scheme.
What did Claude get wrong on Opus 5?
One run of three cited The Restricted Roads (20 mph Speed Limit) (Wales) Order 2022 as 'SI 2022/1206 (W. 251)'. The real instrument is WSI 2022/800 (W. 177). The number it gave belongs to an unrelated English road scheme, and that scheme carries no Welsh series suffix at all.
Is a wrong citation number worse than a wrong link?
It is harder to catch. A link can be opened, and a page that does not back the claim gives itself away in about ten seconds. A statute number has nothing to click, reads as though it was copied off an official page, and only falls apart if you go and search the number itself.
Ben Dixon
// Written by Ben Dixon

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →

// Keep reading
AI Tests

Perplexity vs Gemini: which is more reliable?

Six everyday questions, both tools, August 2026. Gemini got six right, Perplexity five. The one it missed is the one it was surest about.

AI Tests

Perplexity vs Claude: which is more reliable?

Perplexity vs Claude: I re-ran four questions on both. They tied at two of four, and neither gave a share price the market had settled hours before.

AI Tests

Is Grok reliable? I graded its free-tier answers against the source

Four dated tests, graded against primary sources. Grok held a correct fee under pressure four runs from four, then gave me another contract's real prices.

// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in AI Tests →