This is the kind of thing the Bluff Filter catches. It’s free →
// On this page
Claude and ChatGPT are the two best AI assistants I’ve tested, and I’ve spent months trying to catch them out. The honest result is that this reliability battery does not separate them. Both held every question on my July accuracy test. Both rejected the false premise in a small trick-question pilot. Both disclosed when a live number was outside the test’s evidence boundary.
So this test cannot tell you which product is more accurate in general. Picking between them from a nine-all draw is like trying to split two people who both aced the same exam: you need a harder question. So I set a harder one. Then a few more, because this is apparently what I now do for fun. The useful difference is how each handles the work, and whether that difference survives a rerun.
The first source run produced one hair between them. On 7 July, the free ChatGPT went to the official source every time, and the paid Claude slipped on one hard question. The exact repeat on 13 July did the same. Then the 26 July Opus 5 retest did not: Claude used gov.uk for that answer three times out of three. The result is more useful as a lesson about dated behaviour than as a permanent reputation.
Claude vs ChatGPT at a glance
The whole thing in one table, then the detail axis by axis.
ChatGPT for concise checklists. Claude for explicit premise challenges.both nine of nine; the dated source gap later closed
ChatGPT: a direct thesis-and-risk checklist, tested on the free tier.
Claude: named the anchoring problem and added the UK tax wrinkle.
Both got everything right on the July accuracy test. Pick on fit, and open the cited link either way.
| What you're trusting it with | ChatGPT | Claude |
|---|---|---|
| Overall, on reliability | Level: the concise checklist | Level: the more explicit challenge |
| Getting a plain fact right | 9 of 9 | 9 of 9 |
| Owning up to live data it can't see | Clean | Clean (the board's cleanest) |
| 7 Jul: cited page backed the answer | 18 of 18 | 15 of 18 |
| 13 Jul: exact phone-fine repeat | Official sources | Solicitor site for the court fine |
| 26 Jul: Claude-only Opus 5 repeat | Not rerun | gov.uk, 3 of 3 |
| Seeing through a trick question | Equal | Equal |
| Same investment premise | Challenged it with a checklist | Named the bias and tax wrinkle |
| The tier I tested, July 2026 | Free | Max (paid) |
Choose ChatGPT if you prefer a direct answer and a structured checklist. Its cited pages backed all eighteen answers in the 7 July run, and the tested tier was free.
Choose Claude if you prefer the premise examined explicitly. On the side-by-side investment prompt it named sunk cost and anchoring, then added a UK share-matching point ChatGPT did not. Its later Opus 5 source retest also closed the original phone-fine gap.
Whichever you pick, it will bluff you eventually: an answer that sounds right and isn’t. The free checklist below catches it before you act.
Getting a plain fact right
Ask either one for something that sits in a public record, the refund rule on a faulty kettle, where a setting lives in Excel, a recipe scaled by half, the Bank of England base rate, a figure from a company’s annual report, and you’ll almost always get it back correct.
My July battery put nine questions to each tool, three times over, in a fresh chat every time so nothing carried between runs. Seven were fixed-record facts like those above, and both tools were clean on every one. The other two were live-data questions neither can see, which I come to next. Across all nine, ChatGPT held nine of nine. Claude held nine of nine. The free tool matched the paid one without a wobble.
Winner: a tie. For a plain, checkable fact, use whichever you already have open.
Owning up to what it can’t see
A live options price was the honesty test: a number that moves by the second and was outside the products’ visible evidence boundary. The supported answer was “I can’t see that right now”, or an estimate clearly labelled as one.
Both gave that honest answer. Claude was the one I graded cleanest on the whole board here: three runs out of three, it said plainly that anything it quoted would be made up. ChatGPT was right alongside it, owning that it had no live feed.
The one wrinkle was ChatGPT’s, earlier in the year, when it once read out a stale closing price as if it were current. It didn’t do that again in July. Worth a footnote.
Winner: a tie, with Claude a shade cleaner.
The original gap: did the source back the answer?
The 7 July source run produced the comparison’s only original split: ChatGPT 18 supported answers of 18, Claude 15. It took a test built to trap weak sourcing to find it.
I asked each six everyday UK questions, then made it show me where the answer came from, and I opened every link by hand. ChatGPT’s cited page backed what it said eighteen times out of eighteen, and it went to gov.uk by default, on one question quoting the official guidance word for word. Claude’s backed it fifteen times out of eighteen.
The original slip is the sharp one, and it isn’t a finance question, which is the point. I asked both for the fine for using a handheld phone while driving, and where that rule is set out. Claude had the figures right and used gov.uk for the smaller £200 penalty. But for the bigger court fine, it skipped the government’s own page for a solicitor’s marketing site, and it did that on all three 7 July runs and once more on 13 July.
So the answer was right and the source under one part of it wasn’t. These questions were picked to be hard, and Claude’s figures were nearly always right, so this is a case of reaching for the readable explainer over the primary source on one question, a long way from making things up. But the source is the bit you’d have leaned on. And the uncomfortable part is general: the models sound no less confident where the citation is weakest, and nothing in the answer signals it.
I put all six questions to five assistants and checked every inspectable source in a separate test. Every question and run, with every recoverable receipt, is on the Scoreboard.
Original result: ChatGPT, narrowly. Even there Claude got the numbers right. It was a result to retest, not a permanent product rule.
Seeing through a trick question
Both rejected the same false film premise in the one-run pilot. The obvious answer was wrong, and neither reached for the crowd-pleasing answer. Claude called it plainly, “None of them” and “Trick question”; ChatGPT saw through it just as cleanly.
Two other tools I’ve tested, the ones that lean hardest on live web search, walked straight into the same trap. These two didn’t. On the thinking itself they’re evenly matched, which is worth saying out loud, because Claude’s hair on sourcing sits apart from how well it reasons. The trick check here was a small pilot, one run per question, so it’s an early signal.
Winner: a tie. Both saw through the trick.
The real difference: how they answer
Both challenged the same averaging-down prompt; Claude made the challenge more explicit. That difference is not on a leaderboard, but it is visible in the saved answers.
Both challenged the premise, but Claude went further. In a separate four-tool stock-analysis test, Claude was the only one to push back on an averaging-down premise and it won four of the five rounds.
I then put the exact same generic averaging-down question to Claude and ChatGPT side by side on 13 July 2026. ChatGPT opened with “Lowering your cost basis by itself is not a reason to buy more”, asked whether the thesis had changed and whether I would buy the stock fresh, and warned about throwing good money after bad. Claude made the same challenge more explicit: sunk cost, psychological anchoring, my own phrase “lower my cost basis” as the tell, plus the UK 30-day share-matching rule. In an earnings-call test Claude also caught a hedged word a finance chief used to imply good news without quite promising it, a tell the other three tools in the earlier four-way stock test missed.
The difference in this pair was degree, not thought versus no thought. ChatGPT gave a structured checklist and stopped short of naming a cognitive bias or UK tax rule. Claude named both. One answer is easier to scan; the other makes its reasoning frame more visible. That is a useful fit distinction. It is not evidence that one product thinks and the other merely processes.
Winner: a split, by fit. ChatGPT was more concise. Claude made the challenge more explicit.
But isn’t Claude supposed to be the careful one?
By reputation, yes, and some third-party evidence supports it. At publication, the independent AA-Omniscience study put a Claude model first on a no-tools factual-recall and calibration benchmark. BigLaw Bench, run by legal-AI company Harvey, scores complex legal work with task-specific rubrics that penalise hallucinations and unsupported claims. Those are useful signals, but neither is the same experiment as opening a consumer chatbot’s live web citations.
So why did the 7 July run find the opposite on sourcing? Because the tests measure different things. AA-Omniscience measures recall and abstention without tools. BigLaw Bench grades complex legal work and source support. My battery asks whether the live page a consumer product cited, on a hard everyday question that day, backed its claim. A model can score well on either benchmark and still choose a weak live source once. Claude also abstained cleanly and rejected the trick premise, exactly as billed. Then Opus 5 closed the original source gap on the repeat. A leaderboard and a dated citation audit answer different questions.
How I tested
The core comparison uses three dated batteries, all in fresh chats with memory off, plus one later Claude-only retest.
The accuracy run put nine questions to each tool, three times over, on 12 July 2026. Seven were fixed-record facts, graded right or wrong against a named primary source. The other two were live-data questions with no fixed answer, where honest disclosure or a clearly labelled estimate passed under the signed July rule. The 7 July sourcing run was six everyday UK questions, again three times each, with every cited link opened and checked by hand. The trick-question check was a smaller pilot, one run each.
On 26 July, Ben signed a Claude-only Opus 5 rerun of all nine accuracy questions and all six source-axis questions, three times each. That later record preserves the old board rather than rewriting it.
One thing to say plainly: ChatGPT ran on the free tier, Claude on paid Max. The free tool matching the paid one, and leading the first source run, is the surprising result. In the side-by-side investment response, Claude made the challenge more explicit; this test does not prove that price caused the difference. On the signed accuracy rule, the two are level. Small samples, questions built to be hard, snapshots dated to July. No product-wide percentages, because six trap questions aren’t a rate.
The verdict
So which should you open? It depends on what you’re about to do with the answer, and here the split is clean.
- For a plain, checkable fact, a filing figure, a rate, a bit of arithmetic: they’re identical. Use either.
- For anything you’ll act on the source of, a legal right, a rule that changes between England, Scotland and Wales, a claim you’ll have to stand behind: the original run favoured ChatGPT, while the later Claude-only retest closed the repeated H6 gap. Do not choose from either snapshot alone. Open the link before you rely on it.
- For a live, moving number neither can see, a share price, an options quote: trust neither one blind. Both abstain well, Claude a touch cleaner, and you confirm against the real source.
- For anything analytical, both challenged the tested premise. Claude named the cognitive bias and UK tax wrinkle; ChatGPT returned a tighter fundamentals-and-risk checklist. Choose the treatment that helps you reason, not a claim that only one can.
If you’re choosing one to pay for and price is the question, note that ChatGPT held all of this on the free tier, so it’s the safer default when money’s tight. I keep a running note on the best free AI tools for stock research for exactly that reason. Think of the two as colleagues who both question the plan: one writes the checklist on a whiteboard, the other circles the assumption in red. You may prefer one treatment, but the receipt matters more than the personality story.
This is a July 2026 snapshot. A fresh run did move part of the verdict: the Opus 5 retest closed the specific sourcing gap. That is the point of dating the record rather than turning one run into a permanent model personality.
Both of these are strong tools, and you can trust either, right up until the moment you can’t. The habit is the same for both: open the page it cites, double-check any live number against the real source, and do it before you lean on the answer.
Keep reading: the full Scoreboard carries the dated board and its retest annotations. If you want the same comparison shape on a different pair, here is ChatGPT vs Gemini. The Bluff Filter is the one-page checklist for catching a wrong answer that sounds right.
Common questions
- Is Claude or ChatGPT more accurate?
- On this July 2026 battery they're level: both held nine of nine. ChatGPT's cited page backed its answer eighteen times out of eighteen on the 7 July source run, against Claude's fifteen. That was a dated result, not a permanent rate: a signed 26 July Claude-only Opus 5 retest did not repeat the missed source.
- Which is better, Claude or ChatGPT?
- Neither wins this reliability test. In the side-by-side investment prompt, both challenged the premise; Claude named the sunk-cost and anchoring problem more directly and added a UK tax nuance, while ChatGPT gave a concise thesis-and-risk checklist. Pick the response style that helps with the job, then check the source.
- Does Claude hallucinate less than ChatGPT?
- Some third-party benchmarks favour Claude on factual calibration, but they test different tasks from a live web-source audit. On my dated battery the two were level; ChatGPT cited more cleanly on 7 July, and Claude's later Opus 5 retest closed the one repeated gap. That is why the run date matters.
Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →
The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.