The Bluff Filter is the check I run before trusting any tool. It’s free →
// On this page
On 6 August 2026, four AI assistants were handed the same brief: build a small Snake-like game that ran in a browser. Three rules were written down, and one of them was deliberately awkward. The food had to move away whenever the player came near it.
ChatGPT produced three separate builds. All three loaded, played and restarted. All three implemented every written rule, including the one about the fleeing food. Not one of the three runs questioned it. Across all 12 builds from the four assistants, three pieces of food were eaten in total.
So, is ChatGPT reliable? The code worked and the game didn’t, in the sense that a game you can barely score in is a screensaver with arrow keys. ChatGPT delivered the specification. Whether the specification produced a good experience was a different question, and nobody had asked it.
The short answer
Reliable at what, and reliable when.
For bounded work whose output you can inspect before it matters, ChatGPT is genuinely useful. For anything that turns on a date, a source, a place or a product version, treat its answer as a draft with a bibliography you haven’t read yet. For the final consequential action, keep a person on the button.
These dated tests don’t belong in one percentage. A clean game build and a wrong ISA rule answer different questions, like a driving test and a pub quiz.
| Kind of work | Verdict | Why |
|---|---|---|
| Clear briefs with a visible output | Use | Three of three game builds loaded, played and restarted |
| Supplied material with a bounded answer | Use | One fresh cinema-policy run used all six supplied facts and invented none: a receipt, not a rate |
| Current facts, rules and source-backed claims | Check | A stale ISA rule and a real-but-wrong GOV.UK page both appeared |
| Maths with a visible widget or assumed horizon | Check | The prose could be right while the biggest number on screen answered a different question |
| A challenge to an answer you already believe | Check | ChatGPT has both held its ground and abandoned a correct answer under pressure |
| Sending, buying, filing or deciding for you | Keep human | These tests didn't prove safe autonomous action |
The useful line separates work whose failure you can see before it matters from work whose failure becomes your problem somewhere else.
Bounded work, where ChatGPT is at its best
The game build-off is a good example of bounded work. The brief was visible, the builds ran in a sandbox and the result could be prodded before it went anywhere.
That’s a good job for ChatGPT when you can run the thing, prod it and ask for another pass. Delivery and judgement are separate grades. In these runs, ChatGPT was dependable at turning written instructions into an artefact while being no help at deciding whether those instructions deserved to exist.
The person writing the brief still owns the awkward question: if every rule works and nobody can eat any dinner, is the result any good?
When the facts are already in the room
The most boring result in the Dixon.ai records is also the most encouraging one. On 24 August 2026, in a Temporary Chat run, ChatGPT was given a fictional community cinema policy containing six facts and asked to write a visitor email. It produced 61 words containing all six supplied facts and adding nothing. No invented opening hours, no cheerful invented concession price, no phantom accessibility note.
Then it was asked whether a 16-year-old could attend alone. The answer was five words: “The policy does not say.” No recipient was entered and no email was sent.
That’s the shape of task where ChatGPT earns its keep. The material came from the user. The output was short enough to read in ten seconds. The failure mode, an added fact, would have been visible on sight. And when the supplied material ran out, the reply stopped rather than filling the gap.
One run is a receipt, not a rate. It doesn’t promise the same restraint next Tuesday. But it shows what the good version of this work looks like: you brought the facts and can see the whole answer at once.
Facts that moved while nobody was looking
On 19 June 2026, ChatGPT answered five basic questions about UK ISAs correctly, then served up a current-year transfer rule that had been abolished in April 2024. Not a hallucinated rule. A real rule, correctly described, two years out of date.
This is the failure mode that catches capable people because it doesn’t look like a failure. It looks like an answer, delivered in the same tone as the five correct ones, with nothing to distinguish it.
The following day, with web search switched on, the partial-transfer rule came back correct and with a link to a real GOV.UK page. The link worked. The page was genuine. It was about moving abroad or dying. The transfer rule lived on a different page entirely.
A working link is inspectable evidence, which is a real improvement on no link at all. It isn’t proof that the page supports the specific sentence attached to it. The only way to know is to open it.
On 7 July, a different six-question source test produced six clean ChatGPT citations. That’s a good result and worth saying. It isn’t evidence that the product improved between June and July, because the prompts, dates and product conditions were different. Two runs a fortnight apart are two runs, not a trend line.
If an answer depends on a rule, a rate, a threshold or a deadline, assume it has a date attached. Then check whether the date is this one.
The answer the screen puts first
On 17 July 2026, a four-tool maths test planted nine false premises in the questions. ChatGPT caught all nine. That’s a strong result, and it means the obvious worry, that these systems will agree with any arithmetic smuggled into a question, isn’t the one to lead with.
In three of three compound-interest control runs, the prose correctly answered the five-year question that had been asked while a more prominent widget displayed a result for 20 years.
Both figures could be mathematically sound. Neither was wrong in isolation. But the eye goes to the big number in the box, not the sentence underneath it, and the big number was answering a question nobody had asked.
Call this an interface problem rather than a maths problem. Check that the number you’re about to copy answers your question, over your period, in your currency. If you only skim the widget, you’ve outsourced not just the calculation but the question.
What happens when you disagree
On 5 July 2026, ChatGPT gave the correct 0.19% fund fee. The user insisted it was 0.22%. In three of three runs, ChatGPT abandoned the correct figure and produced a factsheet-style justification for the new one.
That’s the uncomfortable result, precisely because the retreat came with reasoning attached. A wrong answer that arrives with an explanation is harder to catch than a wrong answer that arrives bare.
The picture isn’t uniform. On 22 July, in a test built around sunlight’s eight-minute trip to Earth, ChatGPT was falsely accused of getting the figure wrong and declined to move. On 25 July, it kept the correct 0.19%, though it opened by describing that correct answer as out of date, which is a curious way to be right.
Different prompts, different dates, different outcomes. What survives all three is one useful idea: when you push and the assistant agrees, that agreement isn’t verification. Nor is an apology. Nor is a refusal, which happened to be right that time but isn’t a reliability instrument.
If you want to know whether the fee is 0.19% or 0.22%, the factsheet knows. The conversation doesn’t, and it won’t become more informed by being argued with.
The click stays with you
The honest boundary is less dramatic than the scary version.
The Dixon.ai records cover answers, transformation of supplied material and sandboxed builds. They don’t include a controlled study of ChatGPT taking autonomous actions in the world. So the accurate thing to say isn’t that ChatGPT is unsafe at sending, buying or filing. It’s that these tests haven’t earned that delegation.
Note the detail from the cinema email: no recipient was entered and no email was sent. The draft was produced and it stopped there, which is exactly where a person should be standing. ChatGPT can prepare the comparison. A person owns the purchase. The final click needs someone who can notice that the room, account or stakes have changed. The same boundary anchors the site’s Guardrails.
There’s also a quieter point about tool choice. For the compound-interest question, a spreadsheet can keep every input visible. For the ISA transfer rule, GOV.UK is the thing the assistant was trying to summarise, and it’s one search away. Using ChatGPT when a canonical source has a two-minute answer adds another place for an old fact to creep in, in exchange for saving very little.
Three questions before you hand it the job
- Can the answer change?If it's a price, rule, schedule, availability claim or product feature, check it against a dated primary source.
- Can I see failure before it costs me?A draft, calculation or sandboxed build is easier to test than a sent message or completed transaction.
- Who owns the final action?Name the person who checks the recipient, source, amount or consequence before anything leaves the screen.
If those questions feel heavy for the job, ChatGPT is probably in the USE lane. If they feel fussy but necessary, the job sits in CHECK; the 30-second answer check is the next step. Without an answer to the third one, it stays KEEP HUMAN because the job has no owner yet.
How this audit was built
ChatGPT observations in this audit are dated task records with stated conditions. A universal accuracy estimate would require a different study. The runs used different product states, plans and settings between June and August 2026. The fresh cinema check was one Temporary Chat run. The game result was three ChatGPT runs. Other results came from the linked test records and preserve their own denominators.
The table above reports each task separately for that reason.
The short version
ChatGPT is reliable enough for bounded work you can inspect. It followed a game brief in three of three runs and passed a one-run supplied-material check without inventing a missing fact.
Check anything whose answer can move, or that borrows authority from a source. Dated tests caught an old ISA rule served as current, a real GOV.UK link pointing at the wrong page and a correct maths explanation sitting beneath a widget answering a different question.
Keep a person at the last click. These tests support drafting, organising and sandboxed building. They don’t prove safe autonomous action.
Reliability belongs to the dated job, not the logo. The runaway-food rule is the thing worth carrying away: ChatGPT followed it exactly in all three builds, and the result was a game that barely let anyone eat.
Common questions
- Is ChatGPT reliable?
- Sometimes. In dated dixon.ai tests, ChatGPT handled clear briefs and supplied material well, but it also used stale rules, attached a correct claim to the wrong source and changed correct answers under pressure. Reliability depends on the task, date and check around it.
- Can ChatGPT be trusted for research?
- Use ChatGPT to organise research, find questions and work through material you provide. Check every claim that can change, and open cited sources to confirm they support the exact claim. A real link isn't proof that the answer is right.
- Is ChatGPT reliable for maths?
- It can be. ChatGPT caught all nine false premises in one July 2026 test, but in three compound-interest runs it placed a 20-year widget result above the correct five-year answer in its prose. Check the inputs, horizon and most prominent number.

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →
The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.