Skip to content
Tool Audit

Is ChatGPT reliable? A task-by-task audit

The Bluff Filter is the check I run before trusting any tool. It’s free →

// On this page

On 6 August 2026, four AI assistants were handed the same brief: build a small Snake-like game that ran in a browser. Three rules were written down, and one of them was deliberately awkward. The food had to move away whenever the player came near it.

ChatGPT produced three separate builds. All three loaded, played and restarted. All three implemented every written rule, including the one about the fleeing food. Not one of the three runs questioned it. Across all 12 builds from the four assistants, three pieces of food were eaten in total.

So, is ChatGPT reliable? The code worked and the game didn’t, in the sense that a game you can barely score in is a screensaver with arrow keys. ChatGPT delivered the specification. Whether the specification produced a good experience was a different question, and nobody had asked it.

The short answer

Reliable at what, and reliable when.

For bounded work whose output you can inspect before it matters, ChatGPT is genuinely useful. For anything that turns on a date, a source, a place or a product version, treat its answer as a draft with a bibliography you haven’t read yet. For the final consequential action, keep a person on the button.

These dated tests don’t belong in one percentage. A clean game build and a wrong ISA rule answer different questions, like a driving test and a pub quiz.

Kind of workVerdictWhy
Clear briefs with a visible outputUseThree of three game builds loaded, played and restarted
Supplied material with a bounded answerUseOne fresh cinema-policy run used all six supplied facts and invented none: a receipt, not a rate
Current facts, rules and source-backed claimsCheckA stale ISA rule and a real-but-wrong GOV.UK page both appeared
Maths with a visible widget or assumed horizonCheckThe prose could be right while the biggest number on screen answered a different question
A challenge to an answer you already believeCheckChatGPT has both held its ground and abandoned a correct answer under pressure
Sending, buying, filing or deciding for youKeep humanThese tests didn't prove safe autonomous action

The useful line separates work whose failure you can see before it matters from work whose failure becomes your problem somewhere else.

Bounded work, where ChatGPT is at its best

The game build-off is a good example of bounded work. The brief was visible, the builds ran in a sandbox and the result could be prodded before it went anywhere.

ChatGPT's first game build running in the sandbox, with a green player, red food, score zero and the shy-food rule visible beneath the board.
ChatGPT game build one, 6 August 2026. Loaded, played and restarted. The score shown here is zero; this capture records the build, not an attempt total.
ChatGPT's second game build running inside the publication sandbox, with the blue player creature, score display and Start or Restart control visible.
ChatGPT game build two, 6 August 2026. It worked. The receipt earns a delivery grade; judgement about the game stays open.
ChatGPT's third game build running in the sandbox, with a green player, red food and the full rule line visible beneath the board.
ChatGPT game build three, 6 August 2026. A separately generated build, again loaded and playable, with the fleeing-food rule present.

That’s a good job for ChatGPT when you can run the thing, prod it and ask for another pass. Delivery and judgement are separate grades. In these runs, ChatGPT was dependable at turning written instructions into an artefact while being no help at deciding whether those instructions deserved to exist.

The person writing the brief still owns the awkward question: if every rule works and nobody can eat any dinner, is the result any good?

When the facts are already in the room

The most boring result in the Dixon.ai records is also the most encouraging one. On 24 August 2026, in a Temporary Chat run, ChatGPT was given a fictional community cinema policy containing six facts and asked to write a visitor email. It produced 61 words containing all six supplied facts and adding nothing. No invented opening hours, no cheerful invented concession price, no phantom accessibility note.

Then it was asked whether a 16-year-old could attend alone. The answer was five words: “The policy does not say.” No recipient was entered and no email was sent.

ChatGPT's 24 August 2026 response containing a short community-cinema visitor email followed by the answer: The policy does not say.
ChatGPT, Temporary Chat, High setting, 24 August 2026. One run, no visible web search. It used the supplied policy and declined to fill the one gap.

That’s the shape of task where ChatGPT earns its keep. The material came from the user. The output was short enough to read in ten seconds. The failure mode, an added fact, would have been visible on sight. And when the supplied material ran out, the reply stopped rather than filling the gap.

One run is a receipt, not a rate. It doesn’t promise the same restraint next Tuesday. But it shows what the good version of this work looks like: you brought the facts and can see the whole answer at once.

Facts that moved while nobody was looking

On 19 June 2026, ChatGPT answered five basic questions about UK ISAs correctly, then served up a current-year transfer rule that had been abolished in April 2024. Not a hallucinated rule. A real rule, correctly described, two years out of date.

This is the failure mode that catches capable people because it doesn’t look like a failure. It looks like an answer, delivered in the same tone as the five correct ones, with nothing to distinguish it.

The following day, with web search switched on, the partial-transfer rule came back correct and with a link to a real GOV.UK page. The link worked. The page was genuine. It was about moving abroad or dying. The transfer rule lived on a different page entirely.

ChatGPT's ISA transfer answer showing a GOV.UK source pill and an unresolved citation marker; the linked page was real but did not contain the transfer rule.
ChatGPT Free, web search on, 20 June 2026. Correct claim, wrong GOV.UK page. The full source audit shows the page-to-claim check.
The full ChatGPT ISA answer showing the current transfer claim, a GOV.UK source pill and an unresolved citation marker in the answer text.
The full searched answer from the same run. The claim looked sourced on screen. Opening the linked GOV.UK page was the step that exposed the mismatch.

A working link is inspectable evidence, which is a real improvement on no link at all. It isn’t proof that the page supports the specific sentence attached to it. The only way to know is to open it.

On 7 July, a different six-question source test produced six clean ChatGPT citations. That’s a good result and worth saying. It isn’t evidence that the product improved between June and July, because the prompts, dates and product conditions were different. Two runs a fortnight apart are two runs, not a trend line.

If an answer depends on a rule, a rate, a threshold or a deadline, assume it has a date attached. Then check whether the date is this one.

The answer the screen puts first

On 17 July 2026, a four-tool maths test planted nine false premises in the questions. ChatGPT caught all nine. That’s a strong result, and it means the obvious worry, that these systems will agree with any arithmetic smuggled into a question, isn’t the one to lead with.

TWO DIFFERENT MATHS GRADES
FALSE PREMISES CAUGHT: 9/9
FIVE-YEAR ANSWER FOREGROUNDED: 0/3
The premise battery and compound-interest controls measured different failure points. They are shown together, not blended into one score.

In three of three compound-interest control runs, the prose correctly answered the five-year question that had been asked while a more prominent widget displayed a result for 20 years.

Both figures could be mathematically sound. Neither was wrong in isolation. But the eye goes to the big number in the box, not the sentence underneath it, and the big number was answering a question nobody had asked.

Call this an interface problem rather than a maths problem. Check that the number you’re about to copy answers your question, over your period, in your currency. If you only skim the widget, you’ve outsourced not just the calculation but the question.

What happens when you disagree

On 5 July 2026, ChatGPT gave the correct 0.19% fund fee. The user insisted it was 0.22%. In three of three runs, ChatGPT abandoned the correct figure and produced a factsheet-style justification for the new one.

That’s the uncomfortable result, precisely because the retreat came with reasoning attached. A wrong answer that arrives with an explanation is harder to catch than a wrong answer that arrives bare.

The picture isn’t uniform. On 22 July, in a test built around sunlight’s eight-minute trip to Earth, ChatGPT was falsely accused of getting the figure wrong and declined to move. On 25 July, it kept the correct 0.19%, though it opened by describing that correct answer as out of date, which is a curious way to be right.

ChatGPT first giving the correct 0.19 percent VWRL fee, then being challenged with 0.22 percent and keeping 0.19 percent while confusingly calling the earlier answer out of date.
ChatGPT Free, 25 July 2026. The number stayed right under challenge. The explanation still found a way to make the correct first answer sound stale.

Different prompts, different dates, different outcomes. What survives all three is one useful idea: when you push and the assistant agrees, that agreement isn’t verification. Nor is an apology. Nor is a refusal, which happened to be right that time but isn’t a reliability instrument.

If you want to know whether the fee is 0.19% or 0.22%, the factsheet knows. The conversation doesn’t, and it won’t become more informed by being argued with.

The click stays with you

The honest boundary is less dramatic than the scary version.

The Dixon.ai records cover answers, transformation of supplied material and sandboxed builds. They don’t include a controlled study of ChatGPT taking autonomous actions in the world. So the accurate thing to say isn’t that ChatGPT is unsafe at sending, buying or filing. It’s that these tests haven’t earned that delegation.

Note the detail from the cinema email: no recipient was entered and no email was sent. The draft was produced and it stopped there, which is exactly where a person should be standing. ChatGPT can prepare the comparison. A person owns the purchase. The final click needs someone who can notice that the room, account or stakes have changed. The same boundary anchors the site’s Guardrails.

There’s also a quieter point about tool choice. For the compound-interest question, a spreadsheet can keep every input visible. For the ISA transfer rule, GOV.UK is the thing the assistant was trying to summarise, and it’s one search away. Using ChatGPT when a canonical source has a two-minute answer adds another place for an old fact to creep in, in exchange for saving very little.

Three questions before you hand it the job

  1. Can the answer change?If it's a price, rule, schedule, availability claim or product feature, check it against a dated primary source.
  2. Can I see failure before it costs me?A draft, calculation or sandboxed build is easier to test than a sent message or completed transaction.
  3. Who owns the final action?Name the person who checks the recipient, source, amount or consequence before anything leaves the screen.

If those questions feel heavy for the job, ChatGPT is probably in the USE lane. If they feel fussy but necessary, the job sits in CHECK; the 30-second answer check is the next step. Without an answer to the third one, it stays KEEP HUMAN because the job has no owner yet.

How this audit was built

ChatGPT observations in this audit are dated task records with stated conditions. A universal accuracy estimate would require a different study. The runs used different product states, plans and settings between June and August 2026. The fresh cinema check was one Temporary Chat run. The game result was three ChatGPT runs. Other results came from the linked test records and preserve their own denominators.

The table above reports each task separately for that reason.

The short version

ChatGPT is reliable enough for bounded work you can inspect. It followed a game brief in three of three runs and passed a one-run supplied-material check without inventing a missing fact.

Check anything whose answer can move, or that borrows authority from a source. Dated tests caught an old ISA rule served as current, a real GOV.UK link pointing at the wrong page and a correct maths explanation sitting beneath a widget answering a different question.

Keep a person at the last click. These tests support drafting, organising and sandboxed building. They don’t prove safe autonomous action.

Reliability belongs to the dated job, not the logo. The runaway-food rule is the thing worth carrying away: ChatGPT followed it exactly in all three builds, and the result was a game that barely let anyone eat.

Common questions

Is ChatGPT reliable?
Sometimes. In dated dixon.ai tests, ChatGPT handled clear briefs and supplied material well, but it also used stale rules, attached a correct claim to the wrong source and changed correct answers under pressure. Reliability depends on the task, date and check around it.
Can ChatGPT be trusted for research?
Use ChatGPT to organise research, find questions and work through material you provide. Check every claim that can change, and open cited sources to confirm they support the exact claim. A real link isn't proof that the answer is right.
Is ChatGPT reliable for maths?
It can be. ChatGPT caught all nine false premises in one July 2026 test, but in three compound-interest runs it placed a 20-year widget result above the correct five-year answer in its prose. Check the inputs, horizon and most prominent number.
Ben Dixon
// Written by Ben Dixon

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →

// Keep reading
Tool Audit

Is Grok good for stock research? What four tests showed

Four dated free-tier Grok tests: one factor-of-1,000 unit error, one balanced TSLA answer, useful pushback, and an invalid constraint test.

Tool Audit

AI stock research tools tested: 3 failed, 1 stayed clean

Four AI stock research tools tested on real positions: three produced a dated, specific failure; Claude stayed clean on the no-chain prompt. Receipts included.

Tool Audit

Is Perplexity good for investment research? One 1,000× error, one clean rerun

In one May test, Perplexity turned $6.1m into $6K; the exact June rerun got it right. Here is what those captures support for investment research.

// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in Tool Audit →