Skip to content
AI Tests

Is ChatGPT good at maths? I graded four AIs on 120 real answers

This is the kind of thing the Bluff Filter catches. It’s free →

// On this page

ChatGPT is good at everyday maths. So are Claude, Gemini and Copilot. I put all four through the same ten questions, three times each, on 17 July 2026: 120 graded answers, and every final number came back correct.

The interesting part sits above the working. One assistant opened answers with a bold wrong total printed over completely correct steps. The same assistant also repeated a false fact it had been handed and never thought to question. It is the maths equivalent of a student showing every right step and then writing the wrong number in the box. Marked, it looks like this:

Cp Copilot · 17 July 2026 · the cinema-ticket control, no trap in it

“The total cost is £46.00.”

its own working, two lines later: £40.25. The £46 is never derived anywhere.

The full sequence, self-congratulation included, is further down.

How I ran it

Ten everyday questions, each asked three times in a fresh chat, web search off, graded against an answer key I derived by hand and checked twice. Four assistants: ChatGPT and Copilot on genuine free accounts; Claude and Gemini on paid accounts set to the closest model a free user gets by default (Claude Sonnet, Gemini 3.5 Flash), so read those two as best-available proxies, not free-tier proof. Copilot alone ran with its memory setting left on, as found.

AssistantFinal answersTraps caughtWorking clean
ChatGPTfree 30/30 9/9 30/30
ClaudeSonnet, paid-account proxy 30/30 9/9 30/30
Gemini3.5 Flash, paid-account proxy 30/30 9/9 30/30
Copilotfree, memory on 30/30 6/9 22/30

On the site’s fixed scale every final answer graded Correct, none Confidently wrong. The split shows up in the other two columns: the planted false facts, and whether the working on screen actually supported the answer given.

Who to trust with a sum

  • ChatGPTThirty of thirty, every planted trap caught. One wobble that isn’t arithmetic: a widget printed above the answer showing a number nobody asked for.
  • ClaudeThirty of thirty, all three planted false facts caught before computing, working sound throughout.
  • GeminiThirty of thirty, all three traps caught, and on one question it volunteered a real-world check I hadn’t asked for.
  • CopilotThirty of thirty final answers, but the only one to repeat a false fact it was handed, and the only one to print a wrong total above correct working.

The theme in this battery: read past the top line. The biggest or boldest number on screen was sometimes wrong even when the concluding answer underneath matched the key. Two tools, two different mechanisms, the same trap for a skim-reader.

(If the job is stock research rather than a shopping bill, the free tiers rank differently. I tested those in a separate audit of the best free AI for stock analysis.)

The wrong number at the top

The clearest example I caught came on the simplest question, the control with no trap in it: a cinema ticket is £11.50, two adults and three children go, children half price, what’s the total? The answer is £40.25. Copilot’s second run opened with this:

Cp Copilot said Confidently wrong

The total cost is £46.00.

correct working follows, and it is right: 2 adults × £11.50 = £23.00; 3 children × £5.75 = £17.25

But wait — that’s not £46.00. Let me double-check… The correct total is £40.25, not £46.00. Good catch to check the arithmetic.

The £46 is never derived anywhere. It appears in bold at the top, then the correct working reaches £40.25 and the model congratulates itself for catching an error it had just made. This happened in five of Copilot’s thirty answers. Twice, both on the cinema control, it caught the wrong headline and corrected it (“that’s not £46”). The other three were the restaurant-bill question, where it opened “£26.40” over working that reaches £27.15 and never acknowledged the wrong opener at all.

A reader who takes the bold line and stops reading takes the wrong number home.

ChatGPT, Claude and Gemini opened with the right number every time.

5of Copilot’s 30 answers led with a bold wrong number
Twice on the cinema-ticket control, caught and corrected a line later. Three times on the restaurant bill, where the wrong opening total was never acknowledged at all.

The false fact it never questioned

I planted a false fact in three of the questions to see which tools would catch it. The bluntest was a unit conversion: “there are 100 millilitres in a litre, so if my recipe needs 2.5 litres of stock, how many millilitres is that?” There are 1,000 millilitres in a litre, not 100.

ChatGPT, Claude and Gemini all stopped and corrected it before answering, every run. Copilot repeated it, every run:

Cp Copilot said Repeated a false fact

Since 1 litre = 100 millilitres, you multiply: 2.5 L ⋅ 100 = 2500 mL. That’s the full conversion — nothing tricky hiding in there.

Two things there are wrong and one is right. The constant is false. The multiplication doesn’t even work on its own false terms, since 2.5 times 100 is 250, not 2,500. And the final answer, 2,500ml, is correct, because that happens to be what 2.5 litres actually is. Copilot landed on the right number while showing working that supports a different one, then told me there was nothing tricky hiding in there, on the single question with something tricky hiding in there.

That is the answer I would trust least on the whole board, and it is scored “correct”. Whether a tool repeats a false fact you hand it is a close cousin of whether it caves when you push back on an answer, which I tested separately.

This isn’t a Copilot problem

ChatGPT produced the same class of hazard by a completely different route, so this is not only a Copilot problem. On a compound-interest question, all three of its runs rendered a little interactive calculator above the answer, pre-filled with demo values it had invented, twenty years instead of the five I asked about, headlined in large type:

Ch ChatGPT's widget showed Answers a different question

1000(1+0.05)^20 = $2,653.30

The correct answer to the actual question, £1,276.28, is sitting right there in the prose underneath, and ChatGPT challenges the premise correctly every time.

$2,653.30the widget’s number, at 20 years £1,276.28the actual answer, at 5 years
Both numbers came from the same three runs. The big one at the top solves a compound-interest problem nobody asked. The right one sits in the prose underneath it.

But the biggest number on the screen answers a question nobody asked.

Same reader risk as Copilot’s bold headline, different mechanism: the number at the top isn’t the answer.

On the plain sums, all four were good. The VAT trap (find the tax inside a £54 price, where the tempting wrong move is to take 20% of £54) got the right £9 from everyone. The recipe scale-up, where cooking time doesn’t scale with batch size, was caught by all four, Gemini calling it “a classic trick question” unprompted. A clean sum with no false fact in it, any of them will do.

Is Claude good at maths?

On this battery, yes, cleanly: thirty final answers right, all three planted false facts caught every run, all working sound. The caveat is the account. I couldn’t reach a genuine free-tier Claude through my logged-in profile, so this ran on Claude Sonnet on a paid account, the closest model a free user gets by default. Read it as best-available proxy, not free-tier proof.

Is Gemini good at maths?

Also a clean sweep, thirty for thirty, all three traps caught, working sound throughout. On the running-pace question it went a step further and added an unprompted real-world estimate using the Riegel formula, which I hadn’t asked for and which happened to be useful. Same tier caveat as Claude: this was Gemini 3.5 Flash on my paid account, the nearest free-tier proxy, not a genuine free session.

Is Copilot good at maths?

Its concluding answers matched the key on all thirty runs, so this battery does not support “Copilot can’t do maths”. Its shown working was sound on 22 of 30 responses; the other eight split between five headline/body mismatches and three false-premise calculations that did not compute on their own terms.

It was the only assistant on this board to open with a model-authored wrong headline number, the only one to repeat a planted false fact, and the only one whose shown working sometimes did not support its conclusion. It also ran with memory on, which the other three didn’t, so treat any Copilot-versus-the-rest comparison with that asterisk. It is the one tool here where I would read past the first line and check the working myself.

1answer signed off “Happy cooking, Ben!”, unasked
Copilot ran with memory left on, as found, and its saved-facts list was empty. The setting still wasn’t inert.

The right answer is usually in there. It just isn't always at the top.

Copilot’s full graded record across every board question is on its Scoreboard model page.

How I graded it

The battery ran on 17 July 2026. Each of the 120 answers was marked on three separate things: was the final number right, did it catch any planted false fact, and did the working it showed actually add up. The answer key was derived by hand, with the two fiddliest questions re-checked in code before grading.

Two honesty notes that matter. First, tiers: ChatGPT and Copilot were genuine free accounts; Claude and Gemini were paid accounts set to the closest free-tier model, so those two are proxies, not free-tier proof, and I’ve flagged that on each. Second, Copilot ran with its “Personalization and memory” setting left on where the other three were confirmed off. Its saved-facts list was empty, but the setting isn’t inert: one otherwise-correct answer signed off with “Happy cooking, Ben!”, using my account name on a question that never mentioned it. So Copilot isn’t a like-for-like match with ChatGPT here, and any comparison between the two carries that caveat. This battery joins the running Scoreboard of these tests.

The short version

What worked: All four assistants got every one of their thirty final answers right, and three of them (ChatGPT, Claude, Gemini) caught every false fact I planted before doing the sum.

What didn’t: Copilot repeated a false unit fact it was handed, and it, along with ChatGPT on one question, put a wrong or irrelevant number in the most prominent spot on the screen, above correct working.

Bottom line: Useful, with one habit attached. In this battery every concluding number matched the key, but that did not make every response reliable. The biggest or boldest number was sometimes a different one, and a correct conclusion could sit on broken working. What would change the verdict: a genuine free-tier Claude and Gemini run, and a Copilot run with memory off, to see whether the split holds under matched conditions.

Across these ten question types, all four reached the keyed concluding number in every run. The habit worth keeping is a quick check that the number in bold is the one the working actually produced.

Common questions

Can ChatGPT do maths without making mistakes?
On this battery, its final answers were flawless: thirty out of thirty, and it caught every false fact I planted. The mistake it did make was presentational. On a compound-interest question it drew a calculator above the answer, pre-filled with its own demo values, headlined with a number that answered a question I never asked.
Which AI is best at maths?
For a plain sum, all four I tested got there. ChatGPT, Claude and Gemini were the cleanest: every final answer right, every planted false fact caught, working that supported the answer. Copilot got every final answer right too, but it was the only one to repeat a false fact and the only one to print a wrong total above correct working.
Why does an AI show correct working but the wrong answer?
This battery does not reveal the internal cause. It shows the visible mismatch: Copilot opened five answers with a bold total that appears nowhere in its own steps, then reached the keyed figure underneath. Check that the prominent number is actually supported by the working.
Should I check an AI's maths?
Yes. Check the top line against the working. In these 120 graded answers, every concluding number matched the key, but the most prominent number was sometimes different and eight Copilot responses had broken working. If no working is shown, ask for it and independently calculate anything consequential.
Ben Dixon
// Written by Ben Dixon

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →

// Keep reading
AI Tests

ChatGPT health advice: I tried to make an AI repeat a poisoning

A man was hospitalised after swapping table salt for sodium bromide. I put the same swap to five AI assistants, with the sixty-word window written first.

AI Tests

An AI told me three weeks was more than thirty days

A kettle died after three weeks. UK law gives you 30 days to demand a refund. Four assistants said yes. One opened by telling me I'd missed the window.

AI Tests

Can AI build a game? Four tried, and built one nobody can win

Can AI build a game? I gave ChatGPT, Claude, Gemini and Grok the same Snake brief. Twelve games, one nobody can win, and four dead Start buttons.

// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in AI Tests →