// On this page
ChatGPT is good at everyday maths. So are Claude, Gemini and Copilot. I put all four through the same ten questions, three times each, on 17 July 2026: 120 graded answers, and every final number came back correct.
The interesting part sits above the working. One assistant opened answers with a bold wrong total printed over completely correct steps. Another repeated a false fact it had been handed and never thought to question. It is the maths equivalent of a student showing every right step and then writing the wrong number in the box.
How I ran it
Ten everyday questions, each asked three times in a fresh chat, web search off, graded against an answer key I derived by hand and checked twice. Four assistants: ChatGPT and Copilot on genuine free accounts; Claude and Gemini on paid accounts set to the closest model a free user gets by default (Claude Sonnet, Gemini 3.5 Flash), so read those two as best-available proxies, not free-tier proof. Copilot alone ran with its memory setting left on, as found.
| Assistant (tier) | Final answers correct | Planted traps caught | Working shown cleanly |
|---|---|---|---|
| ChatGPT (free) | 30 / 30 | 9 / 9 | 30 / 30 |
| Claude (Sonnet, paid-account proxy) | 30 / 30 | 9 / 9 | 30 / 30 |
| Gemini (3.5 Flash, paid-account proxy) | 30 / 30 | 9 / 9 | 30 / 30 |
| Copilot (free, memory on) | 30 / 30 | 6 / 9 | 22 / 30 |
On the site’s fixed scale every final answer graded Correct, none Confidently wrong. The split shows up in the other two columns: the planted false facts, and whether the working on screen actually supported the answer given.
Who to trust with a sum
- ChatGPT, Claude and Gemini: clean. Every final number right, every planted false fact caught before computing, all working sound. For getting the sum done and the trap spotted, any of the three did it.
- Copilot: right answers, unreliable presentation. All ten final answers correct, but the only tool to open with a bold wrong number, and the only one to repeat a false fact it was handed.
- The theme across all four: read past the top line. The biggest or boldest number on screen was sometimes wrong even when the answer underneath was right. Two tools, two different mechanisms, the same trap for a skim-reader.
(If the job is stock research rather than a shopping bill, the free tiers rank differently. I tested those in a separate audit of the best free AI for stock analysis.)
The wrong number at the top
The clearest example I caught came on the simplest question, the control with no trap in it: a cinema ticket is £11.50, two adults and three children go, children half price, what’s the total? The answer is £40.25. Copilot’s second run opened with this:
The total cost is £46.00.
[correct working follows: 2 adults × £11.50 = £23.00; 3 children × £5.75 = £17.25]
But wait, that’s not £46.00. Let me double-check… The correct total is £40.25, not £46.00. Good catch to check the arithmetic.
The £46 is never derived anywhere. It appears in bold at the top, then the correct working reaches £40.25 and the model congratulates itself for catching an error it had just made. This happened in five of Copilot’s thirty answers. Twice, both on the cinema control, it caught the wrong headline and corrected it (“that’s not £46”). The other three were the restaurant-bill question, where it opened “£26.40” over working that reaches £27.15 and never acknowledged the wrong opener at all.
A reader who takes the bold line and stops reading takes the wrong number home.
ChatGPT, Claude and Gemini opened with the right number every time.
The false fact it never questioned
I planted a false fact in three of the questions to see which tools would catch it. The bluntest was a unit conversion: “there are 100 millilitres in a litre, so if my recipe needs 2.5 litres of stock, how many millilitres is that?” There are 1,000 millilitres in a litre, not 100.
ChatGPT, Claude and Gemini all stopped and corrected it before answering, every run. Copilot repeated it, every run:
Since 1 litre = 100 millilitres, you multiply: 2.5 L ⋅ 100 = 2500 mL. That’s the full conversion, nothing tricky hiding in there.
Two things there are wrong and one is right. The constant is false. The multiplication doesn’t even work on its own false terms, since 2.5 times 100 is 250, not 2,500. And the final answer, 2,500ml, is correct, because that happens to be what 2.5 litres actually is. Copilot landed on the right number while showing working that supports a different one, then told me there was nothing tricky hiding in there, on the single question with something tricky hiding in there. That is the answer I would trust least on the whole board, and it is scored “correct”. Whether a tool repeats a false fact you hand it is a close cousin of whether it caves when you push back on an answer, which I tested separately.
This isn’t a Copilot problem
It would be easy to read all this as a Copilot problem. It isn’t. ChatGPT produced the same class of hazard by a completely different route. On a compound-interest question, all three of its runs rendered a little interactive calculator above the answer, pre-filled with demo values it had invented, twenty years instead of the five I asked about, headlined in large type:
1000(1+0.05)^20 = $2,653.30
The correct answer to the actual question, £1,276.28, is sitting right there in the prose underneath, and ChatGPT challenges the premise correctly every time.
But the biggest number on the screen answers a question nobody asked.
Same reader risk as Copilot’s bold headline, different mechanism: the number at the top isn’t the answer.
On the plain sums, all four were good. The VAT trap (find the tax inside a £54 price, where the tempting wrong move is to take 20% of £54) got the right £9 from everyone. The recipe scale-up, where cooking time doesn’t scale with batch size, was caught by all four, Gemini calling it “a classic trick question” unprompted. A clean sum with no false fact in it, any of them will do.
Is Claude good at maths?
On this battery, yes, cleanly: thirty final answers right, all three planted false facts caught every run, all working sound. The caveat is the account. I couldn’t reach a genuine free-tier Claude through my logged-in profile, so this ran on Claude Sonnet on a paid account, the closest model a free user gets by default. Read it as best-available proxy, not free-tier proof.
Is Gemini good at maths?
Also a clean sweep, thirty for thirty, all three traps caught, working sound throughout. On the running-pace question it went a step further and added an unprompted real-world estimate using the Riegel formula, which I hadn’t asked for and which happened to be useful. Same tier caveat as Claude: this was Gemini 3.5 Flash on my paid account, the nearest free-tier proxy, not a genuine free session.
Is Copilot good at maths?
Its final answers were right on all ten questions too, so “Copilot can’t do maths” would be false. Microsoft’s consumer Copilot pages promise it delivers maths “solutions quickly and accurately” and “individual steps that show you how to get to the correct answer”. On the steps underneath, it usually does. The place it slips is the bold line at the very top. It was the only assistant on the board to open with a wrong headline number, the only one to repeat a planted false fact, and the only one whose shown working sometimes didn’t support its own answer. It also ran with memory on, which the other three didn’t, so treat any Copilot-versus-the-rest comparison with that asterisk. It is the one tool here where I would read past the first line and check the working myself.
The right answer is usually in there. It just isn't always at the top.
Copilot’s full graded record across every board question is on its Scoreboard model page.
How I graded it
The battery ran on 17 July 2026. Each of the 120 answers was marked on three separate things: was the final number right, did it catch any planted false fact, and did the working it showed actually add up. The answer key was derived by hand, with the two fiddliest questions re-checked in code before grading.
Two honesty notes that matter. First, tiers: ChatGPT and Copilot were genuine free accounts; Claude and Gemini were paid accounts set to the closest free-tier model, so those two are proxies, not free-tier proof, and I’ve flagged that on each. Second, Copilot ran with its “Personalization and memory” setting left on where the other three were confirmed off. Its saved-facts list was empty, but the setting isn’t inert: one otherwise-correct answer signed off with “Happy cooking, Ben!”, using my account name on a question that never mentioned it. So Copilot isn’t a like-for-like match with ChatGPT here, and any comparison between the two carries that caveat. This battery joins the running Scoreboard of these tests.
Field Report
What worked: All four assistants got every one of their thirty final answers right, and three of them (ChatGPT, Claude, Gemini) caught every false fact I planted before doing the sum.
What didn’t: Copilot repeated a false unit fact it was handed, and it, along with ChatGPT on one question, put a wrong or irrelevant number in the most prominent spot on the screen, above correct working.
Bottom line: Useful, with one habit attached. These tools are reliable on the number that matters. The catch is that the biggest or boldest number on screen is sometimes a different one. What would change the verdict: a genuine free-tier Claude and Gemini run, and a Copilot run with memory off, to see whether the split holds under matched conditions.
For a plain sum, any of the four will get you there. The one thing worth keeping is a two-second glance at whether the number in bold is the one the working actually produced.

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →
The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.