Skip to content
AI Tests

Best AI for math: which is most reliable with numbers?

What's the best AI for math? On everyday sums the assistants are level. The real test is the number that looks like a sum and isn't, and who catches it.

This is the kind of thing the Bluff Filter catches. It’s free →

// On this page

Ask an AI to multiply a few numbers and it will almost always get them right. That part is settled. So “what’s the best AI for math” turns out to be the wrong question, because on the sums most people type, there isn’t a losing model.

The interesting bit is where the maths only looks like maths. A recipe cooking time. A running total that resets. A number sitting in a list that shouldn’t scale with the rest. That’s where the tools quietly split, and where a confident wrong answer looks exactly like a confident right one. I put the same problems to the main assistants across two dated tests, and the pattern held: level on the basics, and the differences came down to method.

Best AI for maths at a glance

The whole thing in one picture, then the detail. First, the easy case: I asked each tool to scale a recipe from four servings to six, a plain multiply-by-1.5, and checked every quantity against the arithmetic.

SCALE A RECIPE FROM 4 SERVINGS TO 6 (×1.5)
ChatGPT
Claude
Gemini
Perplexity
Grok
Captured 26 June 2026, three runs each in fresh chats (ChatGPT on the one run that finished before a technical limit cut it off). Every tool gave 300g flour, 3 eggs, 450ml milk, 1.5 tbsp sugar. On plain arithmetic there is no wrong choice. The full record is on the Scoreboard.

So if they all pass, why does this post exist? Because the sum above is the control case. The table below is where the real question lives.

The mathsHow reliableWhat to watch
OverallLevel on the basics; the method decides the restShow the working, don’t just trust the answer
Plain arithmetic (scaling, converting, %)All five correctNothing much; check it yourself in ten seconds
Multi-step / word problemsThis is where they split: some catch the trap, some don’tThe number that looks like it scales and doesn’t
A figure it can’t see (a live price)Don’tSome invent a number and cite it
Reasoning about the mathsSplits by tool and by monthWhether it explains why, and that today’s winner may not be next month’s

If the maths is a plain calculation you could check yourself in ten seconds, a recipe, a unit conversion, a percentage: use whatever you already have open. They are level, and two of them (ChatGPT and Grok) are free.

If the maths is multi-step, or you’ll act on the answer: slow it down with the four-line prompt below. That is where the tools differ, and where the one that shows its working earns its place.

Whichever you use, it will hand you a confident wrong number eventually. The free checklist below catches it before you act on it.

Plain arithmetic: every tool I tried got it right

Start with the good news, because it’s real. On a clean calculation, the main assistants are dependable.

The recipe test above put the same scaling problem to all five, three times each, in a fresh chat every time so nothing carried between runs. Multiply four quantities by 1.5. Every tool returned 300g flour, 3 eggs, 450ml milk and 1.5 tablespoons of sugar, on every run. Perplexity even threw in a “1 tablespoon plus 1.5 teaspoons” conversion to save you doing it at the counter, which is the sort of small kindness you don’t expect from a machine.

The baseline here is boring but important. If a tool fumbled multiplying four numbers by 1.5, you’d know to trust nothing harder. None did. So for the maths most people ask, scaling, splitting a bill, working out a discount, the honest answer to “which AI is best” is: the one already open on your phone.

The verdict here: a tie. There is no wrong choice on a plain sum.

The trap: the number that only looks like a sum

Here is where it gets interesting. I gave all five assistants a recipe that serves four (300g flour, 2 eggs, 450ml milk, a teaspoon of baking powder, twenty minutes to cook) and asked each to scale it up to nine. Three of those numbers scale cleanly. The cooking time is a trap: pancakes cook in batches in one pan, so each still needs the same time on the heat whether you’re making four or forty. What grows is the number of goes at the pan, not the minutes. I first ran this in June; because these tools change under you, I re-ran the whole thing fresh on 24 July 2026, a new chat per tool with memory off. On the same day, the five split three ways.

Two handed me a wrong number with full confidence. Perplexity and Grok both multiplied the cooking time along with the flour and told me nine servings would take 45 minutes. Perplexity stated it as flatly as it stated the flour:

Total cooking time: 45 minutes

Perplexity's answer scaling the pancake recipe from 4 to 9 servings, captured 24 July 2026, listing the total cooking time as 45 minutes with no caveat.

Grok went further and justified the mistake, which is worse: it dressed a wrong answer as reasoned fact. It gave 45 minutes and said the figure had been “scaled proportionally to the increased batch size”, stating the wrong relationship as if it were settled.

Grok's answer scaling the same recipe, captured 24 July 2026, giving 45 minutes total and stating it was scaled proportionally to the batch size.

It won’t take 45 minutes. Follow that figure and you’d set a timer for more than double the real time. The tools treated the cooking time like an ingredient. It isn’t one.

3 of 5led with a cooking time that shouldn't scale, one of them right after naming the trap
Re-run 24 July 2026, five assistants, fresh chat each, memory off. Perplexity and Grok stated 45 minutes flat; Gemini named the trap then headlined about 45 minutes anyway; ChatGPT and Claude held it at ~20. The arithmetic was right; the setup was wrong. Original June run: I asked 4 AIs to scale a recipe.

One told me the right rule and then broke it anyway, the most revealing answer of the five. Gemini opened by naming the trap unprompted: “cooking time does not scale linearly with batch size.” Exactly right. Then its own headline figure, under “Single Pan (Batches)”, was about 40–45 minutes, the very number the principle it had just stated says is wrong. The correct answer, roughly 20 minutes, appeared only as a secondary “if you use two pans” option the recipe never called for. Read plainly, Gemini’s answer to “how long does this take now” is still 45 minutes, with the right reasoning sitting one line above it, ignored.

Gemini's answer scaling the same recipe, captured 24 July 2026, stating cooking time does not scale linearly and then giving about 40 to 45 minutes as its headline single-pan figure.

Two caught it cleanly. ChatGPT and Claude both held the cooking time at about 20 minutes and, unprompted, explained why: the extra batter means more goes at the pan, not a longer cook. Claude put it plainly:

With more batter you’ll just be cooking more batches (about 2.25× as many pancakes) to get through it all.

The arithmetic was right. The setup was wrong. And on the same day, the same question, some tools told me which was which and some didn't.

And here is the part that matters more than any single result: this trap doesn’t stay caught. Gemini has bounced across it. In June it named the problem in passing and then headlined 45 minutes anyway; in a mid-July battery it caught the same trap cleanly and called it “a classic trick question”; today it states the rule and still leads with the wrong number. Those runs weren’t a controlled like-for-like: Gemini’s tier shifted across them (Pro, then 3.5 Flash, then Flash) and July used a differently-worded recipe. But the shape is the point: the same tool and the same class of trap, caught one week and missed the next. Perplexity, for its part, has given the flat wrong answer every time I’ve asked, June and today. So “which AI is best at this” has no stable answer. The habit that catches it does.

The verdict here: don’t trust the tally, trust the method. ChatGPT and Claude got it today; that is no guarantee they will next month, and no guarantee Perplexity or Grok won’t. The prompt below is what protects you either way.

The four-line prompt that catches it

When I first ran this trap back in June, I also ran it a second way, same tools, same day as that run, with one change: the prompt told each tool to slow down and check its own work. This is the part worth copying, because it’s the fix that outlasts whichever tool happens to be failing this month.

// The four-line prompt

A pancake recipe serves 4 and uses 300g flour, 2 eggs, 450ml milk, 1 tsp baking powder, and cooks for 20 minutes total. I need it to serve 9. Do four things: (1) state the scaling factor and show the arithmetic for EACH ingredient; (2) flag anything that does NOT scale linearly (e.g. egg counts must be whole, cooking time per batch, pan size); (3) give me one quantity I can sanity-check myself in ten seconds; (4) give the final scaled list with a confidence level (low/medium/high). Don’t just multiply blindly.

It’s the same idea for any maths question. You’re asking the tool to show its working and own up to what it’s unsure about. Swap in your own numbers and the shape holds: show the arithmetic step by step, flag anything that doesn’t scale cleanly, give me one figure I can check myself, and rate your confidence.

The prompt changed the answers across the board. Perplexity, which had flatly stated 45 minutes, reversed itself and told me not to rely on a straight multiplication. ChatGPT tightened its vague hedge into a clean line: cooking time per pancake stays the same. Claude went deepest, rating each part of the recipe by confidence and telling me to cook to doneness and ignore the 45-minute timer.

One honest caveat: it improved Gemini but didn’t fully fix it. Even after being told to flag what doesn’t scale, Gemini still framed 45 minutes as a problem to solve with a bigger pan. The prompt narrowed the error without erasing it. So it does most of the work: it closes most of the gap and, more usefully, tells you where the tool is unsure so you can catch the rest yourself.

The verdict here: the prompt beat the tool swap. Four lines did more for the answer than picking a different model would have.

What no model can do for you

Two limits worth stating plainly, because “best AI for math” is a broad promise and my evidence is narrower than the phrase.

First, I tested everyday arithmetic and one scaling word problem. I haven’t put calculus, a statistics proof or a long multi-variable problem to these tools, so I can’t tell you which is best at those. What I can tell you is the failure mode, and it generalises: the danger is the number that doesn’t behave like the others, handed over just as confidently as the ones that do.

Second, some of these tools will work off a number they can’t really see. Ask for something that depends on a live share price and one or two will reach for a figure and sometimes get it wrong, when the honest move would be to say “I can’t see that”. On a separate test one tool gave an intraday high as if it were the day’s closing price, on every run: the right kind of number, the wrong field, with a citation attached. Any calculation resting on it would be wrong, and nothing about it would look wrong. The weak point is the input the tool started from. Every question and every run for that is on the Scoreboard, and it sits alongside the other ways a confident answer goes wrong in the nine ways AI gets it wrong.

The verdict here: check the working, whichever tool you use. The arithmetic is the reliable part. The setup and the inputs are where you earn your keep.

How I tested

Dated tests across this summer, all in fresh chats with memory off so nothing carried between runs.

The plain-arithmetic run put a recipe-scaling question to all five assistants (ChatGPT, Claude, Gemini, Perplexity and Grok), three times each, on 26 June 2026, graded against the arithmetic. ChatGPT completed one of its three runs before a technical limit cut it off; that run was correct, so it’s marked partial on the board.

The cooking-time trap, the four-to-nine scaling, I first ran on four tools on 19 June 2026, and then re-ran fresh on all five on 24 July 2026, a new chat per tool, each on its default model (ChatGPT free; Claude on Sonnet 5, its free default; Gemini on Flash; Perplexity on “Best”; Grok on “Fast”). Those defaults aren’t a level footing with a top-tier paid model, and the screenshots above are from that 24 July run. I’ve kept the June figures where they add the before-and-after; the tally on screen is today’s.

Small samples, questions built to catch a specific kind of error, snapshots dated across the summer. Three runs tells you whether a behaviour repeats; it can’t tell you how often it happens in general. And these are everyday-arithmetic tests, a long way from a maths exam.

The verdict

So which should you open? It depends on the maths, and the split is clean.

  • For a plain sum, a recipe, a percentage, a unit conversion: use whichever you have open. They’re level, and ChatGPT and Grok are free.
  • For anything multi-step, or a number you’ll act on: use the four-line prompt and check one figure yourself. On my 24 July run ChatGPT and Claude caught the trap unprompted and Perplexity and Grok didn’t, but that split had moved the month before and may move again, which is exactly why the prompt matters more than the pick.
  • For a calculation off a live, moving number none of them can see, a share or options price: trust none of them blind. One or two will reach for a figure and get it wrong before they’ll admit they can’t see it. Confirm the number against the real source first.
  • If cost is the deciding factor: the free tiers handle everyday sums fine, which is exactly why I keep a running note on the best free AI tools for stock research.

The honest headline is that the best AI for maths is mostly a tie, and the thing that protects you is a habit rather than a brand: make the tool show its working, flag what doesn’t scale, and sanity-check one number before you act. Do that and it barely matters which one you picked. Skip it and the tidiest, most confident wrong answer is the one you’ll never think to question.

Keep reading: the full Scoreboard has every question, every run and every source for all five assistants. The recipe test in full is here. For a head-to-head on the two strongest tools, there’s Claude vs ChatGPT. And the Bluff Filter is the one-page checklist for catching a wrong answer that sounds right.

Common questions

Is there a best AI for math?
For everyday arithmetic, no single one wins. On a dated test all five assistants I tried scaled a recipe correctly. They only pull apart on maths that hides a trap: a multi-step question, or a number that looks like it scales and doesn't. There the tool that shows its working beats the one that just answers.
Is ChatGPT good at math?
At plain sums, yes, and so are the others. It got the recipe scaling right. Its weak spot is the same as everyone's: a number dressed up as a simple multiplication when it isn't. On the recipe cooking-time trap, re-run on 24 July 2026, ChatGPT caught it cleanly and explained why the cooking time doesn't scale, as did Claude; Perplexity and Grok both gave a confident wrong figure of 45 minutes. That split changes month to month, so asking the tool to show its working matters more than the brand.
Does AI get maths wrong?
Rarely on a clean calculation, often enough on a word problem to matter. The failure is quiet: it multiplies everything by the same factor, including the one quantity that shouldn't scale, and states the wrong figure with the same confidence as the right ones. That is the number to check, and a four-line prompt gets the tool to flag it for you.
Can I trust an AI to do my maths for me?
For a sum you could check yourself in ten seconds, yes. For anything multi-step or that you'll act on, ask it to show its working and sanity-check one figure yourself. The arithmetic is usually fine. It's the setup, the number that doesn't behave like the others, where a confident answer can be wrong.
Ben Dixon
// Written by Ben Dixon

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →

// Keep reading
AI Tests

Is ChatGPT good at maths? I graded four AIs on 120 real answers

I put four AI assistants through 120 graded everyday sums. Every final answer was right. The mistakes were sitting above the working.

AI Tests

Gemini vs ChatGPT: which one can you actually trust?

Which is better, Gemini or ChatGPT? I tested both. They're level on getting facts right, but one kept pointing me to sources that didn't back its answer.

AI Tests

Claude vs ChatGPT: level on reliability, split on character

Is Claude better than ChatGPT? I tested both against the source. They're level on reliability, both 9 of 9 on accuracy. The gap is character, not trust.

// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in AI Tests →