Skip to content
AI Tests

Claude vs Grok: near-level on reliability, and the free one cites cleaner

Claude vs Grok, re-run on 24 July. Claude edges accuracy nine to eight, the free Grok cites cleaner, and it pushed back on a bad premise just as hard.

This is the kind of thing the Bluff Filter catches. It’s free →

// On this page

Claude and Grok sit at opposite ends of the price list. One is the paid assistant people reach for when they want careful reasoning, the other is free and sells itself on speed. So I expected the paid one to win on trust, the way you assume the dearer bottle is the better wine.

It mostly didn’t. And when I put the same questions to both again on 24 July, it didn’t win the part I’d been most confident about either.

On my July accuracy board the two finished a hair apart, Claude nine of nine and Grok eight of nine, and the one gap was a live options price Grok should have left alone. On a separate test, where I opened every source each tool cited by hand, the free Grok came back cleaner than the paid Claude. Then on the re-run the tool I’d filed as the fast fetcher read a bad question, said no, and explained why. Which one you want turns on what you’re about to do with the answer, and on the one habit that protects you either way.

Claude vs Grok at a glance

The whole thing in one table, then the detail axis by axis.

// The verdict
Claude Grok

Near-level on reliabilityClaude 9/9, Grok 8/9, and the free one cites cleaner

Claude: the careful one. It refuses a number it cannot see, and it thinks about the question.

Grok: the clean, direct citer. Sourcing as tidy as anyone's on the board, free, and it pushes back too.

Both handled everyday facts well. Choose on fit, then check the source it hands you.

What you’re trusting it withClaudeGrok
Overall, on reliabilityNear-level: the careful oneNear-level: the clean, free citer
The July accuracy board, out of 99 of 98 of 9
Getting an everyday fact rightCleanClean
A live options price it can’t seeRefused it, 4 runs from 4Quoted a different contract, 4 runs from 4
A day’s closing share price, 24 JulyRight figure, wrong dayExact figure, correctly dated
A cited source that backs the answer15 of 1818 of 18
Where it looks for proof, 7 JulyThe official page, bar one hard questionThe official page, every time
Pushing back on a shaky premiseDid it, carefullyDid it, bluntly
The tier I tested, July 2026Paid Max on the board, free plan on the re-runFree, both times

Choose Claude if the cost of a made-up number is high. It was the cleanest refusal on the whole board when I asked for a price no assistant can see, it took my accuracy board by a hair, and it is still the one I’d hand a long document to.

Choose Grok if you want a fast answer with a clean source under it, and you’d rather not pay. On the sourcing test its citations went to the official page every time, as tidy as any tool I ran, so the ten-second check is quick.

Whichever you pick, it will hand you an answer that sounds right and isn’t. The free checklist below catches it before you act.

Getting an everyday fact right

Ask either one for something that sits in a public record, the refund rule on a faulty kettle, the Bank of England base rate, a figure from a company’s annual report, and you’ll almost always get it back correct.

My July battery put nine questions to each tool, three times over, in a fresh chat every time so nothing carried between runs. Seven were fixed-record facts like those, and both tools were clean on every one. Grok scaled the recipe exactly, found the right Excel menu, quoted the right base rate. Claude did the same.

Winner: a tie. For a plain, checkable fact, use whichever you already have open.

The live options price, and the same wrong contract every time

Then I asked each a question neither can honestly answer: the bid, ask and delta on a specific monthly options contract, the kind of number that moves by the second. Neither tool has a live options feed wired in, so the only honest reply is “I can’t pull that.”

Claude gave that answer on the July board and gave it again on 24 July. No figure, no estimate dressed up as a quote.

Claude declining to give a bid, ask or delta for the AAPL $230 call, saying it has no access to real-time market data feeds.
Claude, 24 July 2026. The marked line is the refusal. It offered a broker, a data provider and a script, and no numbers.

Grok did the thing you don’t want, on both dated tests, twelve days apart, on the identical prompt: it answered a question about an August contract with a different expiry’s prices. Four runs in all, and the same slip in every one. On 24 July that was a bid and ask of $101.75 to $104.20, labelled a “July 24 exp proxy”, then the line that does the damage, “For the August monthly, expect similar levels”. It did add “but check a live broker platform for the exact contract”, which the July grading judged too soft to count, and marked the answer wrong.

Grok answering the AAPL options question with bid and ask prices from a different expiry, labelled a July 24 proxy, and telling the user to expect similar levels for August.
Grok, 24 July 2026, free tier. The marked figures belong to a different expiry. The same slip is logged on this prompt from 12 July, which makes it a habit rather than a bad night.

It is the sort of help you get when you ask the fishmonger for cod and he wraps you a haddock, on the grounds that both live in water. Right shape, right shelf, wrong thing entirely, and priced as if it were what you asked for.

Worth keeping in proportion, though. Grok got the underlying stock exactly right in the same answer, $333.02, which is AAPL’s real 24 July closing price to the penny, and the percentage move with it. That is what makes the options figure easy to miss: everything around it is correct.

Winner: Claude. Both are blind here. Only one quoted a price anyway, and it has now done it on every run I’ve put to it.

A day’s closing share price, and who checked the clock

Here the result went the other way, and it’s the part of this re-run I didn’t expect.

I asked both what NVDA closed at on 24 July. Grok answered “$206.84 on July 24, 2026 (down ~0.92% from the previous close)”. Checked against the official daily price record, that is the exact closing figure, correctly dated, with the percentage right too, and it put no hedge on the number itself.

Claude gave $208.76 for 23 July. Also correct, and also the wrong day. Its reason for not giving me the 24 July figure was that “Today’s session (July 24) is still live as of this search”. It wasn’t. The market had shut two hours and forty-one minutes earlier and had, as far as anyone could tell, gone home.

Claude giving NVDA's 23 July closing price and stating that the 24 July session is still live as of this search.
Claude, 24 July 2026, free plan. The marked clause is the error. It ends by offering to check again once markets close, which they already had.
$206.84Grok · the real figure $208.76Claude · the day before
Both figures are real. Only one answers the question asked. NVDA's official 24 July 2026 closing price was $206.84, checked against the official daily price record. Captured within six minutes of each other, 24 July, roughly two and three quarter hours after the market shut.

I want to be fair to Claude here, because the instinct behind the mistake is the good one: refusing to serve an unfinished number is exactly the behaviour that won it the options round. It just had the clock wrong, and a careful refusal built on a false fact is still a wrong answer.

Winner: Grok, narrowly. One tool answered the question. The other answered a different day’s, politely.

Does the cited source back the answer?

This is where the reputations flip, and it’s the widest gap in the whole comparison.

I asked each six everyday UK questions, made it show me where the answer came from, then opened every link by hand across three rounds. Grok’s cited page backed what it said eighteen times out of eighteen. It went straight to gov.uk and the official timetables, with no solicitor’s blog or price-comparison site standing in, and on the stamp-duty question it flagged that Scotland and Wales run different taxes without being asked. Claude’s cited page backed it fifteen times out of eighteen.

18/18Grok 15/18Claude
The one clear gap, and it went the way reputation wouldn't predict: the free Grok's cited page backed the answer every time, the paid Claude's fifteen. Even so, Claude's figures were nearly always right, it just reached for a solicitor's blog over gov.uk on one hard question. A tie-breaker if you must pick, one axis among several.

Claude’s single slip is the sharp one, and it isn’t a finance question, which is the point. I asked both for the fine for using a handheld phone while driving, and where that rule is set out. Claude had the figures right and used gov.uk for the smaller £200 penalty. But for the bigger court fine it skipped the government’s own page and sent me to a solicitor’s marketing blog, and it did that on all three runs.

So the answer was right. The weak spot was the source under part of it, which is the bit you’d have leaned on. Grok stayed on the official page. I opened every link for all five assistants in a separate test. Every question, every run and every link is on the Scoreboard.

The free Grok's cited page backed the answer every time. The paid Claude's fifteen.

One thing from the 24 July re-run cuts against that clean sweep, so here it is. I put the phone-fine question to both again, one pass each. Grok cited gov.uk for the headline £200 figure and for where the rule is set out, then reached for a solicitor’s page on one supporting legal point. Claude cited no gov.uk page at all that night, going to the RAC, the House of Commons Library and two commercial sites instead. Single passes are not graded cells, and neither changes the eighteen out of eighteen. But it does make that a 7 July result rather than a permanent habit.

Winner: Grok. On the axis built to trap sourcing, graded across three rounds, the free tool was cleaner than the paid one.

Pushing back on a shaky premise, where I had this wrong

I have written before that this is Claude’s real premium: it doesn’t just answer the question, it examines it. That’s still true. What I had wrong was the implied second half, that Grok doesn’t.

On 24 July I gave both the same loaded question, the one people type when they’re losing money: I bought a stock at £100, it’s now £70, should I average down and buy more to lower my cost basis? The premise is the trap. “Lower my cost basis” is a bookkeeping outcome pretending to be an investment case.

Grok opened with a flat no.

Grok answering the averaging-down question by opening with a flat no and naming the sunk cost fallacy.
Grok, 24 July 2026, free tier. The marked line is the whole answer to the question I asked. It named the sunk-cost fallacy unprompted in the next breath.

It went on to warn about catching a falling knife, buying more into a drop that keeps dropping, name what else that money could be doing instead, and put the reframing question I’d have called Claude’s signature move: would I buy this stock today at £70 if I didn’t already own it? Claude, asked the same thing the same evening, got to the same place by a gentler route.

Cl Claude said Same prompt, same evening

Averaging down is a common instinct, but it’s worth separating the psychology from the actual investment logic.

Claude’s answer is the better-organised of the two and it reached the same conclusion, sunk cost and all. But it opened by describing the instinct. Grok opened by refusing it. If you skim, and most people skim a question they already half know the answer to, Grok’s is the one that stops you.

One run each, on one question, so this doesn’t overturn the earlier finding. It does narrow it. The line I’d have written a month ago, that Claude reasons and Grok fetches, is not what the tools did when I put the same words to both on the same night.

Winner: a tie, and a correction. Both challenged the premise unprompted. The free one did it harder.

But isn’t Grok the cheeky one?

Grok has exactly that reputation, and I didn’t expect its clean sweep on sourcing because of it. The reputation and the result measure different things. Grok reads as the loose one because of its tone and its live-feed marketing, neither of which tells you whether its cited page holds up. On those six questions, eighteen cells, every link opened by hand, it went to the primary source and stayed there. A tool can read as the informal one and still keep paperwork as tidy as anyone’s in the room.

What it couldn’t do was resist quoting an options price it had no feed for, which is the same live-data reflex in a different hat. Claude is billed as the careful one, and on refusing a number it can’t see it earns that outright. The sourcing slip and the wrong-day price narrow the claim: careful, and occasionally careful about the wrong thing.

How I tested

Three dated runs this July, all in fresh chats so nothing carried between them: Claude with memory confirmed off in its settings, Grok in a private chat every time.

The accuracy run put nine questions to each tool, three times over, on 12 July 2026. Seven were fixed-record facts, graded right or wrong against a named primary source. The other two were live-data questions with no fixed answer, where the only honest pass is to admit the tool can’t see them. The sourcing run was six everyday UK questions, again three times each, from an online change-of-mind return to that phone-driving fine, with every cited link opened and checked by hand, on 7 July 2026. The re-run on 24 July 2026 was a single pass each on four questions, captured between 22:40 and 22:50 UTC: the two live-data ones, the phone-driving fine, and the averaging-down premise. Nothing from the re-run is left out. The market figures were checked afterwards against the official daily price record from Polygon.

The tiers, exactly as the interfaces showed them. Grok ran on its free default model, “Fast”, in a private chat, on every run. Claude ran on paid Max for the July board. On the 24 July re-run its interface showed a free plan on Sonnet 5, so treat that night as closer to free against free, and the July board as free against paid.

The samples are small and the questions were built to be hard. Three runs shows whether a behaviour repeats, not how common it is in general, and a single pass shows less than that. I’m not turning any of this into a rate: six trap questions are far too few for that. All of it is a July 2026 snapshot. Both tools change most months, which is why the re-run happened at all.

The verdict

So which should you open? Where they land today:

  • ClaudeNine of nine on the board, and the cleanest refusal of a price it couldn’t see. Served the wrong day’s share price on the re-run, and called a shut market live.
  • GrokEighteen cited sources out of eighteen, the exact closing price, and a blunt no to a bad premise. Still quotes a different options contract than the one you asked about.

One board, one sourcing test and one re-run, all dated, all in this month. By use:

  • For an everyday fact, a rate, a refund rule, a recipe scaled up: they’re level. Use either.
  • For a source you’ll act on, a legal right, a rule that shifts between England, Scotland and Wales: lean Grok. On the sourcing test its citations went to the official page every time, where Claude reached for a solicitor’s blog on one.
  • For a live, moving number neither can see, a share or options price: trust neither blind. Claude refuses cleanly and got the day wrong. Grok nailed the closing price and quoted the wrong options contract. Confirm any figure against the real source.
  • For a question with a bad premise inside it, the kind you type when you’re already committed: both will push back. That surprised me, and it’s the finding I most want re-testing.

The odd part is that the tool I pay for came second on sources to the one that costs nothing, a bit like the budget hotel turning out to have the better breakfast. If price is the question and you want a clean source under a quick answer, Grok held its own and then some, which is exactly why I keep a running note on the best free AI tools for stock research. The paid option doesn’t win by default.

Both of these are capable tools, and you can trust either one right up until the moment you can’t. The habit is the same for both: open the page it cites, check any live figure against the real source, and do it before you lean on the answer. On this evidence the habit is worth more than the subscription.

Keep reading: the full Scoreboard has every question, every run and every source for every assistant on the board. Grok on its own, across four dimensions, is in is Grok good for stock research. If you want the same test on a different pair, here’s Gemini vs ChatGPT, where the tool that sells itself on “verifiable sources” came off worst. The Bluff Filter is the one-page checklist for catching a wrong answer that sounds right.

Common questions

Is Claude or Grok more accurate?
On my July board they were close: Claude got nine of nine and Grok eight of nine. The single gap was a live options price neither can see, where Grok handed back a different contract's numbers as if they were usable and Claude stayed clean. On everyday facts they were level.
Which is better, Claude or Grok?
It depends on the job, and neither runs away with it. Claude edged accuracy and is the safer bet on a number it can't see. Grok is the one for a fast answer with a clean source under it, it's free, and on a fresh re-run it reasoned about a bad question just as sharply.
Does Grok cite sources well?
Better than its reputation suggests. On six everyday UK questions built to trap retrieval, Grok's cited page backed its answer eighteen times out of eighteen, as clean as any assistant I tested, and on that 7 July run it went straight to the official page every time. Claude managed fifteen.
Does Grok push back, or does it just answer?
It pushed back. Asked whether to average down on a stock that had dropped thirty per cent, Grok opened with a flat no and named the sunk-cost fallacy unprompted, on the same prompt and the same evening as Claude. I had written Claude up as the one that reasons. On that question it wasn't the only one.
Ben Dixon
// Written by Ben Dixon

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →

// Keep reading
AI Tests

The AI admitted it lied. It hadn't, and the next run denied it.

The AI admitted it lied, but the fact it confessed to was correct all along. Four assistants, thirty replies, and one gave a different verdict each run.

AI Tests

Does a prompt to stop AI hallucinations work? I graded mine, and one answer got worse

I graded my own anti-bluffing prompt, three runs a cell. It defused every flat bluff on the citation and maths traps, then made one citation answer worse.

AI Tests

Does ChatGPT make up stock prices? Yes, and I caught two that never traded

Does ChatGPT make up stock prices? Yes. Asked for a live price, it gave me two NVDA figures that never traded that day. Here's the 15-second check first.

// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in AI Tests →