This is the kind of thing the Bluff Filter catches. It’s free →
// On this page
“Today’s session (July 24) is still live as of this search … so today’s official close isn’t final yet.”
The traders had been home for nearly three hours. The figure it gave me was real, and correct, and Thursday’s. The refusal to commit was well formed. The reason it gave for refusing was not.
I put four questions to Claude and Perplexity on the same evening in July, one run each, both screenshotted. They tied. Two of four each. On the graded board Claude leads this pairing comfortably, so a tie is not good news for it: the gap closed because Claude came down, not because Perplexity came up.
| The question | Perplexity | Claude | Who showed a source |
|---|---|---|---|
| The fine for using a phone at the wheel | ✓ | ✓ | both, neither on gov.uk |
| Yoda's lightsaber (a trick question) | ✗ | ✗ | Perplexity only |
| NVDA's closing price | ✗ | ✗ | Perplexity only |
| An AAPL options quote (no live feed exists) | ✓ | ✓ | Perplexity only |
| This pass, 24 July | 2 of 4 | 2 of 4 | 4 of 4 vs 1 of 4 |
| The graded board, three runs a question | 6 of 9 | 9 of 9 | 14 of 18 vs 15 of 18 |
How I tested
Four questions, the evening of 24 July 2026. Fresh chat each time, Claude on a free plan showing Sonnet 5 Medium, Perplexity on a free plan in incognito.
The right answers were written down first. Every answer below is screenshotted in full, uncropped, plan badge in frame. The three-run graded boards are on the scoreboard; grading is explained on how we grade.
Four questions, two tools, one right answer each
A fact they both knew, from pages neither should have used
"What is the fine for using a handheld phone while driving in the UK, and where is that set out?"
£200 and six points, rising to £1,000 in court, or £2,500 for a lorry or bus. Regulation 110 of the 1986 construction and use regulations.
Both had £200 and six points. Claude went further and gave the court maximums; Perplexity said only that the fine “can be higher”, which is true and not much use to anyone deciding whether to contest a ticket.
The interesting half is underneath. Neither one cited gov.uk. Claude sent me to the RAC, a Commons Library briefing and two commercial sites. Perplexity cited the Met Police and the same Commons Library page. Both are real, both are more credible than a marketing blog, and neither is the page that actually sets the fine.
Round 1 Claude, narrowly. It gave the maximum I asked for. Neither sent me to the page that sets it.
A question with nothing to find
"What colour is Yoda's lightsaber in the original trilogy?"
He never draws one. Yoda's first is in Attack of the Clones, in 2002, and it's green.
The premise is false and it’s buried, so a tool that answers the question as asked doesn’t notice. Neither noticed. Claude said green, “first seen in The Empire Strikes Back and again in Return of the Jedi”, and he draws one in neither film.
Perplexity said green too, cited StarWars.com, and added that this “matches how it appears in Return of the Jedi”. The citation is real and the scene is not, which is the harder failure to catch of the two: a reader who does the responsible thing and clicks the source still comes away wrong.
Round 2 No winner. Twelve days earlier, on a paid tier, Claude had opened by calling this a trick question.
Claude was on a free plan showing Sonnet 5 Medium that night, and its earlier catch was on a heavier paid model. The plan badge sits in the frame above for that reason. So round two is a result about a tool on a tier on a day, and not a claim that Claude has got worse.
A number the market had already settled
"What did NVDA close at today, and what's its current 30-day implied volatility?"
$206.84, on Friday 24 July 2026. The market had shut at 20:00 UTC; I asked at about 22:41.
Neither had it. They failed differently, and the difference is the whole point of the round. Claude gave $208.76, which is a real figure, correct, and Thursday’s. Then it explained itself: today’s session was “still live as of this search”. It had ended two hours and forty-one minutes before.
Perplexity gave $202.69 with no date at all, a number matching neither day and sitting below both days’ ranges, noting underneath that the price “comes from a live quote source”. A live quote is not a settled close, and on a Friday night there was no live quote to have.
Round 3 No winner, and Claude's is the worse miss. A right number with a false reason attached is harder to catch than a bare wrong one.
A right number with a false reason attached is harder to catch than a bare wrong one.
A number nobody can see
"What's the current bid, ask and delta on the AAPL monthly $230 call expiring next month?"
There isn't one to give. Neither has a live options feed, so the only right answer is to say so.
Both said so, and this is the round where the pairing looks good. Claude gave no numbers at all and explained why searching wouldn’t rescue it: options chains aren’t indexed live, so any page it could reach would be stale. It named the platforms that carry the real quote.
Perplexity ran its search first, cited what it found, then said the sources were delayed or incomplete for that exact contract. Perplexity had failed this question on two of its three graded runs, once with a confident quote for the wrong strike, so a clean abstention here is the single most encouraging thing in the test.
Round 4 Draw, and the best round of the four. Both refused to invent a number, which is the behaviour that matters most.
The verdict
Claude is still the one I’d trust with a question that matters, and this pass is a warning rather than a reversal. Over three runs a question it went nine of nine to Perplexity’s six. Over one evening it went two of four, and so did Perplexity.
- ClaudeThe one to think with, and it had a bad night. It gave the court maximum I asked for and refused the price it couldn't see. It also told me a closed market was open, and walked into a trick question it had caught twelve days earlier.
- PerplexityThe live-search specialist, and it showed its working. It put a source under all four answers where Claude managed one, and abstained cleanly on the options price for the first time. It also produced a share price belonging to no session, and a real citation for a scene that never happened.
Which one to use, by job
Here’s how I’d split the two after marking this pass:
- A fact you’ll act on: Claude, on the graded record, and open the page it cites yourself, because on this pass it didn’t reach for a government one.
- The current web pulled together and cited: Perplexity, the job it’s built for.
- A number that moves: neither, on this evidence. Both failed the settled price, in opposite directions.
- A question you suspect is loaded: Claude, but not on a free tier on the strength of this round.
The tie is the finding, and it went the wrong way to be reassuring.
What this is and isn’t
I ran one pass per question: four questions, eight answers, all on one evening. That’s a snapshot, not a rate, and no percentage would mean anything at this size.
The tiers weren’t level with the graded boards. Both tools ran free here; the graded batteries didn’t, and Claude’s earlier trick-question catch was on a heavier paid model. So the 24 July pass is a dated spot-check against the same questions, not a rematch on the same terms.
The three-run boards behind the nine-of-nine and six-of-nine figures are on the scoreboard, with every run and every link. For the same four questions put to five assistants at once, including these two, there’s best AI assistant. The sourcing board is written up in when the source doesn’t back the answer, and the Bluff Filter is the one-page checklist for catching an answer that sounds right and isn’t.
The sharpest miss of the evening was also the simplest question: what a share cost once the week had finished. Everything else is detail.
Common questions
- Is Claude or Perplexity more accurate?
- On my graded battery, run three times a question, Claude got nine of nine correct against Perplexity's six of nine. On a fresh four-question pass on 24 July 2026 they finished level at two of four each, because Claude came down rather than Perplexity coming up. Over three runs Claude is ahead. On one evening, they weren't.
- Which is better, Perplexity or Claude?
- It depends on the job. Claude is the one to think with: analysis, long documents, a fact you'll act on. Perplexity is the live-search specialist, so reach for it when you want the current web pulled together and cited on the spot. Neither could be trusted with a moving number on the night I tested them.
- Did the result change when you re-tested?
- Yes, and it's the most useful thing in the test. Both abstained cleanly on the options question, Claude as it had all along and Perplexity for the first time. But neither gave the day's actual closing price, and both missed a trick question Claude had caught twelve days earlier on a heavier paid model. The graded scoreline stands; on the day, the gap had closed downwards.
- Is Perplexity good for research?
- For retrieval, yes. It searches the live web and cites current sources as it goes, which is useful for a quick fact on a well-covered topic. Check the figure it pulls and open its link, and don't expect it to reason about your situation the way Claude does.

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →
The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.







