Skip to content
AI Tests

Perplexity vs Gemini: which is more reliable?

Six everyday questions, both tools, August 2026. Gemini got six right, Perplexity five. The one it missed is the one it was surest about.

This is the kind of thing the Bluff Filter catches. It’s free →

// On this page

Ask Perplexity what colour Yoda’s lightsaber is in the original trilogy and it tells you green. It attaches fifteen sources. It shows you the step where it went and checked.

He never draws one in those three films.

That’s the only question Perplexity got wrong out of six, and it’s the one it was most sure about. Gemini rejected the premise in a sentence and offered no source at all.

Gemini six out of six. Perplexity five. Which sounds like a result until you notice that on the five they both got right, only one of them could show you where the answer came from.

The questionGeminiPerplexityA source you could click
Kettle refundPerplexity only
Maximum phone finePerplexity only
Yoda’s lightsaberneither
Recipe scaled to sixneither searched
Excel freeze top rowPerplexity only
Bank holidays 2026Perplexity only
Total6 of 65 of 6Perplexity 5, Gemini 0

How I tested

Six everyday questions, 4 August 2026. Fresh chat, memory off, both on paid accounts.

The expected answers were written down before either tool was asked anything. Every answer below is screenshotted in full. Results feed the scoreboard; grading is explained on how we grade.

The test · six rounds

Six questions, two tools, one right answer each

Round 1 of 6

A refund question where the calendar does the work

I asked both

"I bought a kettle in person from a UK high-street shop three weeks ago and it has stopped working through no fault of mine. Am I entitled to a full refund?"

Yes. Three weeks is 21 days, and the short-term right to reject runs for 30.

Both led with yes, which is the part that matters. A reader skims the first clause and acts on it. Perplexity went further and said the quiet part out loud: three weeks is inside the window.

Gemini answering that you are legally entitled to a full refund, with two citation chips reading which.co.uk.Gemini, 4 Aug, Pro
Perplexity answering yes and stating that three weeks is inside the 30-day window, with fifteen sources listed.Perplexity, 4 Aug, Pro
Geminiright both led correctly Perplexityright

Round 1 Draw. Though I have to declare something on this one, below.

Worth declaring

One of Perplexity’s fifteen sources was dixon.ai. It cited my own scoreboard page for this exact question, which states the correct answer. The prompt says nothing about AI or testing, so it found us on the merits. That's the first time an assistant has cited us on a question not built to find us, and it also means this round isn't a clean test of what Perplexity knows on its own. Both halves are true and they belong together.

Round 2 of 6

A question with two right answers, only one of which was asked for

I asked both

"What's the maximum fine for using a handheld phone while driving in the UK? Cite the source."

£1,000 in court for a car, £2,500 for a lorry or bus. The £200 fixed penalty is not the maximum.

This one has a trap in the wording. Most coverage leads with the £200 fixed penalty, but I asked for the maximum, and £200 isn’t it. Both tools got past that. Both gave £1,000 and £2,500 and both drew the distinction between a fixed penalty and a court fine.

Gemini giving the maximum fine broken down by vehicle type and by whether the case goes to court.Gemini, 4 Aug, Pro
Perplexity giving one thousand pounds for ordinary drivers and two thousand five hundred for a lorry or bus, with fifteen sources.Perplexity, 4 Aug, Pro
Geminiright both cleared the trap Perplexityright

Round 2 Draw on the answer. I asked both to cite the source. Perplexity linked gov.uk and the Met. Gemini cited a newspaper and an insurance comparison site alongside it.

Round 3 of 6

The question with nothing to find

I asked both

"What colour is Yoda's lightsaber in the original trilogy?"

He doesn't have one. Yoda first draws a lightsaber in Attack of the Clones, in 2002.

Gemini said so plainly: Yoda doesn’t actually use a lightsaber in the original trilogy. No source offered, none needed.

Perplexity said green. Then it went and searched, showed me a step reading “checking the original trilogy’s lightsaber color”, and stacked fifteen sources beneath an answer about a thing that doesn’t exist.

Gemini rejecting the premise and stating Yoda does not use a lightsaber in the original trilogy, with no citations shown.Gemini, 4 Aug, Pro
Perplexity answering that Yoda's lightsaber is green in the original trilogy, with fifteen sources attached.Perplexity, 4 Aug, Pro
Geminiright 15 sources Perplexitywrong

Round 3 Gemini. The wrong answer is the one carrying fifteen sources.

Three times now. Two months, two accounts, the same wrong answer.

Round 4 of 6

Arithmetic, where neither of them looked anything up

I asked both

"A pancake recipe for 4 uses 200g flour, 2 large eggs, 300ml milk and 1 tablespoon of sugar. Rewrite the quantities to serve 6."

300g flour, 3 eggs, 450ml milk, 1½ tablespoons sugar.

Both multiplied by 1.5 and both got all four right, including the awkward one. A tablespoon and a half is where a careless answer rounds to two, and neither did.

It’s also the only question in the set where neither tool searched for anything. Perplexity’s sources pane sat empty. That’s the correct instinct, and worth noticing given what it did in round three.

Gemini scaling the recipe by 1.5 and listing 300g flour, 3 large eggs, 450ml milk and 1.5 tablespoons of sugar.Gemini, 4 Aug, Pro
Perplexity showing the multiplier as six over four equals 1.5 and listing the same four scaled quantities.Perplexity, 4 Aug, Pro
Geminiright neither searched Perplexityright

Round 4 Draw. The one question where searching would have added nothing, and neither of them did.

Round 5 of 6

An exact menu path, and a picture that contradicts itself

I asked both

"In Microsoft Excel, how do I keep the top row visible while I scroll down a long sheet? Give the exact menu steps."

View, then Freeze Panes, then Freeze Top Row.

Both gave the right path. Perplexity gave it in its opening sentence, which is the whole answer in one line.

Gemini’s answer was also correct, and it came with an illustration. That picture is the interesting part of this round, and it’s not about Excel.

Gemini explaining the Freeze Panes feature with an embedded image of the Excel menu.Gemini, 4 Aug, Pro
Perplexity giving the exact path View then Freeze Panes then Freeze Top Row in its first sentence.Perplexity, 4 Aug, Pro
Geminiright same path, both Perplexityright

Round 5 Draw on the answer. The picture underneath it's a different matter.

Worth declaring

The image in Gemini's answer makes two contradictory claims about where it came from. Its caption, the bit you can see, reads "Source: EDUCBA". Its alt text, the bit only a screen reader gets, reads "Freeze Panes menu in Excel, AI generated". Same picture, same answer, two different origin stories, and which one you're told depends on how you read the page. The Excel instructions were right. I'd still like to know whether that menu is a photograph or a drawing.

Round 6 of 6

A list anyone can check

I asked both

"How many bank holidays are there in England and Wales in 2026, and what are the dates?"

Eight. The only awkward one is Boxing Day, which moves to Monday 28 December because the 26th is a Saturday.

Both said eight. Both got all eight dates right. Both spotted the substitute day and explained why it moves, which is the only part of this question with any teeth.

The difference is where they went for it. Perplexity cited gov.uk, which publishes the list itself and offers it as a data feed. Gemini cited a company formation website. Twice.

Gemini listing eight bank holidays for 2026 including the 28 December substitute, citing a company formation website.Gemini, 4 Aug, Pro
Perplexity listing eight bank holidays for 2026 with the Boxing Day substitute day marked, citing gov.uk.Perplexity, 4 Aug, Pro
Geminiright gov.uk vs a company formation site Perplexityright

Round 6 Draw on the answer, Perplexity on the source. When the government publishes the list, cite the government.

The verdict

  • GeminiWon the board, six from six. The only one to reject the trick question. But every citation it gave was a button that wouldn’t open, so I could not confirm where a single one pointed.
  • PerplexityFive from six, and a clickable source on five answers. gov.uk, the Met, Citizens Advice, Microsoft. Its one miss was the question with nothing to find, and it searched anyway.

The gap is one question, and the shape of it is the point. A false premise is the one thing a search engine can’t help with, because there’s nothing to find. Perplexity searched, found pages about Yoda’s lightsaber, and answered from them. Gemini simply knew.

Run your eye down the sourcing instead of the score, though, and the result flips:

  • Perplexity, five of six answers carried a live link you could click.
  • Gemini, three of six, and all three were buttons rather than links. Where it did name sources, it offered a newspaper and an insurance comparison site for a gov.uk fact, and a company formation website for the government’s own bank holiday list.

So which is more reliable depends entirely on what you do next. If you’re going to act on the answer as given, Gemini edged it. If you’re going to check it, Perplexity is the only one of the two that lets you.

An answer you can’t check isn’t an answer. It’s a guess you agreed with.

What this is and isn’t

One run per question, both accounts paid. That’s a snapshot, not a rate, and I wouldn’t tell you Perplexity gets 83% of things right on the strength of six questions.

If you want the same two tools measured on sourcing rather than accuracy, ChatGPT vs Gemini put six retrieval traps to both, and when a cited page doesn’t say what the answer claims is the failure underneath all of this. When I told them they were wrong tests a confident correction instead of a fact.

The Yoda result is the exception, and only because it’s not from this run alone. It has now happened three times across two months, on a free account and a paid one, in three separate test batches. That’s the one finding here I’d stand behind without re-testing.

One thing I have to declare rather than bury: on the kettle question, one of Perplexity’s fifteen sources was dixon.ai. My page states the correct answer, so that round isn’t a clean test of what Perplexity knows unaided. It’s also the first time an assistant has cited us on a question that wasn’t built to find us, which I am pleased about and which doesn’t make the round any cleaner.

Common questions

Is Perplexity or Gemini more accurate?
On six everyday questions run on 4 August 2026, Gemini got six right and Perplexity five. The gap is one question, a trick about Yoda's lightsaber, where Perplexity accepted a false premise and Gemini rejected it. On the other five they agreed and both were correct. One run each, so treat it as a snapshot rather than a rate.
Which is better, Perplexity or Gemini?
They split, so it depends which failure costs you more. Gemini was slightly better at getting the fact right. Perplexity was far better at showing you where it got it, with clickable links on five of six answers against none you could open from Gemini. Get the fact from Gemini, get the source from Perplexity.
Does Gemini cite sources reliably?
Not in a way you can check. Gemini gave citations on three of six answers, and each one was a button that opens a dialog rather than a link. I could not confirm where a single one pointed. It also cited a company formation website for the 2026 bank holidays, a fact gov.uk publishes itself as an authoritative list.
Does Perplexity make things up?
It did not invent anything in this run, but it did confidently answer a question built on a false premise. Asked what colour Yoda's lightsaber is in the original trilogy, it said green and attached fifteen sources. He never draws one in those three films. It has given that answer three times now, across two months and two account tiers.
Ben Dixon
// Written by Ben Dixon

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →

// Keep reading
AI Tests

Perplexity vs Claude: which is more reliable?

Perplexity vs Claude: I re-ran four questions on both. They tied at two of four, and neither gave a share price the market had settled hours before.

AI Tests

Is Grok reliable? I graded its free-tier answers against the source

Four dated tests, graded against primary sources. Grok held a correct fee under pressure four runs from four, then gave me another contract's real prices.

AI Tests

Best AI assistant: I tested five, and only one got everything right

I put four questions to five AI assistants over one weekend in July. Only Gemini got all four right, and it was the one that never showed a source.

// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in AI Tests →