This is the kind of thing the Bluff Filter catches. It’s free →
// On this page
Ask Perplexity what colour Yoda’s lightsaber is in the original trilogy and it tells you green. It attaches fifteen sources. It shows you the step where it went and checked.
He never draws one in those three films.
That’s the only question Perplexity got wrong out of six, and it’s the one it was most sure about. Gemini rejected the premise in a sentence and offered no source at all.
Gemini six out of six. Perplexity five. Which sounds like a result until you notice that on the five they both got right, only one of them could show you where the answer came from.
| The question | Gemini | Perplexity | A source you could click |
|---|---|---|---|
| Kettle refund | ✓ | ✓ | Perplexity only |
| Maximum phone fine | ✓ | ✓ | Perplexity only |
| Yoda’s lightsaber | ✓ | ✗ | neither |
| Recipe scaled to six | ✓ | ✓ | neither searched |
| Excel freeze top row | ✓ | ✓ | Perplexity only |
| Bank holidays 2026 | ✓ | ✓ | Perplexity only |
| Total | 6 of 6 | 5 of 6 | Perplexity 5, Gemini 0 |
How I tested
Six everyday questions, 4 August 2026. Fresh chat, memory off, both on paid accounts.
The expected answers were written down before either tool was asked anything. Every answer below is screenshotted in full. Results feed the scoreboard; grading is explained on how we grade.
Six questions, two tools, one right answer each
A refund question where the calendar does the work
"I bought a kettle in person from a UK high-street shop three weeks ago and it has stopped working through no fault of mine. Am I entitled to a full refund?"
Yes. Three weeks is 21 days, and the short-term right to reject runs for 30.
Both led with yes, which is the part that matters. A reader skims the first clause and acts on it. Perplexity went further and said the quiet part out loud: three weeks is inside the window.
Round 1 Draw. Though I have to declare something on this one, below.
One of Perplexity’s fifteen sources was dixon.ai. It cited my own scoreboard page for this exact question, which states the correct answer. The prompt says nothing about AI or testing, so it found us on the merits. That's the first time an assistant has cited us on a question not built to find us, and it also means this round isn't a clean test of what Perplexity knows on its own. Both halves are true and they belong together.
A question with two right answers, only one of which was asked for
"What's the maximum fine for using a handheld phone while driving in the UK? Cite the source."
£1,000 in court for a car, £2,500 for a lorry or bus. The £200 fixed penalty is not the maximum.
This one has a trap in the wording. Most coverage leads with the £200 fixed penalty, but I asked for the maximum, and £200 isn’t it. Both tools got past that. Both gave £1,000 and £2,500 and both drew the distinction between a fixed penalty and a court fine.
Round 2 Draw on the answer. I asked both to cite the source. Perplexity linked gov.uk and the Met. Gemini cited a newspaper and an insurance comparison site alongside it.
The question with nothing to find
"What colour is Yoda's lightsaber in the original trilogy?"
He doesn't have one. Yoda first draws a lightsaber in Attack of the Clones, in 2002.
Gemini said so plainly: Yoda doesn’t actually use a lightsaber in the original trilogy. No source offered, none needed.
Perplexity said green. Then it went and searched, showed me a step reading “checking the original trilogy’s lightsaber color”, and stacked fifteen sources beneath an answer about a thing that doesn’t exist.
Round 3 Gemini. The wrong answer is the one carrying fifteen sources.
Three times now. Two months, two accounts, the same wrong answer.
Arithmetic, where neither of them looked anything up
"A pancake recipe for 4 uses 200g flour, 2 large eggs, 300ml milk and 1 tablespoon of sugar. Rewrite the quantities to serve 6."
300g flour, 3 eggs, 450ml milk, 1½ tablespoons sugar.
Both multiplied by 1.5 and both got all four right, including the awkward one. A tablespoon and a half is where a careless answer rounds to two, and neither did.
It’s also the only question in the set where neither tool searched for anything. Perplexity’s sources pane sat empty. That’s the correct instinct, and worth noticing given what it did in round three.
Round 4 Draw. The one question where searching would have added nothing, and neither of them did.
An exact menu path, and a picture that contradicts itself
"In Microsoft Excel, how do I keep the top row visible while I scroll down a long sheet? Give the exact menu steps."
View, then Freeze Panes, then Freeze Top Row.
Both gave the right path. Perplexity gave it in its opening sentence, which is the whole answer in one line.
Gemini’s answer was also correct, and it came with an illustration. That picture is the interesting part of this round, and it’s not about Excel.
Round 5 Draw on the answer. The picture underneath it's a different matter.
The image in Gemini's answer makes two contradictory claims about where it came from. Its caption, the bit you can see, reads "Source: EDUCBA". Its alt text, the bit only a screen reader gets, reads "Freeze Panes menu in Excel, AI generated". Same picture, same answer, two different origin stories, and which one you're told depends on how you read the page. The Excel instructions were right. I'd still like to know whether that menu is a photograph or a drawing.
A list anyone can check
"How many bank holidays are there in England and Wales in 2026, and what are the dates?"
Eight. The only awkward one is Boxing Day, which moves to Monday 28 December because the 26th is a Saturday.
Both said eight. Both got all eight dates right. Both spotted the substitute day and explained why it moves, which is the only part of this question with any teeth.
The difference is where they went for it. Perplexity cited gov.uk, which publishes the list itself and offers it as a data feed. Gemini cited a company formation website. Twice.
Round 6 Draw on the answer, Perplexity on the source. When the government publishes the list, cite the government.
The verdict
- GeminiWon the board, six from six. The only one to reject the trick question. But every citation it gave was a button that wouldn’t open, so I could not confirm where a single one pointed.
- PerplexityFive from six, and a clickable source on five answers. gov.uk, the Met, Citizens Advice, Microsoft. Its one miss was the question with nothing to find, and it searched anyway.
The gap is one question, and the shape of it is the point. A false premise is the one thing a search engine can’t help with, because there’s nothing to find. Perplexity searched, found pages about Yoda’s lightsaber, and answered from them. Gemini simply knew.
Run your eye down the sourcing instead of the score, though, and the result flips:
- Perplexity, five of six answers carried a live link you could click.
- Gemini, three of six, and all three were buttons rather than links. Where it did name sources, it offered a newspaper and an insurance comparison site for a gov.uk fact, and a company formation website for the government’s own bank holiday list.
So which is more reliable depends entirely on what you do next. If you’re going to act on the answer as given, Gemini edged it. If you’re going to check it, Perplexity is the only one of the two that lets you.
An answer you can’t check isn’t an answer. It’s a guess you agreed with.
What this is and isn’t
One run per question, both accounts paid. That’s a snapshot, not a rate, and I wouldn’t tell you Perplexity gets 83% of things right on the strength of six questions.
If you want the same two tools measured on sourcing rather than accuracy, ChatGPT vs Gemini put six retrieval traps to both, and when a cited page doesn’t say what the answer claims is the failure underneath all of this. When I told them they were wrong tests a confident correction instead of a fact.
The Yoda result is the exception, and only because it’s not from this run alone. It has now happened three times across two months, on a free account and a paid one, in three separate test batches. That’s the one finding here I’d stand behind without re-testing.
One thing I have to declare rather than bury: on the kettle question, one of Perplexity’s fifteen sources was dixon.ai. My page states the correct answer, so that round isn’t a clean test of what Perplexity knows unaided. It’s also the first time an assistant has cited us on a question that wasn’t built to find us, which I am pleased about and which doesn’t make the round any cleaner.
Common questions
- Is Perplexity or Gemini more accurate?
- On six everyday questions run on 4 August 2026, Gemini got six right and Perplexity five. The gap is one question, a trick about Yoda's lightsaber, where Perplexity accepted a false premise and Gemini rejected it. On the other five they agreed and both were correct. One run each, so treat it as a snapshot rather than a rate.
- Which is better, Perplexity or Gemini?
- They split, so it depends which failure costs you more. Gemini was slightly better at getting the fact right. Perplexity was far better at showing you where it got it, with clickable links on five of six answers against none you could open from Gemini. Get the fact from Gemini, get the source from Perplexity.
- Does Gemini cite sources reliably?
- Not in a way you can check. Gemini gave citations on three of six answers, and each one was a button that opens a dialog rather than a link. I could not confirm where a single one pointed. It also cited a company formation website for the 2026 bank holidays, a fact gov.uk publishes itself as an authoritative list.
- Does Perplexity make things up?
- It did not invent anything in this run, but it did confidently answer a question built on a false premise. Asked what colour Yoda's lightsaber is in the original trilogy, it said green and attached fifteen sources. He never draws one in those three films. It has given that answer three times now, across two months and two account tiers.

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →
The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.











