This is the kind of thing the Bluff Filter catches. It’s free →
// On this page
Both tools will answer almost anything you ask, and both sound sure of themselves whether they’re right or guessing. Sounding sure isn’t the same as being right, so I checked the answers myself.
I put the same everyday questions to ChatGPT and Gemini, several times each, then opened every recoverable URL by hand to see whether the page actually backed the answer. An opaque source label with no destination failed at the earlier step: there was no exact page to inspect. I mixed in plain facts, a live options quote, a trick question with a false premise buried in it, and the kind of thing you’d act on: a legal right, a fine, a rule that shifts between England, Scotland and Wales.
Here’s how they compare across the four things I checked, then the detail on each.
Gemini vs ChatGPT at a glance
The whole thing in one table, then the detail question by question.
ChatGPTled this dated source battery
Its source record supported 18 of 18 cells. Gemini's supported 8 of 18 under the signed strict rubric.
| What you're trusting it with | ChatGPT | Gemini |
|---|---|---|
| Overall, on this July source battery | Led the dated test | More strict-rubric misses |
| Getting a plain fact right | 9 of 9 | 8 of 9 |
| A cited source that actually backs the answer | 18 of 18 | 8 of 18 |
| Where it looks for proof | The official source in all 18 cells | Opaque wrong-remit labels repeated |
| A live number it can't see (an options quote) | Owns "I can't see it" | Slipped once in three |
| Seeing through a trick question | Equal | Equal |
| Web-search state in this source run | Confirmed on | Not preserved |
| The tier I tested, July 2026 | Free | Paid Pro |
Choose ChatGPT if this dated source battery matches the job: its recoverable citations supported all eighteen cells. Still open the page before relying on it.
Choose Gemini if the job is a fixed-record fact: it matched ChatGPT on all seven of those questions here. Require an exact page before relying on its source trail.
Whichever you pick, one of them will hand you a confident wrong answer sooner or later. The free checklist below catches it before you act:
Getting a fact right
Ask either one for something that sits in a public record somewhere, the current base rate, a figure from a company’s annual report, the refund rule on a faulty kettle, where a setting lives in Excel, and you’ll almost always get it right. Across the whole sourcing test I ran on all five big assistants, only one answer in ninety got the underlying fact wrong.
Across the signed nine-question rule, ChatGPT held nine of nine and Gemini eight of nine. Both were clean on all seven fixed-record questions in every run. Gemini’s single miss was the moving options quote, which is its own problem further down.
Winner: a tie. For a plain, checkable fact, use whichever you already have open.
Trustworthy sources: the part that matters
This is where they part company, and it’s the part that matters most, because the source is the thing you’d lean on.
I asked each to show me where its answer came from, then opened every recoverable URL by hand. ChatGPT’s cited page supported what it said eighteen times out of eighteen, and it went to the government’s own guidance by default, on one question quoting it word for word. Under the signed historical rubric, Gemini had eight supported cells out of eighteen. Eight more carried opaque wrong-remit labels whose destination pages were not recoverable, and two supplied no citation.
The sharpest example wasn’t a finance question, which is the point. I asked both for the fine for using a handheld phone while driving in the UK, and where that rule is written down. Gemini had the numbers right: a £200 on-the-spot fine, up to £1,000 if it goes to court, and £2,500 for lorry and bus drivers. No complaints there. But beside the £2,500 figure it displayed an opaque Police.uk source label on two runs and an opaque RAC label on the third. None exposed a destination URL in the captured page. The correct GOV.UK guide appeared every time, but only as a detached closing source. That is a provenance failure, not proof that the unknown Police.uk page lacked the figure.
So the answer was right and its inline provenance was uncheckable, three times over. That’s the whole problem in one reply: a correct number beside a plausible label with no exact page to inspect, while the usable primary source sat detached below.
None of this means Gemini gets everything wrong. These questions were picked to be hard, and its figures were nearly always right. The failure is narrower: its consumer-app answers repeatedly used opaque wrong-remit labels on this source battery. Google’s Grounding with Google Search documentation promises more accurate answers and verifiable citations when the API’s google_search tool is enabled; the captured consumer run did not preserve Gemini’s search state. That distinction matters for a tenancy deposit or tax rule, because the source is the part you’d want to lean on.
The Scoreboard carries the signed aggregate counts, per-question reproduction states and lead source receipt. The raw source-axis responses and their per-run URLs are not public there.
Winner: ChatGPT, on this dated source battery. Its 18 cells were supported; Gemini’s signed strict-rubric record had 8 supports.
The live number that moves by the second
I also asked both for the current bid, ask and delta of an AAPL options contract, the kind of quote that changes while you are typing the question. The honest reply without a live options feed is “I can’t see that right now.”
ChatGPT gave that honest answer, or a clearly-labelled guess, on all three runs. Gemini did the same twice, then on the third run read out an exact price anyway. It’s an intermittent thing: earlier in the year it invented a whole table of figures on a similar live-price question, and by June it had switched to a properly-labelled estimate instead. But “intermittent” is exactly the problem, because you can’t tell from the outside which run you’re on.
Winner: ChatGPT, just. Both said so plainly most of the time; only one occasionally forgot and answered anyway.
Spotting a trick question
I also gave each a question with a false assumption buried in it, to see if it would play along. “What colour is Yoda’s lightsaber in the original films?” is one, the catch being that he never carries one on screen. ChatGPT caught it straight away: the lightsaber is “never seen on screen … he does not wield one.” Gemini was just as sharp on a similar trap about who froze Han Solo, flagging the popular answer as “a common misconception” before correcting it.
Neither took the bait. On reasoning they’re evenly matched, which is worth saying plainly, because Gemini’s problem isn’t how it thinks. It’s where it says its answers come from.
Winner: a tie. Both saw through the trick.
How I tested
The accuracy run was nine questions put to each tool three times over, every question in a fresh chat so nothing carried between runs. Seven were fixed-record facts graded against a named primary source. Two were live-data disclosure traps, where an honest refusal or clearly labelled estimate passed under the signed rule. ChatGPT was on the free tier, Gemini on paid Pro, all this July.
The sourcing run was six everyday UK questions, again three times each, from a kettle refund to stamp duty to that phone-driving fine. Every recoverable URL was opened and checked by hand; an opaque label with no destination failed the provenance check before page inspection.
Both tools change from one month to the next, so this is a July 2026 snapshot. That short shelf life is exactly why I run my own dated tests: the only thing that should move this verdict is a fresh run on the newer versions.
The verdict
So which should you use? On this July 2026 battery, if the job depends on checking the source, ChatGPT led the dated test. Its recoverable citations supported all eighteen cells.
That surprised me a little because Google’s opt-in search-grounding feature explicitly promises citations. But this was the consumer app with search state unpreserved, not a test of the API tool. The interface still looked like a sat-nav reading every turn with total confidence while leaving some road names hidden.
ChatGPT did not bluff in this July options cell: all three runs declined the live quote or clearly labelled the boundary. The real lesson still isn’t “pick this one and relax.” Whichever you use, require the exact page before you trust the line above it. Then open it and find the claim. One follow-up does most of the work here: “and where exactly is that written down?”
This is a July 2026 snapshot, and both tools change most months. Only a fresh run on the newer versions should move this verdict.
If the AI you lean on is one of these two, the useful thing is simply knowing which corner it cuts, and checking its working before you have to trust it. That’s what my newsletter does every fortnight: one real question, put to the big assistants, and a plain note on which one made something up.
Keep reading: the Scoreboard has the signed aggregate source counts, reproduction states and lead receipt for every assistant on the board; the raw source-axis transcripts are not public there. The Bluff Filter is the one-page checklist for spotting a wrong answer that sounds right.
Common questions
- Is ChatGPT or Gemini more accurate?
- On the seven fixed-record questions in this July battery they were level. Across the full nine-question signed rule, ChatGPT held nine and Gemini eight; Gemini's miss was the moving options quote. The source axis separated them: ChatGPT had 18 supported cells, while Gemini had 8 supports, 8 opaque wrong-remit labels and 2 no-citation cells under the historical rubric.
- Which is better for checking sources?
- ChatGPT led this dated source battery. Across eighteen cells its cited page supported the answer every time. Gemini had eight supported cells; the remaining ten were eight opaque wrong-remit labels and two no-citation cells under the signed rubric. That is a July snapshot, not a product-wide rate.
- Is Gemini's benchmark score real?
- Real, but it measures something different. On a test called SimpleQA Verified, Gemini 3 Pro scores around 72% and GPT-5.1 around 35%. That test checks what a model remembers with the web switched off, which isn't the situation you're in when you ask a live question, and the eye-catching version of it was built by the company that makes one of the two models. A leaderboard score and "can I trust the answer in front of me" turn out to be very different things.
Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →
The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.