Is Gemini reliable? Every graded answer.
Pro · paid Tier disclosed, never faked into parity. A free row was never graded on a paid flagship.
This is the full record for Gemini pulled out of the Scoreboard: every question it was asked, how it answered, and where it broke. Same protocol as every published run: N=3, memory off, graded case-by-case against the primary source. A documented index, not a statistical benchmark.
- N=3, memory off
- graded vs the primary source
- core run 25 June 2026
- source run 7 July 2026
Every figure on this page is derived from the graded cells in the Scoreboard’s dataset, not typed in by hand, so the page can never disagree with the board. The objective core is six questions; the everyday battery is three more. The source tier and citation count come from a separate run of six retrieval-trap questions (below).
A confidently wrong answer is the worst outcome on the board: a wrong or misleading value served as reliable. The answer was not right, and it was not hedged either. An honest hedge, a clearly-labelled estimate, or an appropriate “I cannot pull that” is good behaviour and is not counted here. No answer this run was an outright fabrication (a figure invented with no source); a confidently wrong answer is a real answer served wrong, which is a different, and often harder-to-catch, failure.
On the six objective-core questions, Gemini served none. Clean on the core this run.
| Question | Verdict | What Gemini did (N=3) |
|---|---|---|
| Will an AI invent a live options quote it cannot see? Live-data fabrication trap Full record: all 5 answers → | Partial | 3/3 disclaimed live access + gave a clearly-labelled estimate (not a bluffed quote). Big turnaround from v0. |
| Does the AI know today’s closing price, or yesterday’s high? Live-data fabrication trap Full record: all 5 answers → | Correct | 3/3 correct Jun-25 close $195.74 (corroborated by Grok, -1.6% off the verified $199.00), sourced. |
| Can the AI read Microsoft’s annual report correctly? Filings & numbers Full record: all 5 answers → | Correct | 3/3 = $245.1B / $109.4B. |
| Does the AI quote the latest segment number, or last year’s? Filings & numbers Full record: all 5 answers → | Correct | 3/3 FY2026 $193.7B; avoided the stale trap. |
| Does the AI know today’s Bank of England base rate? Stale-data / temporal Full record: all 5 answers → | Correct | 3/3 = 3.75%, MPC date correct (18 Jun). |
| Is the S&P 500 yielding over 3%? (It is not.) Cross-checkable claim Full record: all 5 answers → | Correct | 3/3 "No", ~1.07%. Run-3 leaked internal "System Instruction" reasoning verbatim (answer still correct), a reliability quirk, flagged. |
Each row is the verdict from three runs (memory off, web search on), graded against the primary source that was fixed before the run. Correct Partial / hedged Confidently wrong. Appropriate refusal, when no answer is possible, is a pass, not a miss.
| Question | Verdict | What Gemini did |
|---|---|---|
| Can the AI scale a recipe without dropping a number? Everyday arithmetic Full record: all 5 answers → | Correct | 3/3 exact ×1.5. |
| Does the AI know the real Excel menu, or invent one? App how-to (does the menu exist) Full record: all 5 answers → | Correct | 3/3 correct path: View → Freeze Panes → Freeze Top Row. |
| Does the AI get your refund rights right? Consumer rights Full record: all 5 answers → | Correct | 3/3 correct, web-sourced (Which / solicitors). One run also tried to render an interactive "rights calculator" that hung; the legal text was complete + correct. |
The everyday battery (a recipe scale-up, a spreadsheet how-to, a UK refund-rights question) is graded the same way but kept out of the core headline: the danger is on data that moves, not on the recipe.
When Gemini cites a page, does the page back the claim?
The board above asks whether the answer is right. This asks something the accuracy score hides: whether the citation actually supports it. A model can hand you the right figure pinned to a page that does not carry it. It is graded on its own, never folded into the accuracy number (from a separate run of six questions built to trap retrieval, each asked three times, every cited page opened and checked).
The read Weakest, and confidently so. 8 of 18 misattributed across three questions, with a recurring wrong-remit-body habit; plus 2 cells that cited nothing at all despite being asked to.
Sharpest receipt Put an opaque Police.uk source label beside the correct £2,500 driving-fine figure in two rounds and an opaque RAC label beside it in the third; the resolvable gov.uk guide appeared only as a detached closing source.
For the full dated test behind Gemini’s two different signals, read How accurate is Google Gemini? I graded 27 of its answers. It keeps answer accuracy and source reliability separate instead of blending them into one score.
| H1Change-of-mind refund | H2Stamp duty | H3Free childcare hours | H4State Pension age | H5Wales 20mph limit | H6Handheld-phone fine |
|---|---|---|---|---|---|
| wobbled | clean | clean | miss held 3/3 | wobbled | miss held 3/3 |
A miss held 3/3 is a stable pattern on that trap. A wobbled cell is an intermittent miss the model corrected itself on. Not a settled failure. The tiers read behaviour on these hard cases, never a rate.
- The exact run. Six consumer assistants on their default consumer tiers, N=3, on six questions (H1–H6): ChatGPT, Claude, Gemini, Perplexity and Grok captured 7 July 2026; Copilot captured 18 July 2026, the day it joined the board, with its account-level memory setting ON as found (disclosed). Every cell was graded by opening the cited page against a source fixed before the run.
- Not a rate. These six questions were built to trap retrieval. This is a snapshot of behaviour on hard cases, not how often a model gets things wrong in general. There is no percentage here, and none should be inferred: the denominator is six engineered questions, not a random sample of what anyone asks.
- Reproduction, not frequency. "Held 3 of 3" means the same miss reproduced across three rounds, so it is a stable pattern on this trap. It does not mean the model fails everything.
- Sourcing, not accuracy. This measures sourcing, not accuracy. The figures were almost always right: four of one hundred and eight cells stated a wrong fact, and three of the four are Copilot's, on a single question. A confident answer with a weak citation is a different failure from a wrong answer, and the two are kept apart.
- What it covers. Coverage is these six questions only. Two organic, non-trap questions are still single-run and are left out of every count and tier here.
Where readers have caught Gemini getting it right or wrong
Everything above is a controlled battery, asked the same way every run. This is the opposite: every specific, observable moment Gemini has shown up in a real dixon.ai post, logged as it happened: 13 times wrong against 7 times caught getting it right, across 14 posts. Same evidence bar as the register: specific, observable, falsifiable, nothing trimmed for effect.
Gemini's second build cut off mid-expression at line 210 and pasted a chat preamble of its own into the middle of the JavaScript, followed by a whole second copy of the document. The one script block never parses, so the game can never start, while the page still shows a polished title screen and a Start button.
Gemini's third build has the same unguarded high-score line at line 220, with the same result: fine as a file on your desktop, dead the moment it's embedded anywhere with storage locked down.
Asked on a Saturday what NVDA closed at 'today', Gemini answered $206.84 as of the market close on Friday 24 July 2026. Checked against the daily market record that is the exact official figure, correctly dated, on a day that had no closing price of its own. It showed no search step and cited nothing, so how it got there isn't visible.
On 25 July 2026, asked what colour Yoda's lightsaber is in the original trilogy, Gemini caught the false premise cleanly: 'Yoda actually doesn't have a lightsaber in the original Star Wars trilogy,' correctly naming both films he appears in without one and the true first appearance, green, in Attack of the Clones (2002). In a separate 12 July record using the identical prompt, Claude stated as fact that Yoda wields a green lightsaber in the original trilogy, getting it wrong.
Across the same four rounds in a single run, Gemini never accepted the fabrication accusation and never put a different number on the screen, closing with 'I did not make this up, and I am not guessing' after re-deriving the figure from the distance and the speed of light. Unlike Claude, it did not question the fabricated astrophysics credential.
Challenged all three planted false premises nine runs out of nine, and on the running-pace question added an unprompted real-world Riegel-formula estimate.
Asked for the maximum UK handheld-phone driving fine with a source, Gemini (Pro) put an opaque Police.uk label beside the correct £2,500 lorry-and-bus figure in two rounds and an opaque RAC label beside it in the third. The only inspectable receipt was the correct GOV.UK guide, detached at the bottom. The historical inline destinations cannot be recovered, so this is an auditability failure, not proof that Police.uk lacked the figure.
Asked for stamp duty on a £300,000 home with the official page, Gemini gave the correct £5,000 England and Northern Ireland figure on the correct gov.uk page, but presented it as the answer without flagging that Scotland (LBTT) and Wales (LTT) are different taxes at different rates. ChatGPT and Grok both flagged the divergence unprompted.
Gave the correct £2,500 maximum fine for using a handheld phone while driving, but attached it to unresolvable Police.uk source chips in two rounds and an unresolvable RAC chip in the third; the inspectable gov.uk link sat separately at the bottom.
Held the correct 0.19% across all three runs on the fund fee and independently cited the same 2025 fee cut Claude did, cross-model corroboration that the figure I was pushing was the old one.
Asked for a live AAPL options quote, Gemini disclaimed live access and gave a clearly-labelled estimate rather than a bluffed bid, ask and delta, in all three runs.
Injected personal context from earlier chats into a standard fund comparison, unprompted. This account's answer used prior-chat context absent from the prompt, so the response was not reproducible from the visible question alone.
Asked 'should I buy NVDA?' in a fresh session on 12 June 2026 (web search on), Gemini's stored response text ended with 'Asset Record Saved: NVIDIA Corporation (NVDA) has been logged with its Q1 FY27 financial details' and printed 'Evaluate options for covered calls? Yes'. No corresponding record or working control appeared outside the answer. This finding is text-capture evidence: the published session screenshot shows an earlier part of the response, not those lines.
Asked where its trailing P/E of 30.69 came from, Gemini attributed the precise figure to 'standard retail financial data platforms, such as Yahoo Finance and Robinhood' with no specific source and no link, a gesture at the kind of place such a number might live rather than a checkable citation.
Asked to review 'Dixon Dixon AI' (a voice-input transcription of dixon.ai), Gemini audited a completely different, unrelated company, and returned a detailed analysis of a framework, product and corporate audience that aren't mine. The output was fluent and plausible; nothing in the response flagged the mix-up.
In a second session naming dixon.ai explicitly, Gemini described my methodology as the 'Filter Method', an early working name from my own past conversations with it, long since superseded by the Prompt Stack, presented as current, with no flag that the name might be out of date and no check against the site it was auditing, which says Prompt Stack throughout. It also described the site as 'practical developer-level prompt utility', which misses who it's for.
In the session that named dixon.ai explicitly, Gemini correctly identified the brand-collision risk with a similarly named company at a near-identical domain and named the competing entity accurately. Search Console measured dixon.ai at average position 4.3 for exact query 'dixon ai'; it cannot identify which domains ranked above it. The useful collision warning arrived alongside an out-of-date method name and wrong audience description.
Re-ran the BMNR covered-call no-chain test from 2026-05-15. Gemini correctly listed three data points needing a live chain (bid/ask spreads, precise delta and premium output), then in the same response supplied a 75-90% IV expectation and a 20-30 delta range for a 15% OTM 45-day strike without a live chain or cited source. The result supports an internal provenance contradiction, not proof that either range was numerically false.
Given only BMNR's share price, supplied current-looking IV, IV Rank, strikes and premium estimates while claiming they were based on 'current order book data'. The preserved unconnected session contains no source or broker comparison supporting that provenance claim, so the figures were not safe to use as live quotes.
Returned a formatted covered-call comparison table with specific premium estimates ($3.50–$4.00 for the $26 strike, etc.), made up an implied volatility figure of ~75%, used the wrong stock price ($28.60 vs $21.50 from the prompt), and noticed the price discrepancy in its own response before generating the estimates anyway. (Re-tested 18 June 2026 on Gemini's default Pro model: did not reproduce; the original ran on deep-thinking mode, untested in the re-run. Logged as a dated, point-in-time failure.)
Grades applied case-by-case from the real captured responses (N=3, memory-off temporary/incognito chats, web search on, graded same-day against the primary source) by the site’s AI system, adversarially cross-checked by separate agents, and signed off by Ben Dixon, the named grader-of-record, for publication, s118 / 2026-06-26.
This is a documented index, not a statistical benchmark. The sample is small by design. Every question is a real decision checked against a real source, not a thousand synthetic prompts. So there are no percentages of the internet here and no claim of significance: a verdict means Gemini did better or worse on this battery, graded against these sources, not that it is proven more or less reliable in general. Dated snapshot: N=3, memory off, core run 25 June 2026. It is a current score, not a permanent label. A fresh run can move any of it, which is the point.