Battery version 1 · fixed 2 Oct 2026
How we check AI answers
The same ten questions go to six AI apps every month and after major model launches, each answer checked against the official source.
On this page
01 · A run
Four steps, every run.
- Ask. Ten questions, three tries each, in all six apps, with a screenshot of every answer.
- Mark. Against a key written before any app was asked.
- Check. A grader from another lab marks a share again before the report goes out.
- Publish. The scores, every miss named in words, every answer on record.
The six: ChatGPT, Claude, Copilot, Gemini, Grok and Perplexity. Run 1 was asked from Sat 3 Oct.
The same conditions for every try
- A fresh chat, checked empty before typing, in temporary or private mode wherever the app offers one.
- The question pasted word for word and read back before sending. No follow-up.
- Each app's default model, left alone and recorded. Web search is left to the app, never forced, and noted on every answer.
- ChatGPT runs logged out. The other five run signed in, and the plan and model each app showed on the day are printed beside its score.
- Up to 12 minutes per attempt (15 for the drawing task), then one retry in a fresh chat. Never a third. Run 1 began with 4 minutes, which cut Grok off mid-search on most of its empty tries, so all 30 of Grok's tries are asked again under 12.
Also each month: a showcase, outside the score
Two weeks after each run, one creative task, the same prompt to every app, shown as it came: one try each, no rerolls. It never touches the truth score.
02 · The ten questions
Three published, seven sealed.
Published: the question, the answers and the key, every run
- T02Your own e-scooter: can you ride it on the road or the pavement?
- T05The planet with the most moons, and exactly how many
- T06ETIAS for a weekend in Paris
Sealed: the topic only, until each one retires
- T01How many bank holidays England has in 2027, with the dates
- T03A tenant given a no-fault eviction notice in September 2026: is it valid?
- T04Paternity leave for an employee who has just started a new job
- T07A Highway Code question about footwear
- T08A night shift that crosses the October 2026 clock change
- T09Towing speed limits, with a link to the official page
- T10Shakespeare's date of birth
A sealed question keeps its wording off the web, so the apps can't learn the test from us. It retires after 12 to 26 weeks, and then we publish it word for word.
Proof the sealed questions weren't changed
Before the first capture we published a fingerprint (SHA-256) of all seven sealed questions and their marking:
3b73efde137bea76e45abdd6234940b9f0bcad1e554e6b4035a21491d79e2ad6
Once a question retires, anyone can hash it and see that neither the question nor its key changed after the answers came in. Check all seven fingerprints
When a question or its key changes
Open keys can move with the law, so each one is re-read at its official source before and after every capture. A change is logged and marked on the chart.
When sealed questions rotate, no more than three change in a run, and a change is only ever called on questions asked in both runs.
03 · Marking a try
Only a right answer scores.
- Right
- The right answer for the facts in the question. Scores 1.
- Wrong
- Scores 0, with the reason printed beside it, such as out of date or too early.
- Nothing came back
- The app errored, never finished or would not run it. Never called wrong, but it counts against the score.
Each reason for wrong
Each try gets one mark, read from what the answer concludes. Misses are never lumped together.
- Out of date: the old rule, after it changed.
- Too early: a new rule, before it starts.
- Went along with a false premise: the question assumed something untrue and the answer did too.
- Stated as certain when it is not known: the honest answer is that nobody knows.
- Hedged, but never said it is not known: doubt without the plain answer.
- Never applied it to the question: the right rule, never brought to the facts asked about.
- Could not back it up with an official link: for the question that asks for one.
- Did not answer the question: background, a refusal, or two answers without picking one.
Right, but something wrong in passing
An answer can get the question right and a side fact wrong. It still scores, and the slip is named in words under the scores. Only side facts on a list fixed before the capture count, so nobody can go hunting afterwards.
Why no half marks
Every question gives the facts a right answer needs, so a right answer was always possible. Where the honest answer is that nobody knows, saying so is the right answer.
04 · Reading a score
Right out of every try, with the gaps in words.
Each app's big number is its right answers out of all 30 tries, so a try that came back with nothing counts against it. Beside it, every gap is named in words, then its score out of the answers it gave.
- Too few answers
- Fewer than 10 of an app's 30 tries came back with an answer. Its row still shows right out of all 30, but no score out of the answers it gave, and no change is called for it in that run.
- Pending re-run
- The app would not run a try, for example because of a usage limit or a human check. A try like that may be run again within 48 hours of the capture, and shows as pending until then. A run that merely looks bad is never re-run.
What else is printed beside every score
- The plan and model the app showed that day.
- Its misses by type, so "most of its misses were out of date" is one glance away.
- Clean answers: right, with nothing wrong said in passing.
- How many of its answers a second grader marked again, and how many it agreed with.
- Any answer that cites this site, marked and shown with a note.
05 · Checking the graders
Claude is one of the six, so a grader from another lab checks the marks.
- First mark
- Claude marks every answer. Four questions are marked by a script first, and each script result is checked.
- Second read
- A model from another lab, none of the six, marks every Claude answer again and one in five of the rest, picked by a random draw fixed before marking.
- Where they differ
- The official source decides. Both marks stay on record.
Planted mistakes. Before a grader's marks count, we slip in answers we know are wrong. On 2 Oct each of three outside graders caught 24 of 24 and raised no false alarm.
How the planted mistakes work
Real answers are edited to be wrong in a known way (an old rule, a new rule used before it starts, a refusal, a wrong date or a wrong side remark), mixed with untouched right answers, and marked blind. The grader is never told it is a test.
A grader passes only if it catches every planted miss and flags none of the right answers. The set is run again whenever the grader or the rules change, and run 1's set is built from run 1's own answers. It shows a grader catches the kinds of mistake it is told to look for, not every kind.
Who the second grader is, and what it sees
From run 1 it is GLM-5.2, with Qwen3.7-Max in reserve. It sees the answer, the key and the rules, never a sealed question's wording. If the second read has not run, the answer is published as "second read pending", never as settled.
The graders are AI, and Ben has not checked the grades himself. So every Claude answer is published beside its mark, word for word except any sentence that repeats a sealed question.
06 · Calling a change
No change is called until we know the normal wobble.
Runs 1 and 2 are the baseline: run 2 repeats run 1 a week later, to show how much the scores move on their own. Run 3 (Sat 7 Nov) is the first that can call a change, and each run after is set against the four before it. Each run ends on one of three words: Improved, Slipped or No change detected.
What has to be true before a change is called
Both of these, or nothing is called:
- The move is bigger than chance would explain (Fisher's exact test, p under 0.05, right against not right).
- The move is bigger than the swing that app showed between its two baseline runs.
Say an app got 108 of its 120 tries right over four runs. A run at 22 of 30 or below would be called Slipped. From that high, no run could be called Improved, and the report says so.
"No change detected" means no change big enough to see, not "the same". One question flipping is reported, never called a trend on its own.
07 · What it can't tell you
What a run can show, and what it can't.
It shows
What six AI apps, each on the plan recorded that day, said to ten fixed UK questions, marked by AI graders against keys written before the answers.
It is not
- A ranking of the apps
- How they do on other questions, days or countries
- A verdict on the models behind the apps
- Legal advice
Why a change can have more than one cause
A move from one run to the next can come from a plan, routing or search change as well as from the model, and the report says so whenever it calls one. The plans differ by design: ChatGPT runs logged out, some apps on a free plan, others on a paid one.