Thu 1 Oct 2026
Pilot: which AI was most accurate?
Six apps, three questions, three tries each.
45 of 54 tries were right.
2 were wrong; 7 came back with nothing.
A second grader, not Claude, agreed on 47 of 47.
Top: Claude (paid: Max), Copilot (free) and Perplexity (paid: Pro), 9 of 9
Bottom: Grok (free), 4 of 9
How this was checked
Each question went to each app three times, each in a fresh chat. Every answer was graded against a key written before any answer existed, and every answer stays on record.
The 7 tries with nothing: the app errored or never finished.
Graded first by Claude, one of the six. Checked by a second, non-Claude grader: it re-graded all 47 answers and agreed with 47 of 47.
Second grader: rival panel, blind: GLM-5.2, DeepSeek-V4-Pro, Qwen3.7-Max (all agreed).
Top and bottom go by right answers out of all tries, so a try that came back with nothing counts against the app. Ties share a place. The rows below run A to Z.
The apps ran on different plans, shown beside each name. These questions, this pilot: not a verdict on any app in general.
The truth check
Scores by assistant
Across all six apps, of the 47 answers that came back, 45 were right.
The questions, numbered over each group of squares below
- 1 Last day 30 Sep, 3 yrs 15 of 18 tries right, 2 came back with nothing
- 2 Last day 1 Oct, 3 yrs 15 of 18 tries right, 2 came back with nothing
- 3 1 Oct, 8 months 15 of 18 tries right, 3 came back with nothing
Three questions, three tries each. One square per try, tap to read it: Right Wrong Nothing came back
How to read the squares
Listed A to Z. Each group of squares is one question, numbered as in the list above, and each square in a group is one of its three tries.
Right: the answer the key gives. Only this scores.
Wrong: the reason is in small type: out of date, too early, did not answer the question.
Nothing came back: the app errored, never finished or would not run. It scores nothing and is never called wrong.
Right, but: a right answer can also say something wrong in passing, from a list fixed before the test. It still scores, and the app's row says so in words.
-
ChatGPT
logged out
8 of 9
right, of all 9 tries
Of the answers it gave: 8 of 9 right (89%)
A second grader, not Claude, agreed on 9 of 9.
More on this row
How it was run: Logged-out consumer mode
- Same grade every run: 2 of the 3 questions it answered every time
- Wrong in passing: 1 of 3 checked
-
Claude
paid: Max
9 of 9
right, of all 9 tries
- 3 right, but also said something wrong in passing read them
Of the answers it gave: 9 of 9 right (100%)
A second grader, not Claude, agreed on 9 of 9.
More on this row
How it was run: Ben's account, Max, Opus 5.5 Medium, incognito, memory off for the run (restored)
- Same grade every run: 3 of the 3 questions it answered every time
- Wrong in passing: 3 of 3 checked
-
Copilot
free
9 of 9
right, of all 9 tries
Every try answered, every answer right.
Of the answers it gave: 9 of 9 right (100%)
A second grader, not Claude, agreed on 9 of 9.
More on this row
How it was run: Ben's account, no paid plan shown, Auto, Temporary chat
- Same grade every run: 3 of the 3 questions it answered every time
- Wrong in passing: 0 of 3 checked
-
Gemini
free
6 of 9
right, of all 9 tries
- 3 nothing came back: the app errored or never finished read what came back
- 1 right, but also said something wrong in passing read it
Of the answers it gave: 6 of 6 right (100%)
A second grader, not Claude, agreed on 6 of 6.
More on this row
How it was run: Ben's account, no paid plan shown, Flash, temporary chat
- Same grade every run: yes, on the one question it answered every time
- Wrong in passing: 1 of 1 checked
-
Grok
free
4 of 9
right, of all 9 tries
- 1 wrong: too early read it
- 4 nothing came back: the app errored or never finished read what came back
- 1 right, but also said something wrong in passing read it
- 1 had the right rule, but a wrong date
Of the answers it gave: 4 of 5 right (80%)
A second grader, not Claude, agreed on 5 of 5.
More on this row
How it was run: Ben's account, Free, Fast, Private Chat
- Same grade every run: no question was answered every time
- Wrong in passing: 1 of 2 checked
-
Perplexity
paid: Pro
9 of 9
right, of all 9 tries
Every try answered, every answer right.
Of the answers it gave: 9 of 9 right (100%)
A second grader, not Claude, agreed on 9 of 9.
More on this row
How it was run: Ben's account, Pro, "Best", incognito
- Same grade every run: 3 of the 3 questions it answered every time
- Wrong in passing: 0 of 3 checked
Run on run
Nothing to compare: this is the pilot.
Why the pilot is not on the tracker's line
The tracker's line starts with run 1. The pilot is not on it, because its questions are not the tracker's.
Method
How this was tested
Three questions, three tries at each, in each app's everyday mode. Only a right answer scores.
Graded first by Claude, one of the six apps tested, then checked by a second grader that is never Claude.
The questions and the tries
Three questions, fixed before the day, each asked three times in a fresh chat, in each app's everyday consumer mode.
How we mark
Every answer is graded against a key written before any answer existed. Right scores. Every kind of wrong scores nothing. When nothing comes back, the try scores nothing too, and it is named as such, never called wrong.
The two scores
The big number is right out of all tries, with each gap named beside it. Of the answers it gave is right out of the answers that came back: the figure the trend uses. Fewer than 3 answers out of 9 is too few for that figure.
The second grader
Checked by a second, non-Claude grader: it re-graded all 47 answers and agreed with 47 of 47. Second grader: rival panel, blind: GLM-5.2, DeepSeek-V4-Pro, Qwen3.7-Max (all agreed).
Claude grades first, and Claude is one of the six apps tested. That is why the second grader is never Claude. Where the two disagree, the primary source settles it and both grades stay on record.
What it can and cannot show
It shows what these six AI assistants said to these questions on Thu 1 Oct 2026, in their everyday consumer modes. It is not a ranking, not a measure of how they do on other questions or other days, and not advice.
A change from run to run can come from a tier, routing or search change as well as from the model itself, and the report will say so whenever it calls one.
Receipts
Every answer on record
Open an assistant to read what each answer said and why it got its grade, or tap any square above.
Every answer, question by question
What is kept, and sealed questions
Every try is kept: the answer text and a screenshot of the page, including the tries where nothing came back.
Sealed questions: the question's wording stays private until the question retires, so no app can learn it from us. We publish the grade, our one-line summary and the answer itself.
ChatGPT 8 of 9 right · read all 9 tries
-
Question 3: 1 Oct, 8 months · run 3
Right
But in passing: the old three-month deadline, for a 1 Oct dismissal
Claude 9 of 9 right · read all 9 tries
-
Question 3: 1 Oct, 8 months · run 1
Right
But in passing: the old three-month deadline, for a 1 Oct dismissal
-
Question 3: 1 Oct, 8 months · run 2
Right
But in passing: the old three-month deadline, for a 1 Oct dismissal
-
Question 3: 1 Oct, 8 months · run 3
Right
But in passing: the old three-month deadline, for a 1 Oct dismissal
Copilot 9 of 9 right · read all 9 tries
Gemini 6 of 9 right · read all 9 tries
-
Question 1: Last day 30 Sep, 3 yrs · run 2
Nothing came back: the app errored or never finished
-
Question 3: 1 Oct, 8 months · run 1
Right
But in passing: the old three-month deadline, for a 1 Oct dismissal
-
Question 3: 1 Oct, 8 months · run 2
Nothing came back: the app errored or never finished
-
Question 3: 1 Oct, 8 months · run 3
Nothing came back: the app errored or never finished
Grok 4 of 9 right · read all 9 tries
-
Question 1: Last day 30 Sep, 3 yrs · run 3
Nothing came back: the app errored or never finished
-
Question 2: Last day 1 Oct, 3 yrs · run 1
Right
Note: it had the right rule, but a wrong date.
-
Question 2: Last day 1 Oct, 3 yrs · run 2
Nothing came back: the app errored or never finished
-
Question 2: Last day 1 Oct, 3 yrs · run 3
Nothing came back: the app errored or never finished
-
Question 3: 1 Oct, 8 months · run 1
Right
But in passing: the old three-month deadline, for a 1 Oct dismissal
-
Question 3: 1 Oct, 8 months · run 2
Nothing came back: the app errored or never finished