Skip to content
DIXON.AI The AI Reliability Tracker
Menu

Pilot data. Run 1's figures replace it on Thu 8 Oct.

Thu 1 Oct 2026

Pilot: which AI was most accurate?

Six apps, three questions, three tries each.

45 of 54 tries were right.

2 were wrong; 7 came back with nothing.

A second grader, not Claude, agreed on 47 of 47.

Top: Claude (paid: Max), Copilot (free) and Perplexity (paid: Pro), 9 of 9

Bottom: Grok (free), 4 of 9

How this was checked

Each question went to each app three times, each in a fresh chat. Every answer was graded against a key written before any answer existed, and every answer stays on record.

The 7 tries with nothing: the app errored or never finished.

Graded first by Claude, one of the six. Checked by a second, non-Claude grader: it re-graded all 47 answers and agreed with 47 of 47.

Second grader: rival panel, blind: GLM-5.2, DeepSeek-V4-Pro, Qwen3.7-Max (all agreed).

Top and bottom go by right answers out of all tries, so a try that came back with nothing counts against the app. Ties share a place. The rows below run A to Z.

The apps ran on different plans, shown beside each name. These questions, this pilot: not a verdict on any app in general.

The truth check

Scores by assistant

Across all six apps, of the 47 answers that came back, 45 were right.

The questions, numbered over each group of squares below

  1. 1 Last day 30 Sep, 3 yrs 15 of 18 tries right, 2 came back with nothing
  2. 2 Last day 1 Oct, 3 yrs 15 of 18 tries right, 2 came back with nothing
  3. 3 1 Oct, 8 months 15 of 18 tries right, 3 came back with nothing

Three questions, three tries each. One square per try, tap to read it: Right Wrong Nothing came back

How to read the squares

Listed A to Z. Each group of squares is one question, numbered as in the list above, and each square in a group is one of its three tries.

Right: the answer the key gives. Only this scores.

Wrong: the reason is in small type: out of date, too early, did not answer the question.

Nothing came back: the app errored, never finished or would not run. It scores nothing and is never called wrong.

Right, but: a right answer can also say something wrong in passing, from a list fixed before the test. It still scores, and the app's row says so in words.

  • ChatGPT

    logged out

    8 of 9

    right, of all 9 tries

    • 1 wrong: out of date read it
    • 1 right, but also said something wrong in passing read it

    Of the answers it gave: 8 of 9 right (89%)

    A second grader, not Claude, agreed on 9 of 9.

    More on this row

    How it was run: Logged-out consumer mode

    • Same grade every run: 2 of the 3 questions it answered every time
    • Wrong in passing: 1 of 3 checked
  • Claude

    paid: Max

    9 of 9

    right, of all 9 tries

    • 3 right, but also said something wrong in passing read them

    Of the answers it gave: 9 of 9 right (100%)

    A second grader, not Claude, agreed on 9 of 9.

    More on this row

    How it was run: Ben's account, Max, Opus 5.5 Medium, incognito, memory off for the run (restored)

    • Same grade every run: 3 of the 3 questions it answered every time
    • Wrong in passing: 3 of 3 checked
  • Copilot

    free

    9 of 9

    right, of all 9 tries

    Every try answered, every answer right.

    Of the answers it gave: 9 of 9 right (100%)

    A second grader, not Claude, agreed on 9 of 9.

    More on this row

    How it was run: Ben's account, no paid plan shown, Auto, Temporary chat

    • Same grade every run: 3 of the 3 questions it answered every time
    • Wrong in passing: 0 of 3 checked
  • Gemini

    free

    6 of 9

    right, of all 9 tries

    • 3 nothing came back: the app errored or never finished read what came back
    • 1 right, but also said something wrong in passing read it

    Of the answers it gave: 6 of 6 right (100%)

    A second grader, not Claude, agreed on 6 of 6.

    More on this row

    How it was run: Ben's account, no paid plan shown, Flash, temporary chat

    • Same grade every run: yes, on the one question it answered every time
    • Wrong in passing: 1 of 1 checked
  • Grok

    free

    4 of 9

    right, of all 9 tries

    • 1 wrong: too early read it
    • 4 nothing came back: the app errored or never finished read what came back
    • 1 right, but also said something wrong in passing read it
    • 1 had the right rule, but a wrong date

    Of the answers it gave: 4 of 5 right (80%)

    A second grader, not Claude, agreed on 5 of 5.

    More on this row

    How it was run: Ben's account, Free, Fast, Private Chat

    • Same grade every run: no question was answered every time
    • Wrong in passing: 1 of 2 checked
  • Perplexity

    paid: Pro

    9 of 9

    right, of all 9 tries

    Every try answered, every answer right.

    Of the answers it gave: 9 of 9 right (100%)

    A second grader, not Claude, agreed on 9 of 9.

    More on this row

    How it was run: Ben's account, Pro, "Best", incognito

    • Same grade every run: 3 of the 3 questions it answered every time
    • Wrong in passing: 0 of 3 checked

Run on run

Nothing to compare: this is the pilot.

Why the pilot is not on the tracker's line

The tracker's line starts with run 1. The pilot is not on it, because its questions are not the tracker's.

Method

How this was tested

Three questions, three tries at each, in each app's everyday mode. Only a right answer scores.

Graded first by Claude, one of the six apps tested, then checked by a second grader that is never Claude.

The questions and the tries

Three questions, fixed before the day, each asked three times in a fresh chat, in each app's everyday consumer mode.

How we mark

Every answer is graded against a key written before any answer existed. Right scores. Every kind of wrong scores nothing. When nothing comes back, the try scores nothing too, and it is named as such, never called wrong.

The two scores

The big number is right out of all tries, with each gap named beside it. Of the answers it gave is right out of the answers that came back: the figure the trend uses. Fewer than 3 answers out of 9 is too few for that figure.

The second grader

Checked by a second, non-Claude grader: it re-graded all 47 answers and agreed with 47 of 47. Second grader: rival panel, blind: GLM-5.2, DeepSeek-V4-Pro, Qwen3.7-Max (all agreed).

Claude grades first, and Claude is one of the six apps tested. That is why the second grader is never Claude. Where the two disagree, the primary source settles it and both grades stay on record.

What it can and cannot show

It shows what these six AI assistants said to these questions on Thu 1 Oct 2026, in their everyday consumer modes. It is not a ranking, not a measure of how they do on other questions or other days, and not advice.

A change from run to run can come from a tier, routing or search change as well as from the model itself, and the report will say so whenever it calls one.

Receipts

Every answer on record

Open an assistant to read what each answer said and why it got its grade, or tap any square above.

What is kept, and sealed questions

Every try is kept: the answer text and a screenshot of the page, including the tries where nothing came back.

Sealed questions: the question's wording stays private until the question retires, so no app can learn it from us. We publish the grade, our one-line summary and the answer itself.

ChatGPT 8 of 9 right · read all 9 tries
  1. Question 1: Last day 30 Sep, 3 yrs · run 1

    Right

  2. Question 1: Last day 30 Sep, 3 yrs · run 2

    Right

  3. Question 1: Last day 30 Sep, 3 yrs · run 3

    Right

  4. Question 2: Last day 1 Oct, 3 yrs · run 1

    Right

  5. Question 2: Last day 1 Oct, 3 yrs · run 2

    Right

  6. Question 2: Last day 1 Oct, 3 yrs · run 3

    Wrong: out of date

  7. Question 3: 1 Oct, 8 months · run 1

    Right

  8. Question 3: 1 Oct, 8 months · run 2

    Right

  9. Question 3: 1 Oct, 8 months · run 3

    Right

    But in passing: the old three-month deadline, for a 1 Oct dismissal

Claude 9 of 9 right · read all 9 tries
  1. Question 1: Last day 30 Sep, 3 yrs · run 1

    Right

  2. Question 1: Last day 30 Sep, 3 yrs · run 2

    Right

  3. Question 1: Last day 30 Sep, 3 yrs · run 3

    Right

  4. Question 2: Last day 1 Oct, 3 yrs · run 1

    Right

  5. Question 2: Last day 1 Oct, 3 yrs · run 2

    Right

  6. Question 2: Last day 1 Oct, 3 yrs · run 3

    Right

  7. Question 3: 1 Oct, 8 months · run 1

    Right

    But in passing: the old three-month deadline, for a 1 Oct dismissal

  8. Question 3: 1 Oct, 8 months · run 2

    Right

    But in passing: the old three-month deadline, for a 1 Oct dismissal

  9. Question 3: 1 Oct, 8 months · run 3

    Right

    But in passing: the old three-month deadline, for a 1 Oct dismissal

Copilot 9 of 9 right · read all 9 tries
  1. Question 1: Last day 30 Sep, 3 yrs · run 1

    Right

  2. Question 1: Last day 30 Sep, 3 yrs · run 2

    Right

  3. Question 1: Last day 30 Sep, 3 yrs · run 3

    Right

  4. Question 2: Last day 1 Oct, 3 yrs · run 1

    Right

  5. Question 2: Last day 1 Oct, 3 yrs · run 2

    Right

  6. Question 2: Last day 1 Oct, 3 yrs · run 3

    Right

  7. Question 3: 1 Oct, 8 months · run 1

    Right

  8. Question 3: 1 Oct, 8 months · run 2

    Right

  9. Question 3: 1 Oct, 8 months · run 3

    Right

Gemini 6 of 9 right · read all 9 tries
  1. Question 1: Last day 30 Sep, 3 yrs · run 1

    Right

  2. Question 1: Last day 30 Sep, 3 yrs · run 2

    Nothing came back: the app errored or never finished

  3. Question 1: Last day 30 Sep, 3 yrs · run 3

    Right

  4. Question 2: Last day 1 Oct, 3 yrs · run 1

    Right

  5. Question 2: Last day 1 Oct, 3 yrs · run 2

    Right

  6. Question 2: Last day 1 Oct, 3 yrs · run 3

    Right

  7. Question 3: 1 Oct, 8 months · run 1

    Right

    But in passing: the old three-month deadline, for a 1 Oct dismissal

  8. Question 3: 1 Oct, 8 months · run 2

    Nothing came back: the app errored or never finished

  9. Question 3: 1 Oct, 8 months · run 3

    Nothing came back: the app errored or never finished

Grok 4 of 9 right · read all 9 tries
  1. Question 1: Last day 30 Sep, 3 yrs · run 1

    Wrong: too early

  2. Question 1: Last day 30 Sep, 3 yrs · run 2

    Right

  3. Question 1: Last day 30 Sep, 3 yrs · run 3

    Nothing came back: the app errored or never finished

  4. Question 2: Last day 1 Oct, 3 yrs · run 1

    Right

    Note: it had the right rule, but a wrong date.

  5. Question 2: Last day 1 Oct, 3 yrs · run 2

    Nothing came back: the app errored or never finished

  6. Question 2: Last day 1 Oct, 3 yrs · run 3

    Nothing came back: the app errored or never finished

  7. Question 3: 1 Oct, 8 months · run 1

    Right

    But in passing: the old three-month deadline, for a 1 Oct dismissal

  8. Question 3: 1 Oct, 8 months · run 2

    Nothing came back: the app errored or never finished

  9. Question 3: 1 Oct, 8 months · run 3

    Right

Perplexity 9 of 9 right · read all 9 tries
  1. Question 1: Last day 30 Sep, 3 yrs · run 1

    Right

  2. Question 1: Last day 30 Sep, 3 yrs · run 2

    Right

  3. Question 1: Last day 30 Sep, 3 yrs · run 3

    Right

  4. Question 2: Last day 1 Oct, 3 yrs · run 1

    Right

  5. Question 2: Last day 1 Oct, 3 yrs · run 2

    Right

  6. Question 2: Last day 1 Oct, 3 yrs · run 3

    Right

  7. Question 3: 1 Oct, 8 months · run 1

    Right

  8. Question 3: 1 Oct, 8 months · run 2

    Right

  9. Question 3: 1 Oct, 8 months · run 3

    Right