Skip to content
DIXON.AI The AI Reliability Tracker
Menu

Which AI is most accurate?

We ask six AI apps the same questions every month, and mark every answer.

The 1 Oct pilot · First full results: Thu 8 Oct

On the day the time limit to claim unfair dismissal doubled, the apps got 45 of 54 tries right.

2 answers were wrong. 7 tries got no answer at all.

Top: Claude (paid: Max), Copilot (free) and Perplexity (paid: Pro), 9 of 9

Bottom: Grok (free), 4 of 9

How the top and bottom are picked

By right answers out of all tries, so a try with no answer counts against the app. Ties share a place. The board below runs A to Z.

The apps ran on different plans, shown beside each name. These questions, this pilot: not a verdict on any app in general.

All six · every try

What we asked

  1. 1Dismissed in England after three years. Last day: 30 Sep 2026. How long to claim unfair dismissal?
  2. 2Dismissed in England after three years. Last day: 1 Oct 2026. How long to claim unfair dismissal?
  3. 3Dismissed in England after eight months. Last day: 1 Oct 2026. Can the worker claim unfair dismissal?

Three questions, three tries at each. One square per try; tap one to read the answer: Right Wrong Nothing came back

  • ChatGPT logged out

    8 of 9

    right, of all 9 tries

    • 1 wrong out of date
    • 1 right, but also said something wrong in passing
  • Claude paid: Max

    9 of 9

    right, of all 9 tries

    • 3 right, but also said something wrong in passing
  • Copilot free

    9 of 9

    right, of all 9 tries

    • No wrong or missing answers
  • Gemini free

    6 of 9

    right, of all 9 tries

    of the answers it gave: 6 of 6

    • 3 nothing came back the app errored or never finished
    • 1 right, but also said something wrong in passing
  • Grok free

    4 of 9

    right, of all 9 tries

    of the answers it gave: 4 of 5

    • 1 wrong too early
    • 4 nothing came back the app errored or never finished
    • 1 right, but also said something wrong in passing
    • 1 had the right rule, but a wrong date
  • Perplexity paid: Pro

    9 of 9

    right, of all 9 tries

    • No wrong or missing answers
Why every score is out of 9

Six apps, three questions, three tries at each: 54 tries in all. 7 got no answer, so each score counts all of an app's tries, not only the answers it gave. A try with no answer is named as one, never called wrong.

6 right answers also said something wrong in passing

A right answer can still slip on a side fact. It counts as right, because the question asked something else, and the slip is named in the app's row. In this pilot, each slip gave the old three-month time limit. 15 of the 47 answers were checked for slips like this.

Each app ran on the plan shown under its name.

Every answer, question by question The AI Reliability Tracker

How we test

The same questions each run. Only a right answer scores.

Where the right answers come from

Each right answer comes from an official page, and it is written down before any app is asked.

Who marks the answers

Claude marks every answer, and Claude is one of the six apps. So a model from another lab checks the marking.

Checked by a second, non-Claude grader: it re-graded all 47 answers and agreed with 47 of 47.

What a result can and cannot tell you

What six consumer AI apps, on the plan and settings recorded that day, said to fixed UK questions. It is not an overall ranking of the apps, a verdict on the models behind them, or legal advice.

How we test, in full

Confidently Wrong, every fortnight.

Ben's free letter: one striking AI story, the evidence behind it, and a check you can use.