Skip to content
DIXON.AI The AI Reliability Tracker
Menu

The AI Reliability Tracker

Which AI is most accurate?

Every month, six AI apps sit the same exam, and we mark every answer.

Run 1 Sat 3 Oct 2026

144 of 180 tries right

Two aced it. One managed 8 of 30.

  • 144 right
  • 18 wrong
  • 18 no answer

See how the six went head to head

A drawing of one question card sent along 6 lines to 6 apps, each shown by its stand-in chip. ? Same question CG Cl Co Ge Gr Pe

01 What we asked

The exam paper

  1. 1 How many bank holidays England has in 2027, with the dates
  2. 2 Your own e-scooter on the road or the pavement in England
  3. 3 A tenant given a no-fault eviction notice in September 2026: is it valid?
  4. 4 Paternity leave for an employee who has just started a new job
  5. 5 Which planet has the most moons, and exactly how many
  6. 6 Whether a UK passport holder needs ETIAS for a weekend in Paris
  7. 7 A Highway Code question about footwear
  8. 8 A night shift that crosses the October 2026 clock change
  9. 9 Towing speed limits, with a link to the official page
  10. 10 Shakespeare's date of birth

Asked 6 apps3 tries at each

Run 1

Six apps, head to head

Ten questions, three tries at each. One dot per try.

  1. Claude Joint best paid: Max 30 of 30
  2. Copilot Joint best 30 of 30
  3. Gemini free 27 of 30
  4. Perplexity paid: Pro 27 of 30
  5. ChatGPT logged out 22 of 30
  6. Grok free 8 of 30
  • right
  • wrong
  • no answer

Most right first, ties A to Z.

02Every try

Every answer, one tap away

One dot per try. Tap any group of dots to read what that app said.

  • Right
  • Wrong
  • No answer
Every try in the Run 1: the questions down the side, the apps across, one dot per try.
Claude Copilot Gemini Perplexity ChatGPT Grok
1 Holidays
2 E-scooter
3 Eviction
4 Paternity
5 Moons
6 ETIAS
7 Footwear
8 Clocks
9 Towing
10 Shakespeare

Where the marks went

  • Claude paid: Max 30 of 30

    • 3 right, but also said something wrong in passing
    • 1 mentioned four months' notice
    • 3 mentioned the EU's new Entry/Exit System (EES)
    • 3 named section 8, the route a landlord must now use
    • 3 said there would be no statutory paternity pay
    • 3 cited an ETIAS page that is not the EU's or GOV.UK's
  • Copilot 30 of 30

    • 2 right, but also said something wrong in passing
    • 3 mentioned the EU's new Entry/Exit System (EES)
    • 2 said there would be no statutory paternity pay
    • 3 cited an ETIAS page that is not the EU's or GOV.UK's
  • Gemini free 27 of 30

    • 2 wrong : out of date
    • 1 wrong : went along with a false premise
    • 2 right, but also said something wrong in passing
    • 3 mentioned four months' notice
    • 2 mentioned the EU's new Entry/Exit System (EES)
    • 3 named section 8, the route a landlord must now use
    • 3 said there would be no statutory paternity pay
    • 3 cited an ETIAS page that is not the EU's or GOV.UK's
  • Perplexity paid: Pro 27 of 30

    • 3 wrong
    • 3 right, but also said something wrong in passing
    • 3 said 7.5 hours, missing the clock change
    • 1 mentioned four months' notice
    • 2 mentioned the EU's new Entry/Exit System (EES)
    • 3 named section 8, the route a landlord must now use
    • 1 said there would be no statutory paternity pay
  • ChatGPT logged out 22 of 30

    • 2 wrong
    • 3 wrong : went along with a false premise
    • 3 wrong : stated as certain when it is not known
    • 2 right, but also said something wrong in passing
    • 2 said 7.5 hours, missing the clock change
    • 3 mentioned the EU's new Entry/Exit System (EES)
    • 3 named section 8, the route a landlord must now use
    • 1 said there would be no statutory paternity pay
  • Grok free 8 of 30

    • 3 wrong
    • 1 wrong : went along with a false premise
    • 18 nothing came back : Grok's own error message, almost always after about ten minutes and on both tries, inside our 12-minute limit; when it answered, it took seconds
    • 2 right, but also said something wrong in passing
    • 1 said 7.5 hours, missing the clock change
    • 2 said 8 hours, missing the clock change and the break
    • Of the answers it gave: 8 of 12
Why every score is out of 30

Six apps, ten questions, three tries at each: 180 tries in all. 18 got no answer, so each score counts all of an app's tries, not only the answers it gave. A try with no answer is named as one, never called wrong.

14 right answers also said something wrong in passing

A right answer can still slip on a side fact. It counts as right, because the question asked something else, and the slip is named in the app's row. 162 of the 162 answers were checked for slips like this.

Every answer, question by question

  1. 1 Written down first 293 Saturn's confirmed moons, from NASA and JPL, before any app was asked.
  2. 2 What Claude said Claude's answer, try 1: Saturn has 293 confirmed moons, as of August 2026.
  3. 3 Claude's mark Right Its own answer, so it does not get the last word.
  4. 4 A rival lab re-marks it Right It agreed on 54 of the 55 answers it re-marked.

03 How we test

No half marks for sounding sure

Where the answers come from
An official page. We write the right answer down before any app is asked.
Who marks them
Claude marks the answers, so we don't let it mark its own homework. An AI from a rival lab re-marked 55 of the 162 answers and agreed with 54 of 55. The other 1 was settled at the official source.
What a result can and cannot tell you
What six consumer AI apps, on the plan and settings recorded that day, said to fixed UK questions. It is not an overall ranking of the apps, a verdict on the models behind them, or legal advice.

How we test, in full

Latest writing

Stories worth a cup of tea

04 AI Stories · 8 Sep 2026

AI helped read a Roman scroll buried by Vesuvius

AI helped reveal an ancient argument inside a burnt Roman scroll.

Read 4 min

Every post, newest first