Skip to content
// AI reliability, tested

Six AIs, one real question. I show you which answers hold up.

Subscribe and the Bluff Filter is yours, free: a paste-in prompt that makes any AI flag its guesses before you act.

Every fortnight after that, a real case where trusting AI cost someone something, checked against the record, then run again on today’s assistants.

// The scoreboard9 checkable questions · 12–18 Jul 2026

How are the AIs holding up?

Deliberately hard questions, the kind known to trip AI up.

The graded record: Claude held 9 of 9, ChatGPT held 9 of 9, Gemini held 8 of 9, Grok held 8 of 9, Copilot held 8 of 9, Perplexity held 6 of 9.

The full board, every receipt
Ben Dixon
// The author
Ben Dixon

I test how far you can trust the main AI assistants, and I publish exactly where they get things wrong. Everything here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. More about me →

// Caught in the wildthe log
49documented failures

Documented failures by assistant: Claude 5, Gemini 12, Perplexity 16, Grok 4, ChatGPT 9, Copilot 2, Other 1.

A running tally of the times we’ve caught each assistant out. For interest, not a grade: we haven’t asked them the same number of questions, so the counts don’t rank them. “Other” is a tool outside the six, including our own systems.

// free tool Make your AI admit when it’s guessing. The problem isn’t that it’s wrong sometimes. It’s that it sounds just as sure when it’s guessing, with no tell. The fix is a short set of instructions you paste into any AI you use, ChatGPT, Claude, Gemini or Grok. Same model. It just flags its guesses instead of slipping them past you. Get the Bluff Filter, free →