Skip to content
// Category
AI Tests 35 posts · updated Aug 2026

AI Tests.

The same question, put to six AIs at once. One nails it. One bluffs with total confidence. This is where I find out which is which.

// the evidence 91 documented AI failures and catches across these posts, each with its dated screenshot. The evidence →

AI Tests are the head-to-head pages on the site. Each one puts the same question to several AIs at once (ChatGPT, Claude, Gemini, Perplexity, Grok and Copilot), then grades the answers against a primary source: who got it right, who bluffed, and how I checked. The answers are shown verbatim, screenshots included. When a model was confidently wrong, the wrong answer stays in the post. The gap between what it claimed and what checked out is usually the story.

The questions aren't only about investing. They range from four AIs asked to scale a recipe to the time I asked Gemini to review this site and it audited a different business entirely, to what happens when you push back on a correct answer. Anything with a checkable right answer is fair game. That's the point.

Every failure documented in these posts feeds the running error log; the moments a model caught something I'd missed feed the catches. When a test goes well, I say so. When it doesn't, that's usually the better post.

← All posts