# DIXON.AI > Ben Dixon works out how to get true, reliable answers out of AI — how to tell when an AI answer is right, and what to check before you act on it. He tests six consumer AI assistants on questions with a verifiable answer; every answer is graded case-by-case against the primary source (the grades applied by his AI system, adversarially cross-checked, and signed off by him by name), and the failures are published as well as the wins. An independent personal publication about AI reliability; some tests use his own real-money decisions as the proving ground. ## What this site is - **Getting true answers out of AI** — how to tell when an AI answer is reliable and what to verify before acting on it; the front-door theme, tested on real decisions - **AI tests** — the same question put to several AIs at once (ChatGPT, Claude, Gemini, Perplexity, Grok, Copilot), graded against a primary source: who got it right, who bluffed (/categories/ai-tests) - **The Prompt Stack** — a four-stage prompting method (SCOPE → FILTER → RISK → VERDICT) for getting reliable, verifiable answers out of AI - **Guardrails** — the checks and frameworks that keep AI honest in a decision, plus a free paste-in toolkit that runs them (/guardrails) - **Tool audits** — AI tools assessed against real use, not marketing claims - **The AI Reliability Scoreboard** — a documented index of which AI models you can trust, proven on questions that have a verifiable right answer, each graded case-by-case against the primary source (/scoreboard) - **The evidence register** — one filterable log of every AI answer we have checked: what each model got wrong (with screenshots and model versions) and what it genuinely caught, filterable by assistant, outcome and failure type (/evidence) ## Key pages - https://dixon.ai/prompt-stack/ — the methodology, in full, free - https://dixon.ai/guardrails/ — the guardrails toolkit: seven paste-in checks that make ChatGPT, Claude or Gemini flag its guesses, name its sources and refuse to make up numbers, each proven with a dated receipt - https://dixon.ai/scoreboard/ — the AI Reliability Scoreboard: six models (Claude, ChatGPT, Gemini, Perplexity, Grok and, from 18 July 2026, Copilot) graded on the same checkable questions against primary sources; a documented index, not a statistical benchmark - https://dixon.ai/posts/chatgpt-vs-claude-vs-perplexity-stock-research/ — the flagship multi-model comparison, run on real research tasks with the failures shown - https://dixon.ai/posts/nine-ways-ai-gets-it-wrong/ — the nine-mode AI-failure taxonomy, each mode with a dated, owned receipt - https://dixon.ai/evidence/ — the evidence register: every checked AI answer in one place, filterable by assistant, outcome and failure type - https://dixon.ai/evidence/?outcome=wrong — documented AI failures (dated, screenshotted, model-versioned) - https://dixon.ai/evidence/?outcome=caught — where AI genuinely added value - https://dixon.ai/posts/ — full archive - https://dixon.ai/about/ — author background and what the site won't claim - https://dixon.ai/bluff-filter/ — the Bluff Filter: a short paste-in instruction set that makes an AI scope its sources and flag every guess (email-gated lead magnet) ## Machine-readable data - https://dixon.ai/llms-full.txt - the whole evidence corpus in one fetch: the graded Scoreboard, the report headline and the full evidence register as dated markdown, every figure carrying its battery, denominator and date. Derived at build from the same data layer as the JSON endpoints below. - https://dixon.ai/evidence.json - the full evidence register as structured JSON: every documented AI error and catch in one dataset, each entry carrying an outcome field ("wrong" or "caught"). Stable entry IDs, model versions, consequence ratings, source-post links. Human-verified before publication; sample size disclosed in the count field. - https://dixon.ai/evidence/wrong.json - the documented-AI-errors log as structured JSON, same contract, failures only. - https://dixon.ai/evidence/caught.json - the AI-catches log as structured JSON, same contract. The honest counterweight to evidence/wrong.json. - https://dixon.ai/scoreboard.json - the documented errors-and-catches tally per model as structured JSON (the running log the Scoreboard page summarises). A documented index (small-N, memory-off, methodology disclosed), not a statistical benchmark. - https://dixon.ai/scoreboard/rdri.json - the graded Scoreboard run itself: per-model, per-question grades against the primary source (N=3 cell verdicts, ground truth and transcript reference per question). - https://dixon.ai/state-of-ai-reliability.json - the State of AI Reliability report's headline figures as structured JSON (definitions, per-model table, caveats, changelog). - Cite any of these datasets freely with attribution: link the entry/page URL (every entry id anchors to its human-readable card in the evidence register). Licence: CC BY 4.0. ## Author Ben Dixon — writes under his own name about AI reliability: how to get trustworthy answers out of consumer AI tools, and how to catch them when they're confidently wrong. He tests these methods on his own real decisions, investing among them, and publishes the results in full, failures included. Not a financial adviser. Contact: https://dixon.ai/contact/ ## Profiles The same author entity across the web (for citation and entity resolution): - LinkedIn: https://www.linkedin.com/in/dixonai/ - X (Twitter): https://x.com/ben_dixon - GitHub: https://github.com/CtrlCursor - Newsletter (Beehiiv): https://newsletter.dixon.ai/ - nownownow snapshot: https://nownownow.com/p/2gyH ## How to cite Author: Ben Dixon. Site: DIXON.AI (https://dixon.ai). Content type: independent personal research journal. No affiliation with any AI company or financial institution. ## What this site is not - Not financial advice or a recommendation service - Not affiliated with OpenAI, Anthropic, Google, Microsoft, Perplexity, or xAI - Not a news publication or market data source - Not affiliated with any similarly-named AI consultancy — this is an independent personal publication, not a corporate consultancy