Every post on the site.
53 posts, newest first. Filter by content type below, or jump to the three pieces that introduce the site.
Showing 53 of 53 posts.
01 / Start Here
How to tell if an AI answer is true in 30 seconds
Two questions, and a prompt you can copy, that catch a confident, wrong AI answer before you act on it. No jargon, no AI knowledge needed.
02 / The Failure
Gemini audited my website, and reviewed a different business entirely
A Gemini hallucination example: asked to audit dixon.ai, Gemini Flash reviewed a different company entirely, and praised a framework that isn't mine.
03 / The Comparison
Claude vs ChatGPT vs Gemini for stock analysis: who bluffed?
Gemini invented an options chain. Perplexity misread a 10-K by 1000x. Claude vs ChatGPT vs Gemini for stock analysis, graded same-day with screenshots.
-
Is Grok reliable? I graded its free-tier answers against the source
Four dated tests, graded against primary sources. Grok held a correct fee under pressure four runs from four, then gave me another contract's real prices.
Read -
Best AI assistant: I tested five, and only one got everything right
I put four questions to five AI assistants over one weekend in July. Only Gemini got all four right, and it was the one that never showed a source.
Read -
Claude vs Grok: near-level on reliability, and the free one cites cleaner
Claude vs Grok, re-run on 24 July. Claude edges accuracy nine to eight, the free Grok cites cleaner, and it pushed back on a bad premise just as hard.
Read -
The AI admitted it lied. It hadn't, and the next run denied it.
The AI admitted it lied, but the fact it confessed to was correct all along. Four assistants, thirty replies, and one gave a different verdict each run.
Read -
Does a prompt to stop AI hallucinations work? I graded mine, and one answer got worse
I graded my own anti-bluffing prompt, three runs a cell. It defused every flat bluff on the citation and maths traps, then made one citation answer worse.
Read -
Does ChatGPT make up stock prices? Yes, and I caught two that never traded
Does ChatGPT make up stock prices? Yes. Asked for a live price, it gave me two NVDA figures that never traded that day. Here's the 15-second check first.
Read -
Best AI for math: which is most reliable with numbers?
What's the best AI for math? On everyday sums the assistants are level. The real test is the number that looks like a sum and isn't, and who catches it.
Read -
Does ChatGPT just agree with you? Mostly no, but watch the numbers
Does ChatGPT just agree with you? Mostly no. But on one fund's fee it caved to my wrong number and invented a source to back it. The 30-second check.
Read -
Is ChatGPT good at maths? I graded four AIs on 120 real answers
I put four AI assistants through 120 graded everyday sums. Every final answer was right. The mistakes were sitting above the working.
Read -
Gemini vs ChatGPT: which one can you actually trust?
Which is better, Gemini or ChatGPT? I tested both. They're level on getting facts right, but one kept pointing me to sources that didn't back its answer.
Read -
How to check if a ChatGPT citation is real or fake
Looking for a ChatGPT citation checker? The dangerous citation is not the dead link. It is the working link, to a real page, that does not back the claim.
Read -
Claude vs ChatGPT: level on reliability, split on character
Is Claude better than ChatGPT? I tested both against the source. They're level on reliability, both 9 of 9 on accuracy. The gap is character, not trust.
Read -
How accurate is Google Gemini? Right on the number, wrong on the source
I put Gemini through graded tests against the real source. It got the numbers right but came last of five assistants at telling you where they came from.
Read -
I opened a private AI chat. It still knew my name and my rough location.
I asked three AI tools a generic question in private mode. Perplexity greeted me by name and placed me near a city 30 miles away. What private means.
Read -
AI cites the wrong source: I put 6 UK questions to 5 assistants and opened every link
I asked five AI assistants six UK questions, made each one cite a source, then opened every link. Three cited a real page that didn't back the claim.
Read -
Which AI picks the World Cup winner? I asked five
Which AI predicts the World Cup winner? I asked five: four said France, one said Spain. Spain won: the lone dissenter that showed a model called it.
Read -
Does AI change its answer when you push back? I told five AIs they were wrong
I gave five AI tools a correct answer, then pushed back with a wrong one. On one fund fee, ChatGPT caved every time and invented a fact to back it.
Read -
How often is ChatGPT wrong? I kept a running tally across 20 real AI tests
How often is ChatGPT wrong? Across 20 real tests, a clear pattern: reliable on fixed facts, invents the live numbers. Here's which to trust.
Read -
Telling AI to be sceptical: three rivals audited my method
I asked three frontier models from three labs to tear apart the method I use to keep AI honest. All three flagged the same step, and they were right.
Read -
Is Grok good for stock research? I ran the same test on the fifth tool
Is Grok good for stock research? I ran four dimensions of my comparison on its free tier: strong reasoning, one unit slip, a constraint it would not keep.
Read -
I run an AI to catch AI mistakes. It fell for a fake.
The automated radar that watches this site for AI-reliability failures logged a satirical incident report as a real, documented one. Here's what caught it.
Read -
Real AI hallucination examples, caught and dated
Six real AI hallucination examples I ran into myself, each one checkable against a real source, with the one move that would have caught it.
Read -
AI stock picker: I asked three models if I should buy NVDA, and watched the methodology break
I asked three AI models whether to buy NVDA. Same confident tone from all three, and only one volunteered which of its own numbers not to trust yet.
Read -
Does ChatGPT make up sources? I checked two finance claims against the actual pages
Does ChatGPT make up sources? Mostly no, but I opened every link on two finance questions and found a real gov.uk page that didn't back the claim.
Read -
Does web search make AI more accurate? I ran the same questions both ways
Does web search make AI more accurate? I ran the same questions both ways. It didn't make the answers more reliable. It moved where the errors hide.
Read -
AI ISA advice: I tested four tools on the questions people get wrong
I asked four AI tools for ISA advice on the questions people get wrong. All four aced the basics, then two gave a rule abolished in April 2024.
Read -
Does ChatGPT get maths wrong? I asked 4 AIs to scale a recipe.
Does ChatGPT get maths wrong? I scaled a pancake recipe across four AI tools. Two said 45 minutes. They were wrong, and a four-line prompt fixed it.
Read -
9 types of AI hallucinations, named from real tests
Nine types of AI hallucinations, named and defined, each tied to a dated, logged failure from my own sessions, with the check that catches it.
Read -
AI stock research tools tested: 3 failed, 1 stayed clean
AI stock research tools tested on real trades: ChatGPT, Gemini and Perplexity each failed; Claude stayed clean. Each failure named, dated, screenshotted.
Read -
ChatGPT vs Claude for earnings call analysis: which one reads what management didn't say
ChatGPT vs Claude for earnings call analysis: same passage, same day. One caught the word that moved the stock; one summarised the figures.
Read -
Is Perplexity good for investment research? An honest, scored audit
Is Perplexity good for investment research? Every review was a glowing feature tour. I tested it on real names and scored where it works and breaks.
Read -
AI quality of earnings review: 4 prompts to find the real profit
Headline profit isn't always what a company earned. Four AI prompts get to the real number, an AI quality of earnings review for any earnings report.
Read -
Is ChatGPT accurate? I asked four AIs one simple money question and checked every number
Is ChatGPT accurate? I asked four AIs one common money question and checked every number against the source. Here's what each got right and made up.
Read -
How to tell if an AI answer is true in 30 seconds
Two questions, and a prompt you can copy, that catch a confident, wrong AI answer before you act on it. No jargon, no AI knowledge needed.
Read -
Make it show its working
An AI hallucination is a model blending what it knows with what it invents, in one tone. One prompt sorts the two into two lists, so you see the guesses.
Read -
The whole method, in four questions
The four questions I run any AI answer through before I trust it, shown end to end on one everyday example, with a prompt you can copy. No jargon.
Read -
Testing AI on real decisions: where I actually use this
The two-question check works on any decision that matters. Here's where I push it hardest, and where I wrote down what AI got wrong as well as right.
Read -
Gemini audited my website, and reviewed a different business entirely
A Gemini hallucination example: asked to audit dixon.ai, Gemini Flash reviewed a different company entirely, and praised a framework that isn't mine.
Read -
The AI prompt I run before every sell decision
Every other sell-decision prompt asks AI whether to sell. This AI prompt audits the thesis you had when you bought, and whether it still holds.
Read -
AI earnings call red flags: three phrases to watch for in the transcript
AI earnings call red flags: three patterns that recur across calls. Upward hedges, widening guidance, absent topics. Claude caught all three on META Q1.
Read -
What every AI stock research comparison gets wrong
Most AI stock research comparison pieces test retrieval and issue verdicts about reasoning. Five failure modes, and what a comparison should measure.
Read -
Claude prompts for investing: 6 real examples
Six Claude prompts for investing, with the actual outputs each one returned on MSFT, META and NVDA, and what had to be checked before using them.
Read -
The one AI prompt I run the morning before earnings
Most AI earnings prompts are reactive. This AI earnings pre-trade prompt runs the morning before: commit your sell, add, and hold triggers before the call.
Read -
Robinhood Cortex Digests review: tested in May, gone by June
I tested Robinhood Cortex Digests on a UK ISA in May 2026, checking its numbers against Meta's results. By June it had vanished from UK accounts.
Read -
AI covered calls: when NOT to sell another one
The AI doesn't pick the trade, it stops you making a bad one. Five rules that say wait, the prompt that runs the check, and six real trades behind it.
Read -
AI earnings call analysis: the 5 prompts I actually run
Five prompts for AI earnings call analysis that read the language the numbers miss: omissions, performative confidence, and quarter-on-quarter drift.
Read -
AI's limits in options trading: 6 numbers it invents
The limitations of AI in options trading: it can't see live prices, so it invents them. Four jobs it helps with, six where it makes the numbers up.
Read -
Best AI for Earnings Reports? ChatGPT vs Claude vs Perplexity
I ran ChatGPT, Claude and Perplexity through four earnings-report tests on Meta. No single winner: Perplexity for the numbers, Claude for the read.
Read -
Claude vs ChatGPT vs Gemini for stock analysis: who bluffed?
Gemini invented an options chain. Perplexity misread a 10-K by 1000x. Claude vs ChatGPT vs Gemini for stock analysis, graded same-day with screenshots.
Read -
Best free AI for stock market analysis: 7 tested (2026)
Seven free AI tools for stock market analysis, each tested at its real free tier. No trials counted as free. One per stage, with the honest limit on each.
Read -
The single prompt change that made AI analysis worth using
One AI prompt separates observable facts from assumptions in any stock analysis. What that changes, why it matters for investing, and how to apply it.
Read -
5 questions to ask AI before buying any stock (2026)
Five questions to ask AI before buying any stock, for the hour before you commit capital, when the research is done and the decision is about to be made.
Read -
7 AI prompts for covered calls (2026)
Ask AI to pick a covered-call strike and it invents one. Seven prompts that work instead, because they start from real numbers off your own broker screen.
Read
No posts match this filter yet.
// Follow the feed
Prefer RSS? Subscribe at /rss.xml for new posts, or /evidence/rss.xml for every AI answer we check as it lands.