When I get it wrong, it goes here.
Everyone’s evidence looks clean until you ask what they got wrong. This page logs where dixon.ai has been wrong: a number I had to recut, a claim that didn’t hold when I re-tested it, a verdict I’d revise. Dated, kept visible, not quietly edited out of existence. It’s the same standard I hold the AI tools to: the mistakes are the receipt.
This is my own log. The AI tools’ errors live on Lessons, a different page. This one is about me.
- 14 Jul 26
My error count was stricter than my own rule.
14 July 2026. The reliability report's "confident error" count was including results that were honestly hedged, which my own rubric treats as good behaviour, not error. I'd been harder on the tools than my own definition allowed. I enforced the definition in the data layer and recut the headline from 3 of 30 to 1 of 30. The changed number sits in the report's own changelog, because a number that moves is worth more than one that never does.
- 12 Jul 26
A Gemini claim only held two times out of three.
12 July 2026. I'd logged that Gemini disclaimed live market access and gave a clearly-labelled estimate rather than a bluffed price, across three runs. When I re-ran it, two of the three held. The third gave an exact bid and ask on a closed market, with no estimate label. The June record stands as dated; an update note above the claim on the post discloses that it didn't fully reproduce.
Nothing here was forced out of me. The reliability report is re-cut on every fresh run and keeps every prior number in its changelog; the claims logged on posts get re-tested as the models change under them, and when one doesn’t hold, the note goes on the post and a line comes here. If you ever find something wrong I haven’t logged, tell me. That’s a correction too.