// Browse the evidence
56 posts, sorted by what each one does.
Raw tests, tool audits, the method itself, and the checklists. Four kinds of evidence, four different ways to watch an AI show its working, or not. Pick the one that fits what you came for.
AI Tests 33 posts
The same question put to several AIs at once: who got it right, who bluffed, and how I checked. Verbatim answers, graded against a source.
Latest: Do AI model upgrades fix mistakes? It fixed mine, then made a worse one Tool Audit 5 posts
Honest assessments of AI tools used against real positions. What earns its place, what does not.
Latest: Is Grok good for stock research? I ran the same test on the fifth tool Prompt Stack 7 posts
The four-stage method for getting a reliable answer out of AI: scope, filter, risk, verdict.
Latest: AI quality of earnings review: 4 prompts to find the real profit Guardrails 11 posts
The guardrails that keep AI honest in a decision: the checklists and frameworks, and where AI doesn't get a vote.
Latest: Does ChatGPT just agree with you? Mostly no, but watch the numbers