Skip to content
// Guardrails · the toolkit

Run these seven checks before you trust an AI answer.

A guardrail is just a short instruction you paste into your own AI so it flags its guesses. Without one, those guesses slip in alongside the facts. The industry sells companies a fortune in fencing for a cliff edge you can rope off yourself. Below are seven, each one built from a real failure I caught and screenshotted.

Seven checks, seven receipts

None of these makes the model perfect. Nothing does, and someone always has to check the checkers. What they do is make it tell you when it's on thin ice, which is most of what you actually need.

01

The Source Check

For any factual claim, name the source and say whether you retrieved it or are recalling it from training. If you cannot name a real one, say so.

Catches: Fabricated or misattributed citations. A link is not a fact.

ChatGPT quoting the ISA transfer rule as “HMRC says”, with a raw unresolved cite marker in the text, boxed in orange.
ChatGPT, 20 Jun 2026 — the right rule, quoted with confidence, pinned to a gov.uk page about moving abroad. The check above is what catches it.
The receipt → ChatGPT cited the wrong government page for an ISA rule The full write-up → How to check a citation in about ten seconds
02

The Push-Back Test

If I challenge your answer, re-check it against a source before you change it. Do not switch just because I sounded doubtful.

Catches: Caving under pressure, then inventing a new justification for the new answer.

The receipt → I told five AIs they were wrong. Watch which ones folded The full write-up → Does ChatGPT just agree with you? Tested
03

The Made-Up-Number Check

If you do not have live data for a figure, refuse it. Say "I do not have that" rather than generating a plausible number.

Catches: Invented numbers presented as fact. The most expensive failure in the log.

The receipt → Two share prices that never traded, quoted with confidence
04

The Date Check

State the date each figure is good as of, and your training cutoff. For anything that changes over time, tell me you cannot confirm it is current.

Catches: Stale data served with fresh confidence.

The receipt → A pre-2024 ISA rule handed back as if it were still current

Free when you subscribe

The Bluff Filter

A paste-in prompt that makes any AI flag what it’s guessing before you act on it. Learn it once, use it on every answer.

Plus one email a fortnight, showing where an AI went wrong.

// The list continues · 05–07
05

The Wrong-Subject Check

Before you answer, restate exactly what you think I am asking about, so I can catch it if you have the wrong thing.

Catches: Fluent, confident analysis of the wrong thing entirely.

The receipt → Gemini audited my site and reviewed a different business
06

The Unit Check

Quote the unit line verbatim (thousands? millions?) and sanity-check the magnitude before you build anything on a figure.

Catches: A number that is right, in the wrong order of magnitude.

The receipt → Perplexity read $6.1m as $6K, then narrated a fake collapse
07

The Second-Model Habit

Run the same question past a different model. Where they disagree is where you check.

Catches: Over-trusting a single tool. This one is a habit you build.

The receipt → The Scoreboard: six models on the same checkable questions

Want these as one paste-in, instead of seven?

The Bluff Filter folds the checks into a single five-stage instruction set: fix what the answer is for, scope the sources, split fact from guess, name the risk, land a verdict. You paste it once at the top of a chat and it sticks. Same model. It just flags its guesses so you can catch them. It's free. And the asking half of the method, the Prompt Stack, is how to phrase the question before any of this is needed.

Get the Bluff Filter →
// The evidence behind the checks

Every receipt here is a specific, dated moment I can re-run and you can check. That's the whole point: the checks come from watching these tools get real money questions wrong, then working out the one instruction that would have caught it before it caught me. The full write-ups behind these checks live under Guardrails.