Skip to content
Guardrails

How to check if ChatGPT cites your site

The Bluff Filter is this kind of check, on one page. Take it with you →

// On this page

On 27 July I asked ChatGPT the simplest question anyone could ask about this site: what is dixon.ai. It came back with a UK training company that runs workshops and executive sessions for leadership teams, incorporated in 2021, based in Stockport. None of that’s me, and it’s why I now check whether ChatGPT cites your site every month.

There’s another business with a nearly identical name. Watch where the answer joins us together.

Ch ChatGPT · 27 July 2026 · logged out, asked “what is dixon.ai”

“If you’re referring to Dixon AI (dixon.ai / dixonai.com), it’s a UK-based AI consultancy and training company…”

Two unrelated domains, joined by one oblique stroke. From there on there is only one company left in the answer, and it isn’t this one. It ran for three more paragraphs and a bulleted service list.

A diagram showing dixon.ai and dixonai.com as two separate sites being merged into one incorrect ChatGPT company profile.
The entity error, mapped. The slash in ChatGPT’s opening line joined two unrelated domains. Everything after it described the other company.

I only knew because I’d been putting frozen questions to my own site since 11 July, written down in a file I don’t let myself edit between rounds. Three rounds used the same six topical questions. A fourth introduced six branded ones. Two rounds were signed in and two were logged out, so I keep those conditions beside every result. That’s the whole method, and it’s the only way I’ve found to check whether ChatGPT cites your site. It answers the question a lot of people are now asking: is any of this reaching an assistant at all?

The check: eight questions you never reword

1

Write the questions once, then freeze them. Change a word between rounds and you've run a different test, so the comparison you were building is worthless. Mine is on v2, and the version number is the point: changing one becomes a decision rather than a slip.

2

Half with your name on, half without. The branded half tells you whether an assistant knows who you are. The topical half tells you whether anyone finds you cold. Ask only the branded ones and you'll come away far too pleased with yourself.

3

Logged out, one question per fresh chat, web search on. I've visited my own site more than any human alive and a signed-in session knows it, which is why two of my four rounds are ones I can't lean on. And the three questions where ChatGPT never searched cited nothing at all: no links, no sources, just prose. With browsing off there's nothing there to check.

4
the one people skip

Record four things, the same way every time. Whether it linked you. Whether it named you without linking, which is a different result. Which domains it cited instead. And the full answer text with a timestamp, because in a month your memory will invent something tidier.

A blocked question is not a no. The first version I wrote logged Perplexity's sign-in wall as a clean "no citation", which would have shipped the sentence Perplexity doesn't cite us when logged out. It doesn't cite us because it doesn't answer us. A wall is a result. It just isn't that one.

The eight-question set

The fixed citation check arranged as four topical questions and four branded questions, with the wording locked between rounds.
Four cold questions, four named ones. The split stops brand awareness masquerading as topical reach. The lock makes next month comparable with this one.
// The frozen question set - fill it in once, then leave it alone

TOPICAL (what someone types who has never heard of you)

  1. which [TOOL CATEGORY] is most reliable for [YOUR SUBJECT]
  2. how do I check [THE THING YOU TEACH PEOPLE TO DO]
  3. state of [YOUR SUBJECT] 2026
  4. [YOUR SIGNATURE CLAIM, written out with your name left off]

BRANDED (what someone types who has already heard your name)

  1. what is [YOUR SITE]
  2. who is [YOUR NAME] of [YOUR SITE]
  3. is [YOUR SITE] a credible source on [YOUR SUBJECT]
  4. [YOUR SITE] [YOUR NEWSLETTER OR PRODUCT NAME]

Nobody arrives cold, and the file is how I know

One highlighted dot out of six for topical questions compared with four highlighted dots out of six once the site name appeared in the question.
Three topical rounds from 11 July, one branded round on 27 July. The only topical question that ever produced a citation was "AI reliability scoreboard", which is the one closest to being a name.

Of the five topical questions that never produced one, the two that did run a search came back with Stanford, arXiv and a DOI link, and nothing of mine. Here’s the one that worked.

Perplexity answering the query AI reliability scoreboard. A cited card at the top links dixon.ai/scoreboard and carries the site's own meta description word for word, ringed in orange and labelled MY OWN META DESCRIPTION. The answer below opens by naming the scoreboard, then attributes the rest to the site with the phrase It says.
Perplexity, 21 July 2026. The card at the top is my own meta description, quoted back at me word for word. Then, one sentence into the answer, it stops describing and starts attributing: It says the scoreboard uses real questions. Being cited turns out to mean being repeated. And the round was signed in, which is exactly why I can't lean on it.

What changed when I used the name

Then I froze a second set that named the site outright and ran it on 27 July. On four of the six branded questions, ChatGPT went and found me. Asked flat out whether dixon.ai is a credible source, it said yes and then volunteered the objections before I could: no peer review, small samples, testing “largely conducted by one author (Ben Dixon)”. Its verdict was a credible independent technical blog, not an authoritative source on its own. All of that is true, which is a peculiar thing to be marked down for by the machine you’re marking.

Which leaves the control question. I took this site’s own thesis and stripped the name off: has anyone tested which AI assistant bluffs the most. That is precisely what we do here. ChatGPT answered at length and didn’t mention me once. Which is fair enough. I’m not in Nature.

The unbranded control question linked to six sources: ACL Anthology, a GitHub leaderboard, arXiv, Reddit, Tom's Guide and a Nature paper, while dixon.ai is marked not cited.
ChatGPT, logged out, 27 July 2026. Six sources for a question this site exists to answer, and not one of them is me. It reached a forum and a consumer tech site before it reached the site built on the question.

The name opened the door four times in six. The topic managed one.

It’s the name that’s broken, not the ranking

// What the record is telling me

It has us. It just has us filed under somebody else.

Four questions in that same round cited the site happily and described it accurately, so the assistant isn't missing us. The name resolves to the wrong company, and once it's resolved wrongly it answers about that company with total composure.

That decides where the effort goes. A ranking problem sends you back to your own content. An entity problem sends you off your site entirely, to wherever a machine could learn which one of you is which.

The answer names those places for you. The two sources ChatGPT read to describe the other company were that company's own site and Companies House, so those are the two places my name is losing. My about page already carries the structured markup that's meant to prevent this. It carried it the whole time, and it didn't help.

Four faster ways, and what each one misses

Running a frozen set by hand is the slow way, and it’s the one that answers the question properly.

misses most of it

Referral traffic. Every ChatGPT link I saved on these rounds carried utm_source=chatgpt.com, so that's the thing to filter on. The referrer is less reliable than it looks: open a citation inside the phone app and it lands in your analytics as direct traffic. Either way you only ever count the people who clicked.

upstream of the question

Server logs. OpenAI publishes the names and roles: OAI-SearchBot is its search crawler, GPTBot is for model training, and ChatGPT-User fetches pages for some user-requested actions rather than crawling automatically. A hit tells you which route fetched a page, not whether an answer cited it.

this method, billed monthly

A paid tracking tool. Scheduled queries, counted for you at scale. For one site, the twenty minutes is cheaper than the subscription.

google only

Search Console. On 3 June 2026 Google began rolling out a generative-AI report to a subset of sites. It shows impressions in AI Overviews and AI Mode, split by page, country, device and date. It does not show clicks, queries or answer text.

The three passive signals do not preserve what the assistant said about you. A paid tracker can automate repeated questions. This small file does the same job transparently. On my own site, the wording was the entire finding.

A comparison of referral traffic, server logs, paid tracking tools and Search Console showing what each reveals and what it misses.
The gap shared by the shortcuts. They can show a click, a fetch, a repeated query or a Google impression. The frozen record is the one that preserves the answer itself.

Where it doesn’t help, and the bits I got wrong

Four assistants are in my protocol. On the round that matters most, this is as far as it reached:

  • ChatGPT6 answeredThe only engine that answered anything logged out. Every number in this post is its number.
  • Perplexity6 walledSix questions, six sign-up forms, nought answers.
  • Gemininever askedMy protocol wrote off signed-out access without testing its location and device limits.
  • Claudenever askedAlso written off without a test. Untested, not unavailable.
A coverage chart showing ChatGPT with six answered questions, Perplexity with six sign-up walls, and Gemini and Claude as untested.
This is the denominator. Six ChatGPT answers support the numbers in this post. Six Perplexity walls and two untested assistants do not.

The limits are part of the result

Pe Perplexity · 27 July 2026 · logged out, asked “who is Ben Dixon of dixon.ai”

“Sign up and repeat your request.”

That is the whole answer, word for word identical on all six questions, under a heading marked “Answer” and above an empty “Sources” list. Log it as a no and you have invented a finding.

Which is why “four of six” means four of ChatGPT’s six. Google’s current help page says some Gemini web features work without sign-in, while availability depends on location and device. I had written Gemini and Claude off without testing either route. The honest result is not that both were unavailable. It is that I never asked them.

Two of my four rounds were signed in. They stay as dated observations, but not as cold evidence. A signed-in session may be affected by account history or personalisation, so I do not use its citations or misses to claim what a stranger would see. That is the entire reason the signed-out track exists.

One run per question, which is an observation and not a rate. There’s no N of 3 behind any of this, and my own script was keyed in a way that let re-runs overwrite each other, so I can’t tell you how stable that one citation was either. If someone hands you a percentage for AI citation off a handful of runs, ask how many times they ran it.

It measures citations, not crawls. Reading your site and naming you to a reader are two different questions, and I keep them in two different files.

The short version

What worked: Freezing the questions. Three rounds of the same topical six since 11 July produced one comparable record instead of three separate impressions, and the second frozen set caught ChatGPT describing a different company under my name on 27 July. No amount of staring at my analytics would have shown me that.

What I’d do differently: Run every round logged out. Two of mine weren’t, and those are the two I’d have to caveat if anyone asked.

What didn’t: The topical set has returned almost nothing since day one. Across three rounds, one of its six questions produced a citation at least once, and a control question carrying my exact thesis without my name on it fetched academic papers instead. That’s the honest reading of where a young site sits.

Bottom line: Worth twenty minutes a month, on the condition that you keep the wording frozen and log the blocked questions as blocked. It won’t get you cited. It tells you whether you are, which turns out to be the harder thing to find out.

The cue is the first of the month and a browser window with nothing signed in. This check reads the citations an assistant hands other people about you. The companion check runs the other way, on the citations it hands you: how to check if a ChatGPT citation is real or fake. Every mistake I’ve logged, this one included, is in the record of AI errors, and the rest of the paste-in checks are on the guardrails hub.

Next round is the same two frozen sets, the same file, and logged out from the first question to the last. I’ll find out on 1 September whether an assistant has worked out which Dixon is which.

Common questions

How do I know if my content is cited in ChatGPT answers?
Ask a fixed set of questions in fresh chats, signed out where available, and record whether each answer links your site, names it without a link, or does neither. Save the full answer and repeat the same wording every month. Normal analytics will not preserve what the assistant said.
Why does it have to be logged out?
A signed-out fresh chat reduces the chance that account history or personalisation affects the answer. It is an approximation, not proof of what every stranger sees, and availability varies by service, location and device. If the route is unavailable, record it as untestable rather than silently changing the method.
What if the assistant refuses to answer?
Record it as blocked, with the reason, and never as a no. On my last round Perplexity answered none of its six questions: every one came back as a sign-up form. It did not decline to cite me. It declined to answer, and those are different results.
Ben Dixon
// Written by Ben Dixon

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →

// Keep reading
Guardrails

Why does ChatGPT make up sources? Two gov.uk links, only one held the rule

With web search on, ChatGPT's links are real. The failure that gets past you is a working link to a page that doesn't hold the claim. Here's why it happens.

Guardrails

Does ChatGPT just agree with you? Mostly no, but watch the numbers

Does ChatGPT just agree with you? Mostly no. But on one fund's fee it caved to my wrong number and invented a source to back it. The 30-second check.

Guardrails

How to check if a ChatGPT citation is real or fake

Looking for a ChatGPT citation checker? The dangerous citation is not the dead link. It is the working link, to a real page, that does not back the claim.

// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in Guardrails →