AI & LLM Visibility 2026-10-09 5 min read

Getting Cited by AI: How ChatGPT, Perplexity and AI Overviews Choose Sources

There is no position one in an AI answer. A page is either retrieved and quoted, or it does not exist. The mechanics are close to ordinary SEO, but they fail in different ways — and the most common failure is a robots.txt rule that quietly removes you from AI search results.

AI Search Visibility: How to Get Cited

Key takeaways

  • AI assistants select and quote sources rather than rank them, so the question is inclusion or exclusion — not position.
  • The single most common self-inflicted problem is a robots.txt wildcard that blocks every AI crawler, which removes a site from AI search results as well as from model training.
  • Search crawlers and training crawlers are separate agents. Blocking training (GPTBot) does not block being cited in search results (OAI-SearchBot), and vice versa.
  • Pages get quoted when they contain extractable, attributed facts — a number with a source, a dated statement, a definition — inside a document a machine can parse without running JavaScript.
  • Entity clarity matters more than it does in classic SEO: consistent naming, an about page, address details and structured data let a model resolve who you are before it decides to cite you.

AI search visibility is a different problem from ranking, and AI assistants changed its shape. In classic search, the goal was a position on a results page. In an AI answer, there are no positions — a page is either retrieved and quoted, or it is not mentioned at all. That makes the mechanics familiar but the failure modes different.

Most of what determines whether you get cited is ordinary technical hygiene. A few parts are specific to how retrieval-augmented systems work, and one of them is a mistake that removes sites from AI search results without any visible symptom.

Ranking versus retrieval

When someone asks an assistant a question, the system usually does three things: it converts the question into a search, it retrieves a small set of candidate pages, and then it reads those pages and writes an answer with citations. Your page has to survive all three steps.

This is why a page can rank well on Google and still never be quoted. Ranking gets you into the candidate set. Being quotable is what gets you into the answer. They are related properties, but they are not the same one, and the second is the one most sites neglect.

The robots.txt layer, and the mistake in it

AI systems reach pages through named crawlers, and the names matter because they do different jobs:

CrawlerOperatorWhat it feeds
GPTBotOpenAIModel training
OAI-SearchBotOpenAIChatGPT search results
ChatGPT-UserOpenAIOn-demand page fetch when a user asks
ClaudeBot / Claude-UserAnthropicTraining and on-demand fetch
PerplexityBot / Perplexity-UserPerplexityIndex and on-demand fetch
Google-ExtendedGoogleGemini and Vertex grounding — not Google Search
BingbotMicrosoftBing index, and the material Copilot draws on
Applebot-ExtendedAppleApple Intelligence training
CCBotCommon CrawlAn open dataset many models are trained from
BytespiderByteDanceByteDance AI systems

The mistake is treating these as one group. A site that writes a single wildcard rule blocking all of them has decided not only “do not train on our content” — a defensible position — but also “do not show our content in AI answers”, which for most B2B companies is a straightforward commercial loss. If you want to opt out of training but still be cited, block GPTBot, ClaudeBot, CCBot and Applebot-Extended, and leave OAI-SearchBot, PerplexityBot, ChatGPT-User and Bingbot alone.

A related trap: Google-Extended does not control Google Search. Blocking it removes your content from Gemini and Vertex grounding while leaving your ordinary Google rankings untouched. Some publishers block it believing they have opted out of Google entirely, and are surprised to still be ranked.

Note on llms.txt: it is a proposed convention, not an adopted standard, and there is no evidence major assistants use it as a selection signal today. Publishing one is cheap and harmless. Treating it as a strategy is not.

What makes a page quotable

Once a page is in the candidate set, an extraction pipeline reads it and decides what to use. Pages that get quoted share a set of properties:

  • A stated fact early in the document. A number, a date, a definition, a named entity — in the first few hundred words.
  • Attribution. “Approximately 55% of measured search traffic in China, according to StatCounter-derived estimates” is quotable. “Most of the market” is not, because nothing in it can be carried into an answer.
  • Headings that read like questions or topics. H2 and H3 headings that paraphrase how people actually ask things make the relevant passage easy to isolate.
  • Tables and lists. Structured comparison is the single most extractable content format, and it is the format that most often survives into an AI answer.
  • Server-side rendering. If the text only exists after JavaScript runs, many extraction pipelines see an empty document. This is the most common reason a modern site is invisible to AI systems despite looking fine in a browser.
  • Stable URLs and visible dates. Machines weight recency, and they need a URL they can cite and a date they can attach to it.

Entity clarity

Before quoting a page, a model has to know who is behind it. Ambiguity here is expensive, and it is where small companies lose to large ones for reasons that have nothing to do with content quality.

The fixes are unglamorous:

  • Use one exact company name everywhere — on the site, in the footer, in the page title, in press releases, in schema.
  • Publish an about page that states legal name, registration details, founding date, address and contact. Machines resolve entities from exactly the kind of facts a legal page contains.
  • Add Organization schema with sameAs links to profiles you control, so the model can connect the name to a consistent identity.
  • Keep a physical address and a working telephone number on the contact page. This sounds like a formality; it is one of the strongest trust signals available to a small site.

What to do about freshness

Assistants weight recent material, and retrieval systems prefer pages that are updated rather than duplicated. In practice this means updating an existing page with a new date and new figures beats publishing a near-identical second page — the second page competes with the first and neither accumulates authority.

For a company publishing regularly, the compounding asset is not the number of pages. It is the number of pages that state a specific, attributed fact which no one else has published.

Measuring AI search visibility when there are no rankings

There is no equivalent of a rank tracking tool, and the workaround is direct: write down the twenty questions your buyers actually ask, run them monthly across ChatGPT, Perplexity, Copilot and Google's AI results, and record whether you appear and which sources do.

The list of sources is the useful output. It tells you which domains the model trusts on your topic — and presence on those domains is usually a faster route to being cited than optimising your own pages further.

Frequently asked questions

Should I block AI crawlers on my website?

It depends on whether you want to be found by AI assistants. If being cited matters commercially — which it does for most B2B companies — you need the search-side crawlers to have access. Many sites block everything with a single wildcard rule and remove themselves from AI search results without realising it, because the training crawler and the search crawler have different names and different jobs. Decide the policy deliberately, per agent, rather than inheriting it from a template.

Does an llms.txt file improve AI visibility?

It is a proposed convention rather than an adopted standard, and there is no evidence that major assistants currently use it as a ranking or selection signal. Publishing one costs almost nothing and may be useful later, but it should not be part of a plan you expect results from. Crawler access, page structure and entity clarity are where the measurable gains are.

Does structured data help with AI answers?

It helps machines parse who and what your page is about, which is the part that has to work before content selection happens. Organization, Article and FAQ markup do not guarantee citations, but they remove ambiguity — and ambiguity is expensive when a model is deciding which of ten similar pages to trust. The effect is indirect and real: it raises the probability that your page is interpreted correctly.

How do I measure AI visibility if there are no rankings?

Ask the assistants directly. Build a list of the twenty questions your buyers would ask — including your product category and your competitors — and run them monthly across ChatGPT, Perplexity, Copilot and Google's AI results, recording whether you are cited and which sources are. It is a manual process, but it produces something more useful than a rank position: the actual set of pages the model is drawing on, which tells you which sites to be present on.

Why does my page never get quoted even though it ranks on Google?

Usually because the content is not extractable. A page that renders its text with JavaScript, or states its facts across long narrative paragraphs with no headings, is difficult for an extraction pipeline to use. Pages that get quoted tend to state the fact early, put it in a heading or a table, and attribute it. Ranking and quotability overlap, but they are not the same property.

CN NEWSWIRE Editorial Team

Written by the CN NEWSWIRE editorial team in Hangzhou. We place press releases and articles in Chinese media for international clients, and publish what we learn here.

Your story next

Ready to reach Chinese media?

Distribute your announcement today — guaranteed placement on top outlets.