Getting an Evidence-Graded Read on Your AI Citation Readiness
An evidence-graded audit of how AI systems actually see your site.
Guide
What demonstrably influences whether generative engines cite you, what's speculative, and what's pure folklore, from a team that implements this on production platforms.
AI systems cite pages they can crawl, parse, and quote. The working levers: allow their crawlers, serve content in raw HTML, structure pages as direct answers with valid schema, and publish original data worth citing. Nothing guarantees a citation; everything above measurably raises the odds.
The same four steps, in order, every time:
Check robots.txt, your WAF, and your bot manager for rules blocking AI user agents. The ones that decide search visibility are OAI-SearchBot, which OpenAI documents as surfacing sites in ChatGPT search, and PerplexityBot, which Perplexity documents as surfacing sites in its results. GPTBot is the separate training crawler. Many sites block these unknowingly via aggressive bot protection. Decide deliberately per agent; blocking training crawlers while allowing search crawlers is a legitimate, configurable choice.
Fetch your page without JavaScript and see what's there. None of the vendor documentation above promises JavaScript rendering, and in our audits content that only appears after client-side rendering is often missing from what AI systems retrieve. Server rendering or prerendering your key pages is the single highest-impact technical fix, and it is a core check inside the AI Search Readiness Audit.
Engines lift passages, so write liftable passages: a question as the heading, the answer in the first sentence, one idea per paragraph, tables for comparisons, and valid JSON-LD describing the page. This page's own structure is the template.
Engines cite sources that add information: real numbers, real configs, first-hand results. A page restating what fifty pages already say gives an engine no reason to cite it. This is the durable lever, and the only one competitors can't copy overnight.
This site's real robots.txt, allow-all for AI retrieval and training bots by name:
text
User-agent: GPTBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: ClaudeBot
Allow: /Named groups, not a wildcard, so the policy survives a future restrictive default and documents the decision.
An evidence-graded audit of how AI systems actually see your site.
Performance, security, SEO infrastructure, and AI-readiness, in one fixed-fee audit.
The full technical checklist behind AI citation readiness, scored against your site.
A plain-English glossary of the acronyms this guide assumes you already know, plus which one to ignore.
Why scattered name, address, and schema data undercuts citation odds even when the content itself is perfect.
Going deeper
Perplexity says PerplexityBot surfaces websites in its search results and is not used for model training. Allow it, be liftable, and publish original data. There is no documented separate Perplexity SEO beyond that.
GPTBot is OpenAI's training crawler. ChatGPT search results come from OAI-SearchBot, which is controlled separately, so blocking GPTBot does not remove you from ChatGPT search (OpenAI docs). Decide per crawler, deliberately, in robots.txt.
Some, and growing, but the bigger value today is presence: being the named source when a buyer asks an AI about your category. Measurement is immature; that's honest.
Unproven. Google's documentation says no AI text files or special markup are needed to appear in AI Overviews or AI Mode. llms.txt is cheap and harmless, and we deploy it graded as speculative. Steps 1 through 4 above have evidence; llms.txt has a hypothesis.
No. Anyone guaranteeing placement in generative answers is selling folklore. What can be guaranteed is the technical floor and the process, which is exactly what our audit covers.

Guide
The AI Search Readiness Audit checks every step above against your actual site, not a checklist you have to interpret yourself.