The Citation Desk.If the AI search crawlers can't read your site, they can't recommend it — no matter how good your content is. This checks whether your robots.txt is quietly blocking them.
Paste your site's robots.txt below. Find it at yoursite.com/robots.txt. Nothing is sent anywhere — the check runs entirely in your browser.
AI assistants answer questions by reading the live web through their own crawlers. If your robots.txt disallows a crawler, that engine can't fetch your pages — so it can't cite or recommend you, even if you're the best answer available. It's the most common self-inflicted reason a brand is invisible in AI answers.
Each engine uses its own crawler, and they're independent. Blocking one doesn't block the others. And crucially, the crawler that collects training data is separate from the one that fetches pages for live search citations — so it's easy to block the wrong one by accident and remove yourself from answers without realising.
This tool reads your robots.txt the way a crawler would — checking the rules that apply to each bot by name, and the catch-all rules that apply to everyone. It's a quick first check, not a full audit: being allowed to crawl is the floor, not a guarantee of being cited.
robots.txt is a small text file that sits at the root of your website, at yoursite.com/robots.txt. It's a set of instructions telling automated visitors, called crawlers or bots, which parts of your site they're allowed to visit. Search engines and AI engines send crawlers to read the web, and the first thing a well-behaved crawler does is check your robots.txt to see if it's welcome.
Think of it as a sign on the door. It doesn't lock anything, it just tells visitors the rules. Which leads to the honest question most tools skip.
Mostly yes, with honest caveats. robots.txt is a voluntary standard, formally RFC 9309. It works because crawlers choose to respect it, not because it physically blocks anything. The good news is that the major AI crawlers document their bots openly and generally honor robots.txt, OpenAI's GPTBot and OAI-SearchBot, Anthropic's ClaudeBot and Claude-SearchBot, PerplexityBot, and Googlebot all follow standard directives.
Three honest caveats worth knowing. First, compliance is opt-in, a few crawlers, like ByteDance's Bytespider, have a documented history of ignoring robots.txt entirely, so the file is a norm, not a guarantee. Second, real-time fetches are a gray zone, when a person asks ChatGPT or Perplexity to read a specific link, the tool fetches it on the user's behalf, and providers may treat that as user-directed access rather than crawling, so your robots.txt rules may not apply the way you'd expect. Third, for guaranteed blocking you need server or CDN-level rules, robots.txt alone won't stop a determined or non-compliant bot.
The practical takeaway, for the goal most brands care about, being visible to AI, robots.txt works well. The major engines respect it, so if you accidentally disallow their search crawlers, you really will be excluded, and OpenAI states plainly that sites blocking OAI-SearchBot will not appear in ChatGPT search answers. That's exactly the accidental self-block this tool helps you catch.
Most site owners never look at their robots.txt, and when they do, the rules are cryptic. This tool reads your robots.txt the way a crawler would, checking the directives that apply to each AI search bot by name, plus the catch-all rules that apply to everyone, and tells you in plain language whether each major AI engine can currently access your site. It's the fastest way to catch the single most common, and most costly, mistake, quietly blocking the crawlers that decide whether AI can recommend you at all.
No. The AI training and search crawlers, GPTBot, ClaudeBot, PerplexityBot, are separate from Googlebot. Blocking them has no effect on your normal Google search rankings. Blocking Google-Extended only affects Gemini's training, not Google Search.
No, and this trips up a lot of people. Each bot needs its own directive. Allowing ClaudeBot does nothing for Claude-SearchBot or Claude-User. The crawler that gathers training data and the one that fetches pages for live citations are different, so list each one you want to allow.
Yes, and it's a common 2026 strategy. You can block the training crawlers, GPTBot, Google-Extended, ClaudeBot, CCBot, while allowing the search and user crawlers, OAI-SearchBot, Claude-SearchBot, PerplexityBot, so you stay eligible for AI-search citations without feeding model training.
No. Allowing the crawlers is the floor, it makes you eligible to be cited, but it doesn't guarantee it. Whether AI actually names you depends on what it finds, clear answer-first content, original data, and third-party mentions on sources the engines trust. Access is step one, not the whole job.
If you don't have one, most crawlers assume everything is allowed, which is usually fine for AI visibility. The risk is having a robots.txt that accidentally disallows the AI bots, often left over from an old setup or a plugin default. That's the case this tool is built to catch.
Letting the crawlers in is the floor. Whether AI actually recommends you depends on what it finds. The Citation Desk measures where your brand stands across all five AI engines and what to do about it.
See where you rank →