Which AI engines can read your site?
Your robots.txt decides two separate things: whether AI assistants can cite you, and whether they can train on you. The two are easy to mix up. Paste your URL and see where you stand.
Can cite you in AI answers
These crawlers are how assistants find and link you. Blocking one takes you out of that product's answers, and none of them are the crawler that controls model training.
Can train on your content
These feed model training. Blocking them is a legitimate choice and costs you nothing in search or in AI answers.
What this checker reads
It fetches your robots.txt and reads it the way a crawler does, following the rules of RFC 9309. For each AI crawler in the list it works out which group of rules applies: a group that names the crawler, or the User-agent: * group that catches everything else. Then it reports the verdict for that crawler: allowed, blocked, or partly blocked with the paths it is kept out of.
It also asks for /llms.txt and tells you whether one is published. Nothing is scored and nothing is guessed. Every crawler in the list comes with a link to the page where its operator documents it.
Two kinds of AI crawler, and why the split matters
The crawlers that put you in AI answers and the crawlers that collect training data are different user agents, even from the same company. OpenAI says so in its crawler documentation: "Each setting is independent of the others", and gives the example of allowing OAI-SearchBot "to appear in search results while disallowing GPTBot". Blocking the search bot has a cost: "Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers".
Anthropic documents the same split. Disabling Claude-SearchBot "may reduce your site's visibility and accuracy in user search results", while blocking ClaudeBot "signals that the site's future materials should be excluded from our AI model training datasets". Perplexity states that PerplexityBot "is not used to crawl content for AI foundation models".
Google is a special case. Google-Extended is a token in robots.txt, with no crawler of its own, that controls whether pages Google already crawls "may be used for training future generations of Gemini models" and for grounding. Google states that it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search". The checker lists it under training for that reason.
So the top of the page shows two numbers: how many answer crawlers can see you, and how many training crawlers are allowed. Most site owners want the first number high. The second is a choice with no search cost either way.
How to fix what it finds
A blocked answer crawler usually comes from an old rule written against every bot at once. Give the crawlers you want back in a group of their own, and keep the training rules separate. This is the pattern for a site that wants to be cited and does not want to train models:
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: *
Allow: /OAI-SearchBot, Claude-SearchBot and PerplexityBot have no rule of their own in that file, so the User-agent: * group applies and they are allowed. Changes are not instant: OpenAI says "it can take ~24 hours from a site's robots.txt update for our systems to adjust", and Perplexity says "it may take up to 24 hours".
A "partly blocked" verdict names the paths a rule keeps that crawler out of. If the list holds your content, move those paths out of the rule. Cloudflare's /cdn-cgi/ is ignored by the checker because it never holds content.
A "broken" robots.txt means the address returned an HTML page instead of a text file, usually a 404 page served with a 200 status. Crawlers cannot read rules from it. Serve a plain text file, even an empty one.
What robots.txt cannot do
Some fetchers act on a person's request and say they do not follow the file. OpenAI writes of ChatGPT-User that "because these actions are initiated by a user, robots.txt rules may not apply". Perplexity writes of Perplexity-User that "since a user requested the fetch, this fetcher generally ignores robots.txt rules". The checker marks those agents so a green verdict is not read as a guarantee.
It is also not a way to hide a page from search. Google's own words: robots.txt "is not a mechanism for keeping a web page out of Google. To keep a web page out of Google, block indexing with noindex or password-protect the page." Anthropic adds that blocking its IP addresses "may not work correctly or persistently guarantee an opt-out", because that also stops the bot from reading your robots.txt.
About llms.txt
llms.txt is a proposal by Jeremy Howard, first published in September 2024 and revised as v2 in August 2026: "a /llms.txt markdown file to websites to provide LLM-friendly content", with "brief background information, guidance, and links to detailed markdown files". The checker reports whether your site publishes one.
No search engine or AI company has said it reads the file, and we found no evidence that it changes rankings or citations. It is a small courtesy to agents visiting your site, nothing more. We wrote up what we saw across thirty SaaS sites in llms.txt: what it is, and whether your site needs one.
Questions people ask
- Does blocking GPTBot take me out of ChatGPT? Not on its own. OpenAI documents the search opt-out as a separate rule for OAI-SearchBot. Block GPTBot and leave OAI-SearchBot alone to stay in ChatGPT search while opting out of training.
- Will Google-Extended hurt my Google rankings? No. Google states it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal".
- The checker says allowed, but my robots.txt has no AI rules at all. That is correct. A crawler with no rule against it may read everything, so a short robots.txt is an open door. If that is what you want, you are done.
- Which crawlers are on the list? Only agents their operators document publicly: OpenAI, Anthropic, Perplexity, Google, Apple, Meta, Mistral, Common Crawl and Amazon. Each row links to the source.
- Is it free? Yes. No sign-up, no email, and nothing is stored beyond the request.
Read more
- Controlling what AI crawlers see: each crawler, what it feeds, and the rules to write.
- robots.txt in our wiki, with the syntax and the common mistakes.
- Answer Engine Optimization: how to get cited in AI search: being readable is the first step, this is the rest.
Being readable is the easy half.
Letting AI engines in only matters if there is something worth citing when they arrive. BeeRanked publishes fast, clearly structured pages on your own domain, with the structured data search and AI engines read, on every page, automatically.
Start with BeeRankedOr try the structured data analyzer to see what machines understand about your brand.