Most sites that block AI crawlers have no idea they do. We check your robots.txt against all 34 AI user agents, send live requests as real crawlers to see what your server actually returns, and detect noindex directives that suppress your pages after the crawl.
| Agent | robots.txt | Response | Verdict |
|---|---|---|---|
| GPTBot | Allow | 200 | Reachable |
| ClaudeBot | Allow | 200 | Reachable |
| PerplexityBot | Allow | 403 | Blocked |
| Amazonbot | Allow | 429 | Blocked |
| CCBot | Disallow | — | Your rule |
MismatchTwo crawlers you meant to welcome are being turned away.
Indexability/pricing answers every agent with a 200 and carries noindex. Crawlable, still suppressed.
robots.txt is a request, not a gate. Checking the file tells you what you asked for. It tells you nothing about what happened.
User-agent: PerplexityBot Allow: /
The crawler is welcome.
GET / — User-Agent: PerplexityBot 403 Forbidden — via Cloudflare
The crawler never got in.
websites studied in The State of AI Visibility 2026
of 1,597 reports — 11.6% — blocked at least one AI crawler
of those blocks appear nowhere in robots.txt
The gap between the two is the finding.
All 34 agents hit your URL for real — status code, timing, redirect chain, any Retry-After.
Your robots.txt is parsed separately, per user-agent, including wildcards that catch bots you never named.
Every disagreement between the two, with the layer of your stack that caused it.
Published in full rather than truncated. A robots.txt written two years ago does not account for most of these.
Shapes whether you exist in the next generation of models.
Decides whether you can be cited in an answer today.
A person is on the other end, trying to use your page right now.
Classic search infrastructure. Blocking these hits traditional rankings too.
Deep-research features that read sources at length before answering.
Most people assume blocking lives in robots.txt, where you can see it and change it. It usually doesn’t. You open the site in a browser and everything works fine. The AI bot hits a closed door and leaves.
The State of AI Visibility 2026 — 706 websites, 1,597 reports. Read the study →
Run the check on any URL. You get the full 34-agent breakdown, every mismatch between policy and runtime, and the specific fix for each one.
Thirty-four, grouped by the job they do: 5 search crawlers, 8 browsing agents that fetch a page because a person asked for it, 8 training crawlers, 2 research crawlers, and 11 Google crawlers. The group matters more than the count — blocking a training crawler and blocking a search crawler have completely different consequences.
That is half of it. Your robots.txt is parsed for all 34 crawlers. Separately we send live requests using real crawler user agents — one from each category, alongside a plain browser for comparison — so we can see what your server does rather than what your rules say. Those two answers disagree more often than people expect, and when they do, the rules are not the problem.
Yes, and this is the most common version of the problem. Bot protection on a CDN, a WAF rule aimed at “AI scrapers”, a security plugin’s blocklist, a rate limiter tuned for human traffic, or a managed host running mitigation you cannot see from the admin panel. None of it is visible to a visitor, and none of it is in your robots.txt.
That is your call, but make it per group rather than all at once. Blocking training crawlers keeps your content out of model training and costs you nothing in citations. Blocking search and browsing crawlers removes you from the answers themselves. A single User-agent: * rule does both, which is why most blocks we find were never intended.
No. A crawler not mentioned in robots.txt is allowed. Having no file at all means everything is allowed — the risk of an absent robots.txt is that you have no way to be deliberate, not that you are blocking by default.
Robots.txt controls whether a crawler may fetch the page. Noindex tells it not to keep what it found. They fail in different places, so a page can answer every crawler with a healthy 200 and still be suppressed afterwards by a noindex nobody remembers adding — often injected at the CDN as a header rather than written into the page.
Most do. It is a convention, not a technical barrier — there is nothing stopping a crawler that ignores it. Treat robots.txt as the way you state your intent, and infrastructure rules as the way you enforce it.
It depends which kind. Blocked in robots.txt is stated as fact, because a compliant crawler obeys the file whatever a live fetch shows. A live refusal is different: we send from our own infrastructure, and security tools weigh the IP and request fingerprint as well as the user agent. We have seen one page, within seconds, return 403 to a browser, 200 to PerplexityBot and 520 to GPTBot. So we call that a risk worth treating as critical, not a confirmed block — your server logs are what confirm it.
Live checks measure what your infrastructure did at that moment. A 429 or a 5xx can be a transient rate limit or a one-off WAF decision rather than a standing rule. Re-run to see whether it holds.
Blocking the AI-specific crawlers does not affect Google’s ranking of your pages. The mistake that does hurt is blocking Googlebot while allowing Google-Extended: standard indexing has to work before Google’s AI features can include you at all.