Crawl access audit

Can AI crawlers access your site?

Most sites that block AI crawlers have no idea they do. We check your robots.txt against all 34 AI user agents, send live requests as real crawlers to see what your server actually returns, and detect noindex directives that suppress your pages after the crawl.

example.com
Access checkLive scan
Agentrobots.txtResponseVerdict
GPTBotAllow200Reachable
ClaudeBotAllow200Reachable
PerplexityBotAllow403Blocked
AmazonbotAllow429Blocked
CCBotDisallow—Your rule

MismatchTwo crawlers you meant to welcome are being turned away.

Indexability/pricing answers every agent with a 200 and carries noindex. Crawlable, still suppressed.

The problem

Your robots.txt says yes. Your server says no.

robots.txt is a request, not a gate. Checking the file tells you what you asked for. It tells you nothing about what happened.

What you published/robots.txt
User-agent: PerplexityBot
Allow: /

The crawler is welcome.

≠
What your server returnedlive request
GET / — User-Agent: PerplexityBot
403 Forbidden — via Cloudflare

The crawler never got in.

InvisibleNothing in your robots.txt file records this. Bot protection, WAF rules, rate limiters and CDN edge rules all cause it, and every one of them sits above the file you were reading.
706

websites studied in The State of AI Visibility 2026

185

of 1,597 reports — 11.6% — blocked at least one AI crawler

96%

of those blocks appear nowhere in robots.txt

How the check works

We measure both records, then compare them.

The gap between the two is the finding.

  1. 01

    Send live requests

    All 34 agents hit your URL for real — status code, timing, redirect chain, any Retry-After.

  2. 02

    Resolve the policy

    Your robots.txt is parsed separately, per user-agent, including wildcards that catch bots you never named.

  3. 03

    Compare

    Every disagreement between the two, with the layer of your stack that caused it.

Per-agent detail
Blocking each crawler costs you something different.
A training crawler shapes whether you exist in the next model. A search agent decides whether you can be cited in an answer today. A browsing agent is a person, right now, trying to use your page. They are reported separately because the consequence of blocking each one is not the same.
yoursite.com
yoursite.comChecking 34 bots…
🤖GPTBot✓ Allowed
🤖ClaudeBot✓ Allowed
🤖PerplexityBot✗ Blocked
🤖Amazonbot✗ Blocked
2 of 4 blocked — both of them allowed in robots.txt
Coverage

All 34 agents, and what each group costs you

Published in full rather than truncated. A robots.txt written two years ago does not account for most of these.

Training 8

Shapes whether you exist in the next generation of models.

  • GPTBot
  • ClaudeBot
  • Google-Extended
  • CloudVertexBot
  • Amazonbot
  • Applebot-Extended
  • FacebookBot
  • CCBot

Search 5

Decides whether you can be cited in an answer today.

  • OAI-SearchBot
  • PerplexityBot
  • Claude-SearchBot
  • DuckAssistBot
  • Bravebot

Browsing 8

A person is on the other end, trying to use your page right now.

  • ChatGPT-User
  • Perplexity-User
  • Claude-User
  • GoogleAgent-Mariner
  • MistralAI-User
  • facebookexternalhit
  • Meta-ExternalAgent
  • meta-externalfetcher

Crawling 11

Classic search infrastructure. Blocking these hits traditional rankings too.

  • Googlebot
  • Googlebot-Mobile
  • Googlebot-Image
  • Googlebot-Video
  • Googlebot-News
  • Storebot-Google
  • Storebot-Google-Mobile
  • GoogleOther
  • GoogleOther-Mobile
  • GoogleOther-Image
  • GoogleOther-Video

Research 2

Deep-research features that read sources at length before answering.

  • Gemini-Deep-Research
  • Google-NotebookLM

Most people assume blocking lives in robots.txt, where you can see it and change it. It usually doesn’t. You open the site in a browser and everything works fine. The AI bot hits a closed door and leaves.

The State of AI Visibility 2026 — 706 websites, 1,597 reports. Read the study →

Find out what the crawlers see.

Run the check on any URL. You get the full 34-agent breakdown, every mismatch between policy and runtime, and the specific fix for each one.

FAQ

Crawler access, answered

Which AI crawlers do you check?

Thirty-four, grouped by the job they do: 5 search crawlers, 8 browsing agents that fetch a page because a person asked for it, 8 training crawlers, 2 research crawlers, and 11 Google crawlers. The group matters more than the count — blocking a training crawler and blocking a search crawler have completely different consequences.

Isn’t this just reading my robots.txt?

That is half of it. Your robots.txt is parsed for all 34 crawlers. Separately we send live requests using real crawler user agents — one from each category, alongside a plain browser for comparison — so we can see what your server does rather than what your rules say. Those two answers disagree more often than people expect, and when they do, the rules are not the problem.

My site works fine in a browser. Can crawlers still be blocked?

Yes, and this is the most common version of the problem. Bot protection on a CDN, a WAF rule aimed at “AI scrapers”, a security plugin’s blocklist, a rate limiter tuned for human traffic, or a managed host running mitigation you cannot see from the admin panel. None of it is visible to a visitor, and none of it is in your robots.txt.

Should I be blocking AI crawlers at all?

That is your call, but make it per group rather than all at once. Blocking training crawlers keeps your content out of model training and costs you nothing in citations. Blocking search and browsing crawlers removes you from the answers themselves. A single User-agent: * rule does both, which is why most blocks we find were never intended.

If I don’t have a robots.txt file, am I blocking anything?

No. A crawler not mentioned in robots.txt is allowed. Having no file at all means everything is allowed — the risk of an absent robots.txt is that you have no way to be deliberate, not that you are blocking by default.

What’s the difference between robots.txt and noindex?

Robots.txt controls whether a crawler may fetch the page. Noindex tells it not to keep what it found. They fail in different places, so a page can answer every crawler with a healthy 200 and still be suppressed afterwards by a noindex nobody remembers adding — often injected at the CDN as a header rather than written into the page.

Do all AI systems respect robots.txt?

Most do. It is a convention, not a technical barrier — there is nothing stopping a crawler that ignores it. Treat robots.txt as the way you state your intent, and infrastructure rules as the way you enforce it.

If you tell me a crawler is blocked, is that proof?

It depends which kind. Blocked in robots.txt is stated as fact, because a compliant crawler obeys the file whatever a live fetch shows. A live refusal is different: we send from our own infrastructure, and security tools weigh the IP and request fingerprint as well as the user agent. We have seen one page, within seconds, return 403 to a browser, 200 to PerplexityBot and 520 to GPTBot. So we call that a risk worth treating as critical, not a confirmed block — your server logs are what confirm it.

The result changed and I didn’t change anything.

Live checks measure what your infrastructure did at that moment. A 429 or a 5xx can be a transient rate limit or a one-off WAF decision rather than a standing rule. Re-run to see whether it holds.

Will blocking AI crawlers hurt my traditional SEO?

Blocking the AI-specific crawlers does not affect Google’s ranking of your pages. The mistake that does hurt is blocking Googlebot while allowing Google-Extended: standard indexing has to work before Google’s AI features can include you at all.