AI CRAWLER READINESS · MEASURED, NOT CLAIMED
Can AI see your site?
Check reachability, speed, robots.txt and the content available without JavaScript. The free scan compares crawler user-agents with a browser baseline from our own network. Find a possible access problem, get a developer work order, and recheck after the fix. The result is an access diagnostic, not a prediction of AI citations.
By scanning you confirm you own or are authorized to test this domain. ~15 polite requests, cached 24h — ship a fix and re-check anytime from the results below.
Free, no account. Want a weekly recheck and an evidence report when it detects a change? Monitor is $49/mo; Verified — your own server logs — is $149/mo.

This scan is the same method behind the ReadableByAI Index: 661 companies — Y Combinator's Fall 2025 batch and established SaaS firms — measured the same way. A quarter of the startups serve raw HTML that's effectively blank to a non-rendering crawler. Read the findings →
What the scan checks
Four things have to be true before any AI system can cite you, in this order. Almost everything the generative-engine-optimization industry sells — schema markup, entity coverage, citation-friendly phrasing — sits downstream of all four. If you fail gate one, nothing downstream of it runs.
Reachability
Does the crawler get a 200?Bot-management rules ship with defaults, and defaults do not know which crawlers you want. A 403 or a challenge page to a retrieval crawler means you are absent from answers that cite your competitors. This gate fails silently — a crawler that never got a page leaves no trace in your analytics.
Speed
How slow is a cold request?Crawlers land on long-tail URLs your cache has never seen, so the cold-start number is the one that applies to them, not the warm number your monitoring reports. Slow origins get abandoned mid-fetch, which your log records as a 499 and nothing else notices.
Readability
How many words survive with JavaScript off?GPTBot, ClaudeBot and PerplexityBot do not run JavaScript. For most sites, Googlebot (feeding Gemini) and Applebot are the only AI-adjacent crawlers that render it. A client-rendered site can rank first in Google and be a blank page to nearly everything else.
Permission
What does robots.txt actually say?Including the tokens that never fetch. Google-Extended and Applebot-Extended are opt-out signals, not crawlers — they will never appear in your logs because they never make a request. Googlebot and Applebot do the fetching; the token governs the training use.
CONTINUOUS · FREE · CUSTOMER-OWNED
See real AI crawlers in your own PostHog
The outside scan above cannot prove that a request came from a real vendor network. Free hosted monitoring can. It filters your Vercel log drain, verifies known crawler IPs when vendors publish ranges, removes IPs, queries, headers, and raw user-agents, then writes the operational event to a PostHog project you control.
The free service also contributes a smaller domain-level record to ReadableByAI under versioned Data Contribution Terms. Those minimized records build the crawler-intelligence benchmark that pays for the free infrastructure. Original log batches and visitor-level data are not sold.
WEEKLY · $49/MONTH · EVIDENCE, NOT A DASHBOARD
Keep the evidence when client-site access changes
Monitoring checks your homepage every week from a datacenter address — 12 crawler user-agents against a browser baseline, plus robots.txt — and emails the evidence. You get an alert only when a crawler that was reachable becomes blocked on a clean baseline; otherwise a short weekly digest. One domain, $49 a month, no login, no seat licenses, cancel anytime. First run within 24 hours.
Synthetic probes are leads, not proof of vendor identity — pair it with the free log-based monitoring above when you need the real crawlers. Subscribe on the pricing page, or run the free scan above and the button appears in your results.
VERIFIED · $149/MONTH · YOUR OWN LOGS, NOT OUR PROBE
Stop guessing. Read what actually fetched your site.
Everything in Monitor, plus your server's own records: which AI crawlers fetched what, how many arrived from the vendor's published IP ranges versus impostors wearing the user-agent, which requests returned 4xx/5xx or took over a second, and the paths they hit most. One Vercel drain, about three minutes to connect; credentials arrive in your first report. IPs and user-agents are never retained, and paid log data never enters our public benchmark.
Evidence is labelled verified (your logs) or simulated (our probes) — never blended. This is the tier that settles an argument a probe can only start. Monitor and Verified are separate subscriptions rather than an upgrade path; if you are already on Monitor, email us to move instead of starting a second one. How the hookup works · Compare the tiers
Requires Vercel Drains on Pro/Enterprise or Cloudflare HTTP Logpush on Enterprise, with permission to configure the connection. Provider charges may apply. Check compatibility before subscribing. Neither plan measures AI-answer citations or guarantees visibility. The free scan and post-fix recheck remain free.
The 12 crawlers we probe
Each is sent as a separate request and compared against a baseline browser request to the same URL, seconds apart. A difference between the two is the finding — that is what distinguishes bot-specific filtering from a site that is simply down or slow.
| User-agent | Type | Operator and purpose |
|---|---|---|
| GPTBot | training | OpenAI — model training corpus |
| OAI-SearchBot | retrieval | OpenAI — index behind ChatGPT search |
| ChatGPT-User | user fetch | OpenAI — live fetch when a user shares a link |
| ClaudeBot | training | Anthropic — model training corpus |
| Claude-SearchBot | retrieval | Anthropic — search index |
| Claude-User | user fetch | Anthropic — live fetch on request |
| PerplexityBot | retrieval | Perplexity — answer index |
| Perplexity-User | user fetch | Perplexity — live fetch on request |
| bingbot | retrieval | Microsoft — feeds Copilot as well as Bing |
| Amazonbot | retrieval | Amazon — Alexa and shopping surfaces |
| CCBot | training | Common Crawl — upstream of most open datasets |
| meta-externalagent | training | Meta — Llama training corpus |
Google-Extended and Applebot-Extended are checked in robots.txt but never probed — they are permission tokens that never issue a request.
What this found across 18 major sites
Measured 7 August 2026. The full dataset is committed in the open-source repo, including the caveats and the sites where the probe could not reach a conclusion.
The Guardian: 200 to GPTBot, 403 to ClaudeBot
The Guardian served GPTBot, OAI-SearchBot and ChatGPT-User a clean 200 — and returned 403 to ClaudeBot, PerplexityBot and CCBot. The Guardian has a content deal with OpenAI; the status codes mirror the contract. The New York Times, in litigation with OpenAI, returned 403 to nearly the entire field — only bingbot and Amazonbot got through. Licensing deals are visible in HTTP status codes: the firewall is the policy document, whether or not anyone wrote it down.
Reddit disallows 14 AI tokens — 5 crawlers got a 200 anyway
Reddit's robots.txt disallows fourteen AI tokens with no exceptions. In practice, GPTBot got a 403 and ClaudeBot and CCBot got 429s — while OAI-SearchBot, Claude-SearchBot, PerplexityBot, Amazonbot and meta-externalagent all received a 200 from the same IP seconds apart. Figma runs the mirror image: seven tokens disallowed in robots.txt, a 200 returned to every one of them. robots.txt is a request; the WAF is the enforcement — and the two frequently disagree.
Airbnb: 94 words. LinkedIn: 23. Reddit: 1. Stripe: 1,957.
That is how many words of raw HTML each homepage serves before JavaScript runs — which is all that GPTBot, ClaudeBot and PerplexityBot ever see. Airbnb's homepage carries 94 words; LinkedIn served the probe a 23-word bot-check interstitial; Reddit's homepage contains one word. For contrast, Stripe served 1,957 words and a 200 to every user-agent in the test, and Anthropic answered fastest at 0.102 seconds warm. Famous sites are blank to AI — some deliberately, some without knowing it.
The full write-up, including two findings that did not survive review of the data, is at Your firewall is your AI policy.
Human investigation when automation cannot settle it — from $2,500
The free scan and hosted monitor should settle most access incidents without a person. The Verified AI Search Audit is the escalation path when the evidence remains inconclusive or implementation needs hands-on work: your ten most valuable URLs probed the way machines fetch them, verified against your real crawler logs, joined with Search Console demand, and classified across seven stages — access, indexing, intent fit, authority, outcome, variance, and business impact. You get the root cause with evidence, paste-ready developer tickets, and a re-check after you ship the fixes. Written by Alex Bouchard, forwardable to your team or client. Single domain from $2,500; multi-domain and portfolio engagements from $10,000. Run the free scan above — the request form is in your results.
What the free scan can tell you, and what it can't
It can tell you what's on your page. Whether your words are actually there for a machine to read, what your robots file permits, and how quickly your site answers. Those answers are the same no matter who asks, so you can rely on them.
It can't tell you who's really getting in. To test your door, we knock while saying we're ChatGPT's crawler — but we aren't really it. If your security turns us away, that might mean it turns the real one away too, or it might mean it spotted an impostor and let the genuine one through. From outside, those look identical. So we treat it as a question, never an answer.
Your server records settle access. They show which AI crawlers genuinely visited, which pages they took, and which ones gave up waiting before your site replied. Turn that evidence on continuously with free customer-owned monitoring, or analyze the logs yourself. The paid audit is reserved for ambiguous cases and implementation work.
Read the technical version
The scan fetches your page from a single datacenter address, declaring each crawler's identity in the request header. Content findings — visible word count in raw HTML, rendering classification, robots.txt directives, time to first byte — are identity-independent: they do not vary with who issues the request, so they need no further confirmation.
Access findings are different. We are an unverified requester asserting a crawler identity; the genuine crawler arrives from the vendor's published IP ranges, and bot-management products authenticate on exactly that basis. A challenge or 403 to our probe is therefore ambiguous between a real block and correct impostor handling, and we report every such result as a lead requiring log confirmation — never as a finding.
Server-log analysis resolves it: each hit verified against vendor IP ranges, separating authenticated crawlers from spoofed traffic, with per-path retrieval counts and 499 abandonment rates that never appear in JavaScript analytics. The open-source engine performs this analysis (Mode B); the audit supplies the interpretation, the Search Console join, and the fixes.
Roadmap, not yet built and therefore not sold: dual-origin probing — the same request issued from residential and datacenter networks simultaneously, exposing IP-sensitive bot management without requiring logs.
Questions
Why doesn't my analytics show this?
Because analytics runs in JavaScript, and the crawlers that matter here do not execute it. A bot that receives a 403, a challenge page, or an empty shell never fires your tracking. The failure is invisible from inside the browser — it is only visible from the request side, which is what this scan reads.
Isn't this just SEO?
It overlaps at the reachability layer and diverges at rendering. Googlebot renders JavaScript, so a client-rendered page can rank perfectly well in Google while being empty to GPTBot and ClaudeBot. Passing a Lighthouse or Search Console check tells you nothing about whether an AI crawler could read the page.
Does llms.txt help?
Almost certainly not, and this scan weights it at zero. Measured adoption data shows the overwhelming majority of llms.txt files receive no AI-crawler requests at all. It costs nothing to publish and it is not a substitute for being fetchable. We check for it and report it as information, not as a score component.
Is a spoofed user-agent the same as the real crawler?
No, and this is the honest limit of any active probe. We send requests from our own IP with a claimed identity in the header. Vendors verify their crawlers by published IP range, so a site can treat a real GPTBot differently from a request calling itself GPTBot. When our own baseline browser request is filtered, we withhold the score instead of reporting a number we cannot stand behind. Server-log analysis is the only way to settle what real crawlers received.
What is a 499 and why does it matter?
It is the status your log records when the client hung up before your origin answered. For AI crawlers it means the fetch was abandoned — no error page, no alert, no citation, and nothing in your uptime monitoring. It is the most common silent failure on slow origins and it cannot be detected by probing from outside.
What does the Verified AI Search Audit include, and what does it cost?
Ten commercially important URLs probed the way machines actually fetch them, verification against your server or CDN logs where available, a Search Console join showing which valuable pages are affected, repeated answer-engine sampling reported with run counts rather than one-run claims, root-cause classification across seven stages, paste-ready fixes for your developer, and one post-fix re-verification within 30 days. Every finding is labeled verified, simulated, or sampled — never blended. A single-domain audit starts at $2,500; multi-domain and portfolio engagements start at $10,000 and are scoped to the number of properties and whether server logs are available.