ReadableByAI
Menu ▾

THE READABLEBYAI INDEX

A quarter of YC Fall 2025 startups are blank pages to AI crawlers.

25.5% of 145 YC Fall 2025 homepages ship an empty client-rendered shell to GPTBot, ClaudeBot and PerplexityBot — no headline, no product description, nothing but a <div id="root">. Among 450 established SaaS companies measured the same way, it's 2%. That's a 12.8x gap. Measured 8 August 2026.

What was measured, and what wasn't

Every homepage in this Index was fetched twice: once with a baseline browser user-agent, and once each with the twelve crawler user-agents (eleven AI-vendor crawlers plus bingbot as a non-AI control) ReadableByAI probes, all from the same datacenter IP. The readability numbers below — visible word count, rendering classification, robots.txt contents — come from the baseline fetch, and they are identity-independent: whether a page ships its content in raw HTML or hides it behind a client-side render is true for every visitor, bot or browser, and needs no further confirmation.

What we deliberately did not publish is bot-specific access results — which user-agents got a 200, a 403, or a challenge page. A datacenter IP claiming to be ClaudeBot is an unverified probe, not the real crawler; vendors authenticate their bots by published IP range, so a challenge to our probe cannot distinguish a genuine block from a WAF simply doubting an impostor. Only the site's own server logs settle that, and we don't have them. So this Index reports what any fetch can prove — content and permission — and leaves access claims out entirely.

One further limit, stated plainly because it is the strongest objection to this method: a site can serve different HTML to different requesters. The common implementation keys on the user-agent string, and this Index detects it — every company named below returned the same content to every crawler identity tested, verified by comparing the response returned to each identity against the baseline (four were byte-for-byte identical across all thirteen fetches; the other two varied by under one percent in size, consistent with per-request timestamps). The two companies that did serve crawlers a materially different page are reported separately rather than counted as failures — they receive more, not less. What an outside probe cannot rule out is a site that varies its content by verified crawler IP range rather than by user-agent. That is rare, and nothing in this dataset suggests it, but it cannot be disproven from outside — which is the honest reason server logs matter and probes alone are never the last word.

That limit is not a flaw in this Index; it is the boundary of what any outside probe can establish, ours included. Resolving it needs the one record we do not have: the target's own server logs, with every hit checked against each vendor's published crawler IP ranges. That separates verified crawlers from impostors, shows which pages they actually retrieved, and surfaces the fetches they abandoned midway — none of which is visible from the outside, and all of which is what the paid audit reads. A roadmap item, not yet built and therefore not sold: dual-origin probing from residential and datacenter networks simultaneously, which would expose IP-sensitive bot management without logs.

The comparison

145 of 147 YC Fall 2025 companies and 455 of 514 established SaaS companies returned a clean 200 to the baseline fetch and are counted below. The rest are excluded from every percentage on this page, not silently dropped: 2 YC domains (2 non-200 responses) and 59 SaaS domains (51 non-200, 8 unreachable).

A further 5 SaaS domains are excluded from every percentage below because the baseline fetch returned a WAF challenge or interstitial rather than the site's own page — an empty <title> with a body under 2 KB, or a title naming a challenge page. Counting a challenge page as a content measurement is the error this rule exists to prevent; four of these rows were published as findings before 2026-08-27 and have been withdrawn. The same rule was applied to the YC population, where no domain matched. That leaves 450 SaaS and 145 YC homepages in the readability population.

PopulationClean baselineCSR_SHELLSSR_THINSSR_FULLMedian wordsrobots.txtllms.txt
YC Fall 2025145 / 14737 (25.5%)14 (9.7%)94 (64.8%)630111 (76.6%)48 (33.1%)
Established SaaS455 / 51411 (2%)11 (2.4%)428 (94.1%)1241441 (96.9%)245 (53.8%)

CSR_SHELL = under 150 visible words in raw HTML. SSR_THIN = 150–399. SSR_FULL = 400+. Median visible words: 630 for YC vs. 1241 for SaaS — the typical SaaS homepage ships roughly double the readable content of the typical YC Fall 2025 homepage, before either one is judged by whether it's reachable at all.

The comparison isn't one-directional. 4 companies in the established-SaaS set serve AI crawlers more content than they served our baseline fetch — bot-aware, dynamic rendering that detects a non-browser request and responds with fuller markup instead of a thinner one. The two largest gaps we measured: a payroll platform serving roughly 220x the byte size to a crawler that it serves to a plain browser fetch, and an ML-ops platform at roughly 141x (both public companies; unnamed here because per-company detail belongs to the named-examples policy above). Some teams have already solved this deliberately — it's evidence this is a solvable engineering problem, not an unavoidable one.

The 145 YC Fall 2025 homepages, anonymized

These are the 145 Y Combinator Fall 2025 companies that returned a clean baseline fetch, published anonymized. The distribution is the finding, not any single row in it: 25.5% of a whole recent YC batch under 150 visible words is the headline of this Index. Individual companies aren't named here, for the same reason the SaaS roster isn't published in full — the point of this Index is to get pages fixed, not to build a wall of names. If your company is in this batch, the free scan on the ReadableByAI homepage checks your own homepage the same way we checked this one, so you can find out where you stand without waiting on us to reach out.

The yc-001yc-145 ids are stable within this dataset — the same company keeps the same id release over release, so a re-verified row can be tracked across waves — but they carry no mapping we publish. There is no lookup table connecting an id back to a domain anywhere on this site.

Anonymized IDVisible wordsClassification
yc-0011CSR_SHELL
yc-0021CSR_SHELL
yc-0031CSR_SHELL
yc-0042CSR_SHELL
yc-0055CSR_SHELL
yc-0065CSR_SHELL
yc-0075CSR_SHELL
yc-0086CSR_SHELL
yc-0096CSR_SHELL
yc-0106CSR_SHELL
yc-0116CSR_SHELL
yc-0127CSR_SHELL
yc-0137CSR_SHELL
yc-0148CSR_SHELL
yc-0158CSR_SHELL
yc-0169CSR_SHELL
yc-0179CSR_SHELL
yc-0189CSR_SHELL
yc-0199CSR_SHELL
yc-02010CSR_SHELL
yc-02112CSR_SHELL
yc-02214CSR_SHELL
yc-02315CSR_SHELL
yc-02418CSR_SHELL
yc-02519CSR_SHELL
yc-02620CSR_SHELL
yc-02727CSR_SHELL
yc-02829CSR_SHELL
yc-02930CSR_SHELL
yc-03032CSR_SHELL
yc-03136CSR_SHELL
yc-03241CSR_SHELL
yc-03352CSR_SHELL
yc-03467CSR_SHELL
yc-03588CSR_SHELL
yc-03693CSR_SHELL
yc-037130CSR_SHELL
yc-038166SSR_THIN
yc-039167SSR_THIN
yc-040189SSR_THIN
yc-041212SSR_THIN
yc-042223SSR_THIN
yc-043304SSR_THIN
yc-044309SSR_THIN
yc-045314SSR_THIN
yc-046314SSR_THIN
yc-047369SSR_THIN
yc-048386SSR_THIN
yc-049388SSR_THIN
yc-050391SSR_THIN
yc-051395SSR_THIN
yc-052401SSR_FULL
yc-053440SSR_FULL
yc-054442SSR_FULL
yc-055454SSR_FULL
yc-056454SSR_FULL
yc-057515SSR_FULL
yc-058519SSR_FULL
yc-059522SSR_FULL
yc-060523SSR_FULL
yc-061534SSR_FULL
yc-062539SSR_FULL
yc-063540SSR_FULL
yc-064541SSR_FULL
yc-065554SSR_FULL
yc-066568SSR_FULL
yc-067574SSR_FULL
yc-068585SSR_FULL
yc-069596SSR_FULL
yc-070607SSR_FULL
yc-071612SSR_FULL
yc-072629SSR_FULL
yc-073630SSR_FULL
yc-074638SSR_FULL
yc-075644SSR_FULL
yc-076651SSR_FULL
yc-077651SSR_FULL
yc-078652SSR_FULL
yc-079655SSR_FULL
yc-080659SSR_FULL
yc-081690SSR_FULL
yc-082691SSR_FULL
yc-083706SSR_FULL
yc-084728SSR_FULL
yc-085729SSR_FULL
yc-086742SSR_FULL
yc-087748SSR_FULL
yc-088757SSR_FULL
yc-089775SSR_FULL
yc-090777SSR_FULL
yc-091786SSR_FULL
yc-092796SSR_FULL
yc-093835SSR_FULL
yc-094850SSR_FULL
yc-095876SSR_FULL
yc-096906SSR_FULL
yc-097908SSR_FULL
yc-098941SSR_FULL
yc-099943SSR_FULL
yc-100954SSR_FULL
yc-101962SSR_FULL
yc-102976SSR_FULL
yc-103981SSR_FULL
yc-1041003SSR_FULL
yc-1051049SSR_FULL
yc-1061065SSR_FULL
yc-1071069SSR_FULL
yc-1081096SSR_FULL
yc-1091101SSR_FULL
yc-1101119SSR_FULL
yc-1111154SSR_FULL
yc-1121160SSR_FULL
yc-1131169SSR_FULL
yc-1141187SSR_FULL
yc-1151188SSR_FULL
yc-1161213SSR_FULL
yc-1171216SSR_FULL
yc-1181279SSR_FULL
yc-1191292SSR_FULL
yc-1201323SSR_FULL
yc-1211368SSR_FULL
yc-1221377SSR_FULL
yc-1231420SSR_FULL
yc-1241423SSR_FULL
yc-1251509SSR_FULL
yc-1261522SSR_FULL
yc-1271533SSR_FULL
yc-1281538SSR_FULL
yc-1291566SSR_FULL
yc-1301589SSR_FULL
yc-1311615SSR_FULL
yc-1321635SSR_FULL
yc-1331641SSR_FULL
yc-1341642SSR_FULL
yc-1351656SSR_FULL
yc-1361689SSR_FULL
yc-1371748SSR_FULL
yc-1381756SSR_FULL
yc-1391812SSR_FULL
yc-1401943SSR_FULL
yc-1412198SSR_FULL
yc-1422299SSR_FULL
yc-1432895SSR_FULL
yc-1443868SSR_FULL
yc-1453952SSR_FULL

Six companies, named

This Index doesn't publish the roster of 661 domains it measured — see below for why. What it can publish, without that tradeoff, are results any reader can reproduce in ten seconds: six large, well-resourced, publicly recognizable companies, each checked exactly the way every other homepage in this dataset was checked.

These aren't obscure or under-resourced teams — that's the point. A company with Palantir's or Duolingo's engineering headcount isn't serving zero words to AI crawlers because it can't afford server-side rendering. It's serving zero words because a popular frontend framework's default configuration ships a client-rendered shell unless someone deliberately turns on server rendering, and nobody happened to. This is a framework-default problem, not a competence problem — and none of the six disallows any AI crawler in robots.txt, which is what you'd expect to see if this were policy instead of an accident. A page that's deliberately kept off-limits to GPTBot says so in robots.txt; a page that's just empty says nothing, because nobody meant for it to be empty.

CompanyVisible wordsClassification
duolingo.com1CSR_SHELL
palantir.com4CSR_SHELL
qualys.com8CSR_SHELL
grindr.com118CSR_SHELL
blackberry.com216SSR_THIN
substack.com284SSR_THIN

Reproduce any of these yourself — raw HTML, JavaScript disabled, visible word count only:

curl -sL https://duolingo.com | python3 -c "import re,sys; h=sys.stdin.read(); h=re.sub(r'<(script|style)(\s[^>]*)?>.*?</\1>','',h,flags=re.S|re.I); h=re.sub(r'<[^>]+>',' ',h); print(len(h.split()))"

What we're not publishing, and why

Earlier versions of this page listed all 600 clean-baseline domains. We took that table down. This dataset doubles as a private outreach list — companies we contact directly when we find their homepage invisible to AI crawlers — and publishing the full roster would trade a company's incentive to fix the problem for our incentive to publish a bigger list. We'd rather have the fix. Beyond the six SaaS companies named above: thirteen other companies in this dataset received the same or a closely related finding. We are notifying each of them privately rather than publishing a wall of names. That's not concealment — it's the actual policy, stated plainly: the point of this Index is to get pages fixed, not a leaderboard of who's been caught. If you run a company and want to know where you stand, the free scan on the ReadableByAI homepage checks your own site the same way, right now, without waiting for us to reach out.

Method and reproducibility

Every homepage was fetched with a plain HTTP GET — no headless browser, no JavaScript execution — because that's what GPTBot, ClaudeBot and PerplexityBot do. Visible word count is text content extracted from the raw HTML response, minus script and style contents. The three classification bands are fixed thresholds: CSR_SHELL under 150 visible words, SSR_THIN 150–399, SSR_FULL 400 or more. robots.txt is read directly and reported as a literal fact — present or absent, nothing inferred about intent. The engine behind this Index is open source at github.com/abouchard11/geo-crawl-audit. Anyone can clone it, point it at their own list, and check our numbers or produce their own.

Fixed it? We'll re-verify free.

If your homepage came back CSR_SHELL or SSR_THIN when we measured it — named above or notified privately — and you've since shipped server-rendered content, tell us and we'll re-probe it at no charge — no catch, no upsell attached. The goal of this Index is an accurate public record, not a leaderboard anyone stays stuck on. Tell us: alex+reverify@midnightdev.dev.

Named here? Talk to us directly.

If your company appears in this report — named in the table above or described anonymously — this is your direct line: alex+named@midnightdev.dev. A real person answers, same week. If we didn't have a working public contact channel for your company before publication, this address is that channel — no form, no gatekeeper. Disputes and reproduction mismatches follow the process on the Corrections page; re-verification after a fix is free, always.

Related reading

Four case files go deeper on the findings behind this Index, and the free scan on the ReadableByAI homepage checks your own site the same way.

Invisible without JavaScript

What a client-rendered shell looks like to a crawler that never runs the script — the raw HTML, side by side with what a browser paints.

Startups are 12.8x more invisible

The YC Fall 2025 vs. established-SaaS comparison behind this Index, in full — including the two findings that didn't survive review.

A menu on a locked door

Why publishing an llms.txt file does nothing if the homepage behind it is a blank shell — permission without content to permit.

“Why they block GPTBot” is usually wrong

The difference between a robots.txt disallow, a WAF challenge, and an empty page — and why only one of the three is in this dataset.

This Index is measured periodically, not once. The next wave will show what changed — which domains fixed a CSR_SHELL homepage, which didn't, and whether the gap between YC Fall 2025 and established SaaS narrowed or widened.