Recrawled every night. Method published in full. Dataset open.
Most of the web never decided how AI may read it.
Someone decided anyway. We measure the most-visited sites on the internet every night and publish exactly what we find: which AI crawlers each one blocks, what it publishes for agents to read, whether it serves a crawler something different from what it serves you, and who actually made that call.
What we foundMeasure your siteBrowse the index
What the last crawl found
3,618 domains measured on 2026-09-22. Sites we could not observe are excluded rather than counted as failures, which is why these denominators are smaller than the corpus. See the whole funnel.
Block a crawler that answers questions
A crawler whose output reaches a person as an answer today: the ones behind ChatGPT, Claude, Perplexity, Gemini, Apple Intelligence and Meta AI. Blocking one has an immediate, visible cost. MoreSay one thing and do another
robots.txt permits GPTBot and the server refuses GPTBot anyway. The operator published one policy and a different one is being enforced, almost always by an edge rule switched on above them. MorePublish an llms.txt
A markdown file at the site root that points an AI agent at the pages worth reading. A community convention rather than a ratified standard, and adoption is still small. MoreMean readiness score
Zero to one hundred, from three bands: whether AI crawlers are allowed at all (45 points), whether the site publishes machine-readable surfaces (25), and whether its content is structured enough to be read (30). Arithmetic over archived evidence. No model is involved. MoreThe gap between what sites say and what they do
robots.txt is a published promise. What a server does when an AI crawler actually knocks is a separate fact. Every index in this category publishes the first one. We measure both on every domain, and 682 sites turn out to be enforcing a policy they never published.
| Group | Sites | Share |
|---|---|---|
| Says yes, does no | 682 | 18.9% |
| Open, and means it | 2479 | 68.5% |
| Closed, and means it | 128 | 3.5% |
| Says no, does yes | 194 | 5.4% |
Who is actually deciding
1,807 measured sites have a robots.txt that names no AI crawler at all, against 641 that name at least one. For most of the web, the AI policy is a side effect of a default somebody else shipped. Sites behind varnish block an answer-surface crawler 25.7% of the time against 3.7% behind azure-frontdoor, a spread far wider than anything the sites themselves publish explains.
Policy posture
Whether anyone actually decided. Deliberate means robots.txt names AI crawlers by token. Inherited means it names none, so whatever AI policy exists is a side effect of generic rules. Blanket means one rule for everyone. Absent means no robots.txt at all. More
Policy posture
Whether anyone actually decided. Deliberate means robots.txt names AI crawlers by token. Inherited means it names none, so whatever AI policy exists is a side effect of generic rules. Blanket means one rule for everyone. Absent means no robots.txt at all. More- Inherited1,807 (50%)
- Absent1,021 (28%)
- Deliberate641 (18%)
- Blanket149 (4%)
Blocking rate by edge network
The CDN or reverse proxy in front of the origin: Cloudflare, Akamai, Fastly, CloudFront. It can block a crawler before the site ever sees the request, which is why blocking correlates better with the CDN than with anything the operator published. More
edge network
The CDN or reverse proxy in front of the origin: Cloudflare, Akamai, Fastly, CloudFront. It can block a crawler before the site ever sees the request, which is why blocking correlates better with the CDN than with anything the operator published. More| Group | Value (%) |
|---|---|
| cloudflare | 10.1 |
| cloudfront | 21.1 |
| akamai | 16.0 |
| fastly | 23.9 |
| 3.8 | |
| vercel | 6.0 |
| varnish | 25.7 |
| azure-frontdoor | 3.7 |
All edge networks . By publishing platform . The full argument
Readiness by platform
What a site is built on predicts how legible it is to an agent, because the defaults come with the box.
| Platform | Sites | Blocking AI | Proportion blocking | Mean score |
|---|---|---|---|---|
| WordPress | 301 | 40 (13.3%) | 70.7 | |
| Next.js | 291 | 48 (16.5%) | 66 | |
| Adobe Experience Manager | 127 | 5 (3.9%) | 65.9 | |
| Drupal | 105 | 3 (2.9%) | 63.3 | |
| HubSpot CMS | 82 | 5 (6.1%) | 71.3 | |
| Contentful | 53 | 2 (3.8%) | 66.9 |
Which crawlers get shut out
Share of measured sites whose robots.txt denies each crawler the site root. The first 11 are answer-surface crawlers
A crawler whose output reaches a person as an answer today: the ones behind ChatGPT, Claude, Perplexity, Gemini, Apple Intelligence and Meta AI. Blocking one has an immediate, visible cost. More
| Crawler | Operator | Blocked by | Proportion |
|---|---|---|---|
| CCBot | Common Crawl | 489 (13.5%) | |
| Bytespider | ByteDance | 460 (12.7%) | |
| GPTBot | OpenAI | 457 (12.6%) | |
| ClaudeBot | Anthropic | 434 (12.0%) | |
| meta-externalagent | Meta | 397 (11.0%) | |
| Google-Extended | 384 (10.6%) | ||
| Diffbot | Diffbot | 371 (10.3%) | |
| cohere-ai | Cohere | 370 (10.2%) | |
| Applebot-Extended | Apple | 364 (10.1%) | |
| PerplexityBot | Perplexity | 354 (9.8%) | |
| Amazonbot | Amazon | 350 (9.7%) | |
| YouBot | You.com | 327 (9.0%) |
How the web scores
Every fully measured site, in ten-point bands, tinted by the grade each band falls under. The shape is the finding: 77% of the web sits between 50 and 79. Not hostile to agents, not ready for them either. Hover or tab through a band to see what is in it.
| Score band | Grade | Sites | Share |
|---|---|---|---|
| 0 to 9 | F | 0 | 0.0% |
| 10 to 19 | F | 3 | 0.1% |
| 20 to 29 | F | 18 | 0.7% |
| 30 to 39 | F | 33 | 1.3% |
| 40 to 49 | D | 251 | 9.8% |
| 50 to 59 | D | 755 | 29.4% |
| 60 to 69 | C | 775 | 30.2% |
| 70 to 79 | C and B | 451 | 17.6% |
| 80 to 89 | B | 266 | 10.4% |
| 90 to 100 | A | 15 | 0.6% |
Scores 0 to 9 . grade F
0 sites, 0.0% of the indexBlocks most answer-surface crawlers, or serves almost nothing a crawler can read.
No measured site currently scores in this band.
Scores 10 to 19 . grade F
3 sites, 0.1% of the indexBlocks most answer-surface crawlers, or serves almost nothing a crawler can read.
Scores 20 to 29 . grade F
18 sites, 0.7% of the indexBlocks most answer-surface crawlers, or serves almost nothing a crawler can read.
For examplelaunchpad.net 20usatoday.com 24amazon.ca 22
Scores 30 to 39 . grade F
33 sites, 1.3% of the indexBlocks most answer-surface crawlers, or serves almost nothing a crawler can read.
For exampleamazonvideo.com 37cnn.com 31bbc.com 34
Scores 40 to 49 . grade D
251 sites, 9.8% of the indexSubstantially closed or substantially unreadable. Often a platform default rather than a decision.
For examplegoogle.com 49yahoo.com 47nic.ru 45
Scores 50 to 59 . grade D
755 sites, 29.4% of the indexSubstantially closed or substantially unreadable. Often a platform default rather than a decision.
For exampleyoutube.com 56bing.com 52sharepoint.com 55
Scores 60 to 69 . grade C
775 sites, 30.2% of the indexReadable but undeclared. Typically no llms.txt, thin structured data, and a robots.txt that names no AI crawler.
Scores 70 to 79 . grades C and B
451 sites, 17.6% of the indexReadable but undeclared. Typically no llms.txt, thin structured data, and a robots.txt that names no AI crawler.
For exampleapple.com 77gandi.net 77zoom.us 77
Scores 80 to 89 . grade B
266 sites, 10.4% of the indexBroadly legible to agents with room left on the table. Usually one or two changes short of an A.
Scores 90 to 100 . grade A
15 sites, 0.6% of the indexOpen to the crawlers that answer questions, publishing machine-readable surfaces, and serving content an agent can read without running JavaScript.
For exampleprotothema.gr 91therapservices.net 90yotpo.com 93
Least agent-ready right now
Fully measured sites with the lowest scores, one row per operator. Partial assessments
Some checks could not be observed, usually because a bot wall answered instead of the site, so those points were removed from the total rather than failed. The remaining points are renormalised to one hundred. A partial score is not comparable with a complete one, which is why partial sites are kept out of ranked lists. More
| Rank | Domain | Score | Policy | Stack |
|---|---|---|---|---|
| 54 | tiktok.com | Score 15 out of 100, grade F | WalledDeliberate | unidentified |
| 849 | amazon.com.auand 4 more domains with the same policy | Score 15 out of 100, grade F | SelectiveDeliberate | Amazon CloudFront |
| 277 | launchpad.net | Score 20 out of 100, grade F | SelectiveDeliberate | unidentified |
| 586 | usatoday.com | Score 24 out of 100, grade F | WalledDeliberate | unidentified |
| 1287 | themoviedb.org | Score 24 out of 100, grade F | WalledDeliberate | Amazon CloudFront |
| 1404 | dw.com | Score 24 out of 100, grade F | SelectiveDeliberate | Akamai |
| 2635 | tmdb.org | Score 24 out of 100, grade F | WalledDeliberate | Amazon CloudFront |
| 3688 | politico.eu | Score 24 out of 100, grade F | WalledDeliberate | WordPress . Cloudflare |
| 692 | dailymail.co.ukand 1 more domain with the same policy | Score 26 out of 100, grade F | WalledDeliberate | Akamai |
| 1113 | theconversation.com | Score 26 out of 100, grade F | WalledDeliberate | Next.js . Fastly |
Recent movements
Only real changes are recorded. A site's first measurement is a baseline, not an event, and nothing is diffed across a change to our own probe.
- zscaler.comScore fell 6 points to 77.
- zscaler.comzscaler.com: llms.txt removed.
- zoho.euScore fell 3 points to 57.
- zoho.comScore fell 3 points to 57.
- zillow.comScore fell 28 points to 50.
- zaloapp.comWent unreachable: ETIMEDOUT.
The short answers
Every figure on this page with its denominator and its measurement date attached, so quoting one correctly takes no work. Free to reuse under CC BY 4.0 with credit.
- What percentage of websites block AI crawlers?
- 15.7% of measured sites block at least one AI crawler that answers questions today, 567 of 3,618 domains. Measured 2026-09-22 by Fidget Labs BV and published as CrawlIndex. 15.7%
- How many websites block every AI crawler?
- 4.1% of measured sites block every answer-surface AI crawler, 148 of 3,618 domains. Measured 2026-09-22 by Fidget Labs BV and published as CrawlIndex. 4.1%
- Do websites enforce the AI crawler policy they publish?
- Often not. 18.9% of measured sites, 682 of 3,618, permit GPTBot in robots.txt and then refuse a request from GPTBot at the server. The published policy and the enforced one disagree, usually because of an edge rule applied above the site operator. Measured 2026-09-22 by Fidget Labs BV and published as CrawlIndex. 18.9%
- How many websites publish an llms.txt file?
- 13.4% of measured sites publish an llms.txt, 484 of 3,618 domains. Adoption remains very small even across the most-visited sites on the web. Measured 2026-09-22 by Fidget Labs BV and published as CrawlIndex. 13.4%
- How many websites publish an agents.md file?
- 0.9% of measured sites publish an agents.md, 31 of 3,618 domains. Measured 2026-09-22 by Fidget Labs BV and published as CrawlIndex. 0.9%
- Do website owners actually choose their AI crawler policy?
- Mostly not. 78.2% of measured sites, 2,828 of 3,618, name no AI crawler in robots.txt at all, either because the file names none or because there is no file. Only 641 name one explicitly. For most of the web the AI policy arrived as a platform or CDN default rather than as a decision. Measured 2026-09-22 by Fidget Labs BV and published as CrawlIndex. 78.2%
- Are any websites charging AI crawlers for access?
- Yes, but very few. 10 of 3,618 measured sites answer an unpaid agent with HTTP 402 Payment Required, metering access rather than refusing it. Measured 2026-09-22 by Fidget Labs BV and published as CrawlIndex. 0.3%
- How ready is the average website for AI agents?
- The mean agent readiness score is 64.26 out of 100 across 3,618 measured domains. Measured 2026-09-22 by Fidget Labs BV and published as CrawlIndex. 64.26
Why this exists
Publishers are deciding, one robots.txt at a time, whether AI systems may read the web. Those decisions are made quietly, changed without announcement, and are individually trivial to check but collectively invisible.
CrawlIndex checks them on a schedule and keeps the receipts. The rubric is published, every score is arithmetic over archived evidence, and no language model touches the numbers. The whole dataset is downloadable. If you disagree with a result you can read exactly how it was reached and recompute it yourself.
Read the methodologyWhat every term meansDownload the datasetAdd a domain
crawl 2026-09-22 08:34 UTCprobe 3.0.0rubric 2.0.0registry 1.0.0vantage gha-ubuntu3,618 of 5,006 reachable
Using these figures
Free to reuse in research, journalism or a product under CC BY 4.0, with credit to Fidget Labs BV. Quote the measurement date so the claim stays checkable as the index moves.
CrawlIndex by Fidget Labs BV. "The state of AI crawler access." https://crawlindex.org (measured 2026-09-22). Licensed CC BY 4.0.