Findings
What the index actually shows
Four things stand out from 3,646 measured domains, and three of them are not what the category usually reports. Every figure here recomputes from the published dataset, so it moves when the web does.
Finding one
Thousands of sites enforce a policy they never published
robots.txt is a published promise. What a server does when a crawler carrying an AI user agent actually knocks is a separate fact, and the two do not have to agree. Every other index in this category publishes the first one. This index has measured both on every domain since the first crawl.
530 measured sites, or 14.5%, permit GPTBot in robots.txt and refuse a request from GPTBot at the server. Nothing in their published policy asked for that. It is almost always an edge rule switched on above the operator, and the operator usually does not know.
The measurement is deliberately narrow: only an outright refusal counts, never a response that merely looks thin, because a dynamic page varies legitimately and a false accusation here would be expensive.
| Group | Sites | Share |
|---|---|---|
| Says yes, does no | 530 | 14.5% |
| Open, and means it | 2572 | 70.5% |
| Closed, and means it | 220 | 6.0% |
| Says no, does yes | 190 | 5.2% |
The most-visited sites where this is happening
| Rank | Domain | Score | Answer-surface crawlers | Stack |
|---|---|---|---|---|
| 35 | appsflyersdk.comrefused GPTBot | Score 84 out of 100, grade B, partial assessment | all allowed | Amazon CloudFront |
| 64 | cloudflare.netrefused GPTBot | Score 84 out of 100, grade B, partial assessment | all allowed | Cloudflare |
| 108 | okcdn.rurefused GPTBot | Score 84 out of 100, grade B, partial assessment | all allowed | unidentified |
| 115 | openai.comrefused GPTBot | Score 60 out of 100, grade C | all allowed | Contentful . Cloudflare |
| 127 | nih.govrefused GPTBot | Score 84 out of 100, grade B, partial assessment | all allowed | Cloudflare |
| 155 | tumblr.com | Score 46 out of 100, grade D | 4 blocked | WordPress |
| 156 | epicgames.comrefused GPTBot | Score 84 out of 100, grade B, partial assessment | all allowed | Cloudflare |
| 213 | sciencedirect.comrefused GPTBot | Score 84 out of 100, grade B, partial assessment | all allowed | Cloudflare |
| 219 | researchgate.netrefused GPTBot | Score 91 out of 100, grade A, partial assessment | all allowed | Cloudflare |
| 276 | springer.com | Score 100 out of 100, grade A, partial assessment | all allowed | Fastly |
| 311 | shalltry.comrefused GPTBot | Score 84 out of 100, grade B, partial assessment | all allowed | Amazon CloudFront |
| 312 | wiley.comrefused GPTBot | Score 63 out of 100, grade C | all allowed | Adobe Experience Manager . Cloudflare |
| 318 | etsy.comrefused GPTBot | Score 91 out of 100, grade A, partial assessment | all allowed | Fastly |
| 323 | gnu.orgrefused GPTBot | Score 58 out of 100, grade D | all allowed | unidentified |
| 332 | un.orgrefused GPTBot | Score 53 out of 100, grade D | all allowed | Amazon CloudFront |
| 381 | cookiedatabase.orgrefused GPTBot | Score 84 out of 100, grade B, partial assessment | all allowed | Cloudflare |
| 383 | indeed.comrefused GPTBot | Score 91 out of 100, grade A, partial assessment | all allowed | Cloudflare |
| 404 | 163.comrefused GPTBot | Score 57 out of 100, grade D | all allowed | Alibaba Cloud CDN |
| 419 | amazonalexa.comrefused GPTBot | Score 50 out of 100, grade D | all allowed | Amazon CloudFront |
| 427 | expireddomains.comrefused GPTBot | Score 84 out of 100, grade B, partial assessment | all allowed | Cloudflare |
| 439 | pixabay.comrefused GPTBot | Score 84 out of 100, grade B, partial assessment | all allowed | Cloudflare |
| 461 | ietf.orgrefused GPTBot | Score 59 out of 100, grade D | all allowed | Cloudflare |
| 467 | featureassets.orgrefused GPTBot | Score 84 out of 100, grade B, partial assessment | all allowed | Google Cloud / GFE |
| 494 | stackoverflow.comrefused GPTBot | Score 84 out of 100, grade B, partial assessment | all allowed | Cloudflare |
| 509 | ft.comrefused GPTBot | Score 58 out of 100, grade D, partial assessment | 6 blocked | Cloudflare |
Finding two
Most of the web never made a decision about AI at all
Coverage of AI crawler blocking treats it as a choice publishers are making. For most of the web it is not a choice, it is a default.
1,833 measured sites have a robots.txt that names no AI crawler at all, and another 949 have no robots.txt whatsoever. Against that, only 709 name a single one. So for roughly 76.3% of the sites in this index, whatever AI policy exists arrived with the platform or the CDN.
Policy posture
Whether anyone actually decided. Deliberate means robots.txt names AI crawlers by token. Inherited means it names none, so whatever AI policy exists is a side effect of generic rules. Blanket means one rule for everyone. Absent means no robots.txt at all. More
Policy posture
Whether anyone actually decided. Deliberate means robots.txt names AI crawlers by token. Inherited means it names none, so whatever AI policy exists is a side effect of generic rules. Blanket means one rule for everyone. Absent means no robots.txt at all. More| Group | Value (%) |
|---|---|
| Inherited | 50.3 |
| Absent | 26.0 |
| Deliberate | 19.4 |
| Blanket | 4.3 |
- Inherited
- robots.txt exists and names no AI crawler at all. Whatever AI policy this site has is a side effect of generic rules it inherited, most often from its platform or CDN default.
- Absent
- No robots.txt. Crawler policy is undeclared, so every crawler applies its own default.
- Deliberate
- robots.txt names AI crawlers by token. Somebody at this organisation decided what AI systems may do with the site.
- Blanket
- One rule for every crawler, allow nothing. This is a decision, but it is not a decision about AI specifically.
Blocking rate by edge network
The CDN or reverse proxy in front of the origin: Cloudflare, Akamai, Fastly, CloudFront. It can block a crawler before the site ever sees the request, which is why blocking correlates better with the CDN than with anything the operator published. More
edge network
The CDN or reverse proxy in front of the origin: Cloudflare, Akamai, Fastly, CloudFront. It can block a crawler before the site ever sees the request, which is why blocking correlates better with the CDN than with anything the operator published. More| Group | Value (%) |
|---|---|
| cloudflare | 17.2 |
| cloudfront | 19.7 |
| akamai | 15.0 |
| fastly | 24.8 |
| 3.1 | |
| vercel | 6.3 |
| varnish | 30.8 |
| alibaba | 13.5 |
| azure-frontdoor | 3.7 |
| netlify | 4.0 |
The spread between the most and least restrictive edge networks is far wider than anything the sites themselves publish accounts for. Which CDN a site sits behind predicts its AI policy better than anything about the site.
Finding three
Blocking is not one behaviour, and the interesting group wants to be cited
A single "blocks AI" percentage flattens six different positions into one. Separating them turns up a segment that is usually invisible: sites that block the training crawlers and allow every crawler that answers a live question. They are not refusing AI. They are refusing to be absorbed while staying available to be quoted.
Open
56.2%2,049 sites
Every answer-surface crawler is allowed. An agent asked about this site can read it.
Undeclared
26.0%949 sites
No robots.txt, so nothing is stated either way.
No training
6.3%229 sites
Training crawlers are blocked and the crawlers that answer live questions are not. This site wants to be cited without being absorbed.
Selective
5.9%215 sites
Some answer-surface crawlers are blocked and others are not.
Walled
4.2%154 sites
Every answer-surface crawler is blocked. This site is invisible to AI answers by choice.
Assistant only
1.1%39 sites
Index builders are blocked but crawlers fetching a page on behalf of a person are allowed. Readable on request, not in bulk.
Metered
0.3%11 sites
Access is sold rather than refused. An unpaid agent gets HTTP 402 Payment Required.
Measured across 11 answer-surface crawlers. Per-crawler detail.
Finding four
Almost nobody publishes anything for agents to read
There is a lot of writing about llms.txt, agent cards and machine-readable licensing. There is very little deployment. These are the adoption rates across the most-visited domains on the web, which is the most favourable population you could pick.
These signals were added in probe 3 and appear here after the next nightly crawl. They are shown as unmeasured rather than as zero, because we have not asked yet.
Declared authorship reaches 0.0%. An answer engine deciding whether to quote a page weighs when it was written and who wrote it, and both are cheaper to add than anything else on this list.
By stack
Cohorts under 25 measured sites are not published, because a 100% blocking rate over three sites is noise presented as a finding.
Edge network
| Edge network | Sites | Blocking AI | Proportion blocking | Mean score |
|---|---|---|---|---|
| Cloudflare | 1,080 | 186 (17.2%) | 74.7 | |
| Amazon CloudFront | 416 | 82 (19.7%) | 70.1 | |
| Akamai | 266 | 40 (15.0%) | 73.6 | |
| Fastly | 218 | 54 (24.8%) | 69.6 | |
| Google Cloud / GFE | 160 | 5 (3.1%) | 62.1 | |
| Vercel | 48 | 3 (6.3%) | 79.6 | |
| Varnish | 39 | 12 (30.8%) | 64.4 | |
| Alibaba Cloud CDN | 37 | 5 (13.5%) | 67.2 | |
| Azure Front Door | 27 | 1 (3.7%) | 66.6 | |
| Netlify | 25 | 1 (4.0%) | 77.2 |
Publishing platform
| Platform | Sites | Blocking AI | Proportion blocking | Mean score |
|---|---|---|---|---|
| WordPress | 300 | 52 (17.3%) | 74.8 | |
| Next.js | 282 | 44 (15.6%) | 71.6 | |
| Adobe Experience Manager | 134 | 7 (5.2%) | 71.9 | |
| Drupal | 117 | 3 (2.6%) | 69.5 | |
| HubSpot CMS | 81 | 3 (3.7%) | 77.6 | |
| Nuxt | 55 | 4 (7.3%) | 69.5 | |
| Contentful | 49 | 1 (2.0%) | 74.4 | |
| Sanity | 30 | 3 (10.0%) | 80.6 |
Top-level domain
Check any of this yourself
Every number on this page is arithmetic over an archived observation, no model involved at any point, and the whole dataset is a download. How each one is measured . What the index does not cover . Download it
crawl 2026-08-10 04:42 UTCprobe 2.0.0rubric 1.0.0registry 1.0.0vantage gha-ubuntu3,646 of 5,006 reachable
Using these figures
Free to reuse in research, journalism or a product under CC BY 4.0, with credit to Fidget Labs BV. Quote the measurement date so the claim stays checkable as the index moves.
CrawlIndex by Fidget Labs BV. "What the CrawlIndex data shows." https://crawlindex.org (measured 2026-08-10). Licensed CC BY 4.0.