crawlindex

Search the index. A domain that is not here can still be measured live on the check page.

Findings

What the index actually shows

Four things stand out from 3,646 measured domains, and three of them are not what the category usually reports. Every figure here recomputes from the published dataset, so it moves when the web does.

Finding one

Thousands of sites enforce a policy they never published

robots.txt is a published promise. What a server does when a crawler carrying an AI user agent actually knocks is a separate fact, and the two do not have to agree. Every other index in this category publishes the first one. This index has measured both on every domain since the first crawl.

530 measured sites, or 14.5%, permit GPTBot in robots.txt and refuse a request from GPTBot at the server. Nothing in their published policy asked for that. It is almost always an edge rule switched on above the operator, and the operator usually does not know.

The measurement is deliberately narrow: only an outright refusal counts, never a response that merely looks thin, because a dynamic page varies legitimately and a false accusation here would be expensive.

Sites grouped by what robots.txt states against what the server does when asked as GPTBot
GroupSitesShare
Says yes, does no53014.5%
Open, and means it257270.5%
Closed, and means it2206.0%
Says no, does yes1905.2%

The most-visited sites where this is happening

Sites permitting GPTBot in robots.txt and refusing it at the server
RankDomainScoreAnswer-surface crawlersStack
35appsflyersdk.comrefused GPTBotScore 84 out of 100, grade B, partial assessmentall allowedAmazon CloudFront
64cloudflare.netrefused GPTBotScore 84 out of 100, grade B, partial assessmentall allowedCloudflare
108okcdn.rurefused GPTBotScore 84 out of 100, grade B, partial assessmentall allowedunidentified
115openai.comrefused GPTBotScore 60 out of 100, grade Call allowedContentful . Cloudflare
127nih.govrefused GPTBotScore 84 out of 100, grade B, partial assessmentall allowedCloudflare
155tumblr.comScore 46 out of 100, grade D4 blockedWordPress
156epicgames.comrefused GPTBotScore 84 out of 100, grade B, partial assessmentall allowedCloudflare
213sciencedirect.comrefused GPTBotScore 84 out of 100, grade B, partial assessmentall allowedCloudflare
219researchgate.netrefused GPTBotScore 91 out of 100, grade A, partial assessmentall allowedCloudflare
276springer.comScore 100 out of 100, grade A, partial assessmentall allowedFastly
311shalltry.comrefused GPTBotScore 84 out of 100, grade B, partial assessmentall allowedAmazon CloudFront
312wiley.comrefused GPTBotScore 63 out of 100, grade Call allowedAdobe Experience Manager . Cloudflare
318etsy.comrefused GPTBotScore 91 out of 100, grade A, partial assessmentall allowedFastly
323gnu.orgrefused GPTBotScore 58 out of 100, grade Dall allowedunidentified
332un.orgrefused GPTBotScore 53 out of 100, grade Dall allowedAmazon CloudFront
381cookiedatabase.orgrefused GPTBotScore 84 out of 100, grade B, partial assessmentall allowedCloudflare
383indeed.comrefused GPTBotScore 91 out of 100, grade A, partial assessmentall allowedCloudflare
404163.comrefused GPTBotScore 57 out of 100, grade Dall allowedAlibaba Cloud CDN
419amazonalexa.comrefused GPTBotScore 50 out of 100, grade Dall allowedAmazon CloudFront
427expireddomains.comrefused GPTBotScore 84 out of 100, grade B, partial assessmentall allowedCloudflare
439pixabay.comrefused GPTBotScore 84 out of 100, grade B, partial assessmentall allowedCloudflare
461ietf.orgrefused GPTBotScore 59 out of 100, grade Dall allowedCloudflare
467featureassets.orgrefused GPTBotScore 84 out of 100, grade B, partial assessmentall allowedGoogle Cloud / GFE
494stackoverflow.comrefused GPTBotScore 84 out of 100, grade B, partial assessmentall allowedCloudflare
509ft.comrefused GPTBotScore 58 out of 100, grade D, partial assessment6 blockedCloudflare

Finding two

Most of the web never made a decision about AI at all

Coverage of AI crawler blocking treats it as a choice publishers are making. For most of the web it is not a choice, it is a default.

1,833 measured sites have a robots.txt that names no AI crawler at all, and another 949 have no robots.txt whatsoever. Against that, only 709 name a single one. So for roughly 76.3% of the sites in this index, whatever AI policy exists arrived with the platform or the CDN.

Policy postureWhether anyone actually decided. Deliberate means robots.txt names AI crawlers by token. Inherited means it names none, so whatever AI policy exists is a side effect of generic rules. Blanket means one rule for everyone. Absent means no robots.txt at all. More

Share of measured sites in each policy posture
GroupValue (%)
Inherited50.3
Absent26.0
Deliberate19.4
Blanket4.3
Inherited
robots.txt exists and names no AI crawler at all. Whatever AI policy this site has is a side effect of generic rules it inherited, most often from its platform or CDN default.
Absent
No robots.txt. Crawler policy is undeclared, so every crawler applies its own default.
Deliberate
robots.txt names AI crawlers by token. Somebody at this organisation decided what AI systems may do with the site.
Blanket
One rule for every crawler, allow nothing. This is a decision, but it is not a decision about AI specifically.

Blocking rate by
edge networkThe CDN or reverse proxy in front of the origin: Cloudflare, Akamai, Fastly, CloudFront. It can block a crawler before the site ever sees the request, which is why blocking correlates better with the CDN than with anything the operator published. More

Share of sites behind each edge network blocking at least one answer-surface crawler
GroupValue (%)
cloudflare17.2
cloudfront19.7
akamai15.0
fastly24.8
google3.1
vercel6.3
varnish30.8
alibaba13.5
azure-frontdoor3.7
netlify4.0

The spread between the most and least restrictive edge networks is far wider than anything the sites themselves publish accounts for. Which CDN a site sits behind predicts its AI policy better than anything about the site.

Finding three

Blocking is not one behaviour, and the interesting group wants to be cited

A single "blocks AI" percentage flattens six different positions into one. Separating them turns up a segment that is usually invisible: sites that block the training crawlers and allow every crawler that answers a live question. They are not refusing AI. They are refusing to be absorbed while staying available to be quoted.

Open

56.2%

2,049 sites

Every answer-surface crawler is allowed. An agent asked about this site can read it.

Undeclared

26.0%

949 sites

No robots.txt, so nothing is stated either way.

No training

6.3%

229 sites

Training crawlers are blocked and the crawlers that answer live questions are not. This site wants to be cited without being absorbed.

Selective

5.9%

215 sites

Some answer-surface crawlers are blocked and others are not.

Walled

4.2%

154 sites

Every answer-surface crawler is blocked. This site is invisible to AI answers by choice.

Assistant only

1.1%

39 sites

Index builders are blocked but crawlers fetching a page on behalf of a person are allowed. Readable on request, not in bulk.

Metered

0.3%

11 sites

Access is sold rather than refused. An unpaid agent gets HTTP 402 Payment Required.

Measured across 11 answer-surface crawlers. Per-crawler detail.

Finding four

Almost nobody publishes anything for agents to read

There is a lot of writing about llms.txt, agent cards and machine-readable licensing. There is very little deployment. These are the adoption rates across the most-visited domains on the web, which is the most favourable population you could pick.

These signals were added in probe 3 and appear here after the next nightly crawl. They are shown as unmeasured rather than as zero, because we have not asked yet.

Declared authorship reaches 0.0%. An answer engine deciding whether to quote a page weighs when it was written and who wrote it, and both are cheaper to add than anything else on this list.

By stack

Cohorts under 25 measured sites are not published, because a 100% blocking rate over three sites is noise presented as a finding.

Edge network

Blocking rate by edge network
Edge networkSitesBlocking AIProportion blockingMean score
Cloudflare1,080186 (17.2%)74.7
Amazon CloudFront41682 (19.7%)70.1
Akamai26640 (15.0%)73.6
Fastly21854 (24.8%)69.6
Google Cloud / GFE1605 (3.1%)62.1
Vercel483 (6.3%)79.6
Varnish3912 (30.8%)64.4
Alibaba Cloud CDN375 (13.5%)67.2
Azure Front Door271 (3.7%)66.6
Netlify251 (4.0%)77.2

Publishing platform

Blocking rate by publishing platform
PlatformSitesBlocking AIProportion blockingMean score
WordPress30052 (17.3%)74.8
Next.js28244 (15.6%)71.6
Adobe Experience Manager1347 (5.2%)71.9
Drupal1173 (2.6%)69.5
HubSpot CMS813 (3.7%)77.6
Nuxt554 (7.3%)69.5
Contentful491 (2.0%)74.4
Sanity303 (10.0%)80.6

Top-level domain

Blocking rate by top-level domain
Top-level domainSitesBlocking AIProportion blockingMean score
.com1,957374 (19.1%)70.6
.org21926 (11.9%)70.2
.net17936 (20.1%)65.3
.ru14312 (8.4%)67.5
.edu835 (6.0%)70.3
.io746 (8.1%)72.5
.de7321 (28.8%)66.1
.gov683 (4.4%)71.3
.jp345 (14.7%)66.6
.fr3218 (56.3%)62.5

Check any of this yourself

Every number on this page is arithmetic over an archived observation, no model involved at any point, and the whole dataset is a download. How each one is measured . What the index does not cover . Download it

crawl 2026-08-10 04:42 UTCprobe 2.0.0rubric 1.0.0registry 1.0.0vantage gha-ubuntu3,646 of 5,006 reachable

Using these figures

Free to reuse in research, journalism or a product under CC BY 4.0, with credit to Fidget Labs BV. Quote the measurement date so the claim stays checkable as the index moves.

CrawlIndex by Fidget Labs BV. "What the CrawlIndex data shows." https://crawlindex.org (measured 2026-08-10). Licensed CC BY 4.0.