crawlindex

Search the index. A domain that is not here can still be measured live on the check page.

Recrawled every night. Method published in full. Dataset open.

Most of the web never decided how AI may read it.

Someone decided anyway. We measure the most-visited sites on the internet every night and publish exactly what we find: which AI crawlers each one blocks, what it publishes for agents to read, whether it serves a crawler something different from what it serves you, and who actually made that call.

What we foundMeasure your siteBrowse the index

16 in every 100 measured sites block at least one crawler that answers questions today567 of 3,618 domains, measured 2026-09-22

What the last crawl found

3,618 domains measured on 2026-09-22. Sites we could not observe are excluded rather than counted as failures, which is why these denominators are smaller than the corpus. See the whole funnel.

16%
Block a crawler that answers questionsA crawler whose output reaches a person as an answer today: the ones behind ChatGPT, Claude, Perplexity, Gemini, Apple Intelligence and Meta AI. Blocking one has an immediate, visible cost. More
567 of 3,618 sites
19%
Say one thing and do anotherrobots.txt permits GPTBot and the server refuses GPTBot anyway. The operator published one policy and a different one is being enforced, almost always by an edge rule switched on above them. More
682 permit GPTBot in robots.txt and refuse it at the server
13%
Publish an llms.txtA markdown file at the site root that points an AI agent at the pages worth reading. A community convention rather than a ratified standard, and adoption is still small. More
484 sites, out of 3,618
64
Mean readiness scoreZero to one hundred, from three bands: whether AI crawlers are allowed at all (45 points), whether the site publishes machine-readable surfaces (25), and whether its content is structured enough to be read (30). Arithmetic over archived evidence. No model is involved. More
Out of 100, across fully scored sites

The gap between what sites say and what they do

robots.txt is a published promise. What a server does when an AI crawler actually knocks is a separate fact. Every index in this category publishes the first one. We measure both on every domain, and 682 sites turn out to be enforcing a policy they never published.

Sites grouped by what robots.txt states against what the server does when asked as GPTBot
GroupSitesShare
Says yes, does no68218.9%
Open, and means it247968.5%
Closed, and means it1283.5%
Says no, does yes1945.4%

Which sites, and why it happens

Who is actually deciding

1,807 measured sites have a robots.txt that names no AI crawler at all, against 641 that name at least one. For most of the web, the AI policy is a side effect of a default somebody else shipped. Sites behind varnish block an answer-surface crawler 25.7% of the time against 3.7% behind azure-frontdoor, a spread far wider than anything the sites themselves publish explains.

Policy postureWhether anyone actually decided. Deliberate means robots.txt names AI crawlers by token. Inherited means it names none, so whatever AI policy exists is a side effect of generic rules. Blanket means one rule for everyone. Absent means no robots.txt at all. More

  • Inherited1,807 (50%)
  • Absent1,021 (28%)
  • Deliberate641 (18%)
  • Blanket149 (4%)

Blocking rate by
edge networkThe CDN or reverse proxy in front of the origin: Cloudflare, Akamai, Fastly, CloudFront. It can block a crawler before the site ever sees the request, which is why blocking correlates better with the CDN than with anything the operator published. More

Share of sites behind each edge network that block at least one answer-surface crawler
GroupValue (%)
cloudflare10.1
cloudfront21.1
akamai16.0
fastly23.9
google3.8
vercel6.0
varnish25.7
azure-frontdoor3.7

All edge networks . By publishing platform . The full argument

Readiness by platform

What a site is built on predicts how legible it is to an agent, because the defaults come with the box.

AI blocking rate by publishing platform
PlatformSitesBlocking AIProportion blockingMean score
WordPress30140 (13.3%)70.7
Next.js29148 (16.5%)66
Adobe Experience Manager1275 (3.9%)65.9
Drupal1053 (2.9%)63.3
HubSpot CMS825 (6.1%)71.3
Contentful532 (3.8%)66.9

Which crawlers get shut out

Share of measured sites whose robots.txt denies each crawler the site root. The first 11 are

answer-surface crawlersA crawler whose output reaches a person as an answer today: the ones behind ChatGPT, Claude, Perplexity, Gemini, Apple Intelligence and Meta AI. Blocking one has an immediate, visible cost. More
, whose output reaches a reader today.

AI crawlers ranked by how many indexed sites block them
CrawlerOperatorBlocked byProportion
CCBotCommon Crawl489 (13.5%)
BytespiderByteDance460 (12.7%)
GPTBotOpenAI457 (12.6%)
ClaudeBotAnthropic434 (12.0%)
meta-externalagentMeta397 (11.0%)
Google-ExtendedGoogle384 (10.6%)
DiffbotDiffbot371 (10.3%)
cohere-aiCohere370 (10.2%)
Applebot-ExtendedApple364 (10.1%)
PerplexityBotPerplexity354 (9.8%)
AmazonbotAmazon350 (9.7%)
YouBotYou.com327 (9.0%)

All 23 tracked crawlers

How the web scores

Every fully measured site, in ten-point bands, tinted by the grade each band falls under. The shape is the finding: 77% of the web sits between 50 and 79. Not hostile to agents, not ready for them either. Hover or tab through a band to see what is in it.

Number of measured sites in each ten-point score band
Score bandGradeSitesShare
0 to 9F00.0%
10 to 19F30.1%
20 to 29F180.7%
30 to 39F331.3%
40 to 49D2519.8%
50 to 59D75529.4%
60 to 69C77530.2%
70 to 79C and B45117.6%
80 to 89B26610.4%
90 to 100A150.6%

Scores 60 to 69 . grade C

775 sites, 30.2% of the index

Readable but undeclared. Typically no llms.txt, thin structured data, and a robots.txt that names no AI crawler.

For examplemicrosoft.com 62github.com 68wikipedia.org 66

All 775 sites scoring 60 to 69

Find out where your site sits . How the score is built

Least agent-ready right now

Fully measured sites with the lowest scores, one row per operator.

Partial assessmentsSome checks could not be observed, usually because a bot wall answered instead of the site, so those points were removed from the total rather than failed. The remaining points are renormalised to one hundred. A partial score is not comparable with a complete one, which is why partial sites are kept out of ranked lists. More
are excluded, because a renormalised score is not comparable with a complete one.

Lowest scoring operators in the index
RankDomainScorePolicyStack
54tiktok.comScore 15 out of 100, grade FWalledDeliberateunidentified
849amazon.com.au
and 4 more domains with the same policy
Score 15 out of 100, grade FSelectiveDeliberateAmazon CloudFront
277launchpad.netScore 20 out of 100, grade FSelectiveDeliberateunidentified
586usatoday.comScore 24 out of 100, grade FWalledDeliberateunidentified
1287themoviedb.orgScore 24 out of 100, grade FWalledDeliberateAmazon CloudFront
1404dw.comScore 24 out of 100, grade FSelectiveDeliberateAkamai
2635tmdb.orgScore 24 out of 100, grade FWalledDeliberateAmazon CloudFront
3688politico.euScore 24 out of 100, grade FWalledDeliberateWordPress . Cloudflare
692dailymail.co.uk
and 1 more domain with the same policy
Score 26 out of 100, grade FWalledDeliberateAkamai
1113theconversation.comScore 26 out of 100, grade FWalledDeliberateNext.js . Fastly

Full leaderboard, best and worst

Recent movements

Only real changes are recorded. A site's first measurement is a baseline, not an event, and nothing is diffed across a change to our own probe.

All recorded changes

The short answers

Every figure on this page with its denominator and its measurement date attached, so quoting one correctly takes no work. Free to reuse under CC BY 4.0 with credit.

What percentage of websites block AI crawlers?
15.7% of measured sites block at least one AI crawler that answers questions today, 567 of 3,618 domains. Measured 2026-09-22 by Fidget Labs BV and published as CrawlIndex. 15.7%
How many websites block every AI crawler?
4.1% of measured sites block every answer-surface AI crawler, 148 of 3,618 domains. Measured 2026-09-22 by Fidget Labs BV and published as CrawlIndex. 4.1%
Do websites enforce the AI crawler policy they publish?
Often not. 18.9% of measured sites, 682 of 3,618, permit GPTBot in robots.txt and then refuse a request from GPTBot at the server. The published policy and the enforced one disagree, usually because of an edge rule applied above the site operator. Measured 2026-09-22 by Fidget Labs BV and published as CrawlIndex. 18.9%
How many websites publish an llms.txt file?
13.4% of measured sites publish an llms.txt, 484 of 3,618 domains. Adoption remains very small even across the most-visited sites on the web. Measured 2026-09-22 by Fidget Labs BV and published as CrawlIndex. 13.4%
How many websites publish an agents.md file?
0.9% of measured sites publish an agents.md, 31 of 3,618 domains. Measured 2026-09-22 by Fidget Labs BV and published as CrawlIndex. 0.9%
Do website owners actually choose their AI crawler policy?
Mostly not. 78.2% of measured sites, 2,828 of 3,618, name no AI crawler in robots.txt at all, either because the file names none or because there is no file. Only 641 name one explicitly. For most of the web the AI policy arrived as a platform or CDN default rather than as a decision. Measured 2026-09-22 by Fidget Labs BV and published as CrawlIndex. 78.2%
Are any websites charging AI crawlers for access?
Yes, but very few. 10 of 3,618 measured sites answer an unpaid agent with HTTP 402 Payment Required, metering access rather than refusing it. Measured 2026-09-22 by Fidget Labs BV and published as CrawlIndex. 0.3%
How ready is the average website for AI agents?
The mean agent readiness score is 64.26 out of 100 across 3,618 measured domains. Measured 2026-09-22 by Fidget Labs BV and published as CrawlIndex. 64.26

Why this exists

Publishers are deciding, one robots.txt at a time, whether AI systems may read the web. Those decisions are made quietly, changed without announcement, and are individually trivial to check but collectively invisible.

CrawlIndex checks them on a schedule and keeps the receipts. The rubric is published, every score is arithmetic over archived evidence, and no language model touches the numbers. The whole dataset is downloadable. If you disagree with a result you can read exactly how it was reached and recompute it yourself.

Read the methodologyWhat every term meansDownload the datasetAdd a domain

crawl 2026-09-22 08:34 UTCprobe 3.0.0rubric 2.0.0registry 1.0.0vantage gha-ubuntu3,618 of 5,006 reachable

Using these figures

Free to reuse in research, journalism or a product under CC BY 4.0, with credit to Fidget Labs BV. Quote the measurement date so the claim stays checkable as the index moves.

CrawlIndex by Fidget Labs BV. "The state of AI crawler access." https://crawlindex.org (measured 2026-09-22). Licensed CC BY 4.0.