crawlindex

Search the index. A domain that is not here can still be measured live on the check page.

Coverage

What this index does not cover

Every measurement project has a funnel between the domains it tries and the domains it can make a claim about. Most do not publish theirs. Ours loses roughly a third of the corpus before scoring, and you cannot judge any percentage on this site without knowing that.

The funnel

Everything on the site that says "of measured sites" uses the third number as its denominator. Everything ranked uses the fourth. Neither uses the first, because counting a domain we could not reach as a failing site would be a claim we have no evidence for.

Why 73 domains could not be reached

A ranking of the most-visited domains contains a surprising amount of rubble: hosts that have moved, parked, expired, or never served a website in the first place. A domain failing three consecutive crawls is demoted out of the published population with the reason recorded.

Transport-level reasons a domain could not be measured
ReasonDomainsWhat it means
timeout31A connection was made but no response arrived in time.
UND_ERR_CONNECT_TIMEOUT15The host accepted no connection before the timeout.
ENOTFOUND5DNS has no record. The domain is parked, expired, or was never a website.
redirect count exceeded5A redirect loop, or a chain longer than any browser would follow.
homepage returned HTTP 5033A transport failure reported by the runtime.
homepage returned HTTP 4053A transport failure reported by the runtime.
homepage returned HTTP 5022A transport failure reported by the runtime.
homepage returned HTTP 4042A transport failure reported by the runtime.
ECONNRESET2The connection was closed mid-response.
ETIMEDOUT1The connection attempt itself timed out.
homepage returned HTTP 5111A transport failure reported by the runtime.
homepage returned HTTP 4011A transport failure reported by the runtime.

Why 962 more are measured but not ranked

These sites answered, but something stopped us observing part of the rubric honestly. Those checks are removed from the total rather than failed, which leaves a score renormalised over fewer points, which is not comparable with a complete one. So they are published on their own pages and kept out of ranked lists.

  • 611A bot wall answered instead of the site
  • 185Measured before the current probe, so newer checks were never asked
  • 155robots.txt bars every crawler, so no page was fetched
  • 11HTTP 402, access is metered rather than free

This is the rule that keeps the index honest and it has been wrong in both directions. Reddit once scored 3 because our crawler hit a proof-of-work wall, which would have been a lie about Reddit. Amazon scored 8 across twenty country domains because a 2KB anti-automation stub slipped past the challenge detector and every check derived from that stub was charged to Amazon. Both are now excluded rather than counted.

Limits that no amount of crawling fixes

One vantage point
Everything is measured from GitHub's runners in the US and EU. Origins genuinely serve differently by geography and IP reputation, so this index describes what an agent on that network sees. The vantage is recorded on every record and change detection refuses to compare across a change to it, but the geographic bias is real.
One page per site
The probe reads the homepage, robots.txt and three well-known paths. A site with a superb, well-structured article section and a thin homepage scores the thin homepage. Six requests per domain is what makes a nightly crawl of thousands of sites free and polite; fifteen would not be.
Cloaking detection is conservative
Only an outright refusal or a response under a quarter the size of the browser response counts. Dynamic pages vary legitimately, so treat an individual result as indicative and the aggregate as sound.
Fingerprints drift
Platforms and CDNs change their headers. A wrong label poisons a cross-tab, so the rule is never to guess: an unrecognised stack is null, not a guess, and cohorts under 25 sites are not published at all.
The corpus is the popular web
Tranco ranks the most-visited domains. That is the right population for a longitudinal index and the wrong one for a claim about the web as a whole, which is mostly small sites nobody ranks.

Something missing?

Add a domain and it is measured on the next crawl. Check one live without adding it. Or take the dataset and compute your own denominators.

crawl 2026-08-10 04:42 UTCprobe 2.0.0rubric 1.0.0registry 1.0.0vantage gha-ubuntu3,646 of 5,006 reachable

Using these figures

Free to reuse in research, journalism or a product under CC BY 4.0, with credit to Fidget Labs BV. Quote the measurement date so the claim stays checkable as the index moves.

CrawlIndex by Fidget Labs BV. "CrawlIndex coverage and limits." https://crawlindex.org. Licensed CC BY 4.0.