crawlindex

Search the index. A domain that is not here can still be measured live on the check page.

Coverage

What this index does not cover

Every measurement project has a funnel between the domains it tries and the domains it can make a claim about. Most do not publish theirs. Ours loses roughly a third of the corpus before scoring, and you cannot judge any percentage on this site without knowing that.

The funnel

Everything on the site that says "of measured sites" uses the third number as its denominator. Everything ranked uses the fourth. Neither uses the first, because counting a domain we could not reach as a failing site would be a claim we have no evidence for.

Why 40 domains could not be reached

A ranking of the most-visited domains contains a surprising amount of rubble: hosts that have moved, parked, expired, or never served a website in the first place. A domain failing three consecutive crawls is demoted out of the published population with the reason recorded.

Transport-level reasons a domain could not be measured
ReasonDomainsWhat it means
UND_ERR_CONNECT_TIMEOUT15The host accepted no connection before the timeout.
timeout9A connection was made but no response arrived in time.
ETIMEDOUT7The connection attempt itself timed out.
ECONNREFUSED2Something is at that address and it actively refused the connection.
homepage returned HTTP 5021A transport failure reported by the runtime.
UND_ERR_SOCKET1A transport failure reported by the runtime.
homepage returned HTTP 5041A transport failure reported by the runtime.
redirect count exceeded1A redirect loop, or a chain longer than any browser would follow.
UNABLE_TO_VERIFY_LEAF_SIGNATURE1The certificate chain could not be verified.
SELF_SIGNED_CERT_IN_CHAIN1A transport failure reported by the runtime.
UNABLE_TO_GET_ISSUER_CERT_LOCALLY1A transport failure reported by the runtime.

Why 1,026 more are measured but not ranked

These sites answered, but something stopped us observing part of the rubric honestly. Those checks are removed from the total rather than failed, which leaves a score renormalised over fewer points, which is not comparable with a complete one. So they are published on their own pages and kept out of ranked lists.

  • 693A bot wall answered instead of the site
  • 155A stub with no readable content was served to our crawler
  • 151robots.txt bars every crawler, so no page was fetched
  • 17Measured before the current probe, so newer checks were never asked
  • 10HTTP 402, access is metered rather than free

This is the rule that keeps the index honest and it has been wrong in both directions. Reddit once scored 3 because our crawler hit a proof-of-work wall, which would have been a lie about Reddit. Amazon scored 8 across twenty country domains because a 2KB anti-automation stub slipped past the challenge detector and every check derived from that stub was charged to Amazon. Both are now excluded rather than counted.

What happens when a crawl goes wrong

Every guard in this project used to protect against the crawl failing. None protected against the crawl succeeding in a changed world. If a network starts refusing our crawler, several hundred origins become unreachable, drop out of the denominator, and the surviving subset gets published as the headline with tonight's date on it. Every process exits zero and nothing says a word. Checking a vendor status page cannot catch that, because the infrastructure is fine and the measurement environment is not.

So each run is now compared against the last one that passed. Reachability, the size of the measured population, the size of the published population, the mean score, and the volume of change records all have to stay within a sane range of the previous night. A run that fails any of those is quarantined.

  • The observations are still writtenThey are the evidence needed to work out what went wrong.
  • No change records are recordedA change record is a claim about a named site, and a run we do not trust must not make claims about anybody.
  • The day is excluded from monthly reportsA sealed report is never rewritten, so this is the one consequence a later re-crawl could not undo.
  • The site keeps showing the last day that passedWith a banner naming the reasons, rather than quietly serving old numbers under a new date.

Recovery is a re-run. Nothing is deleted and nothing is hidden, and the quarantine clears the moment a clean crawl replaces the day.

There is one automatic way out, and it exists because of a mistake. Every check compares against the last day that passed, which is right for a fault that persists and wrong for a one-time step. Changing the rubric legitimately moved the mean score five points, and the gate then quarantined every night against a baseline that no longer existed. A gate that can never unstick itself is a gate somebody eventually switches off, which is worse than having none. So a run that reproduces the previous quarantined run within a tight tolerance is accepted, the baseline moves, and the site says so. A fault that is still changing does not agree with itself and keeps tripping, and a run that is broken on its own terms, having measured nothing or produced an impossible volume of change, is never accepted however consistently it repeats.

Limits that no amount of crawling fixes

One vantage point
Everything is measured from GitHub's runners in the US and EU. Origins genuinely serve differently by geography and IP reputation, so this index describes what an agent on that network sees. The vantage is recorded on every record and change detection refuses to compare across a change to it, but the geographic bias is real.
One page per site
The probe reads the homepage, robots.txt and three well-known paths. A site with a superb, well-structured article section and a thin homepage scores the thin homepage. Six requests per domain is what makes a nightly crawl of thousands of sites free and polite; fifteen would not be.
Cloaking detection is conservative
Only an outright refusal or a response under a quarter the size of the browser response counts. Dynamic pages vary legitimately, so treat an individual result as indicative and the aggregate as sound.
Fingerprints drift
Platforms and CDNs change their headers. A wrong label poisons a cross-tab, so the rule is never to guess: an unrecognised stack is null, not a guess, and cohorts under 25 sites are not published at all.
The corpus is the popular web
Tranco ranks the most-visited domains. That is the right population for a longitudinal index and the wrong one for a claim about the web as a whole, which is mostly small sites nobody ranks.

Something missing?

Add a domain and it is measured on the next crawl. Check one live without adding it. Or take the dataset and compute your own denominators.

crawl 2026-09-24 08:18 UTCprobe 3.0.0rubric 2.0.0registry 1.0.0vantage gha-ubuntu3,653 of 5,006 reachable

Using these figures

Free to reuse in research, journalism or a product under CC BY 4.0, with credit to Fidget Labs BV. Quote the measurement date so the claim stays checkable as the index moves.

CrawlIndex by Fidget Labs BV. "CrawlIndex coverage and limits." https://crawlindex.org. Licensed CC BY 4.0.