Coverage
What this index does not cover
Every measurement project has a funnel between the domains it tries and the domains it can make a claim about. Most do not publish theirs. Ours loses roughly a third of the corpus before scoring, and you cannot judge any percentage on this site without knowing that.
The funnel
- Domains attempted on the last crawl5,006 (100%)
- Published records3,719 (74%)
- Reachable and observed3,646 (73%)
- Fully scored and comparable2,684 (54%)
Everything on the site that says "of measured sites" uses the third number as its denominator. Everything ranked uses the fourth. Neither uses the first, because counting a domain we could not reach as a failing site would be a claim we have no evidence for.
Why 73 domains could not be reached
A ranking of the most-visited domains contains a surprising amount of rubble: hosts that have moved, parked, expired, or never served a website in the first place. A domain failing three consecutive crawls is demoted out of the published population with the reason recorded.
| Reason | Domains | What it means |
|---|---|---|
| timeout | 31 | A connection was made but no response arrived in time. |
| UND_ERR_CONNECT_TIMEOUT | 15 | The host accepted no connection before the timeout. |
| ENOTFOUND | 5 | DNS has no record. The domain is parked, expired, or was never a website. |
| redirect count exceeded | 5 | A redirect loop, or a chain longer than any browser would follow. |
| homepage returned HTTP 503 | 3 | A transport failure reported by the runtime. |
| homepage returned HTTP 405 | 3 | A transport failure reported by the runtime. |
| homepage returned HTTP 502 | 2 | A transport failure reported by the runtime. |
| homepage returned HTTP 404 | 2 | A transport failure reported by the runtime. |
| ECONNRESET | 2 | The connection was closed mid-response. |
| ETIMEDOUT | 1 | The connection attempt itself timed out. |
| homepage returned HTTP 511 | 1 | A transport failure reported by the runtime. |
| homepage returned HTTP 401 | 1 | A transport failure reported by the runtime. |
Why 962 more are measured but not ranked
These sites answered, but something stopped us observing part of the rubric honestly. Those checks are removed from the total rather than failed, which leaves a score renormalised over fewer points, which is not comparable with a complete one. So they are published on their own pages and kept out of ranked lists.
- 611A bot wall answered instead of the site
- 185Measured before the current probe, so newer checks were never asked
- 155robots.txt bars every crawler, so no page was fetched
- 11HTTP 402, access is metered rather than free
This is the rule that keeps the index honest and it has been wrong in both directions. Reddit once scored 3 because our crawler hit a proof-of-work wall, which would have been a lie about Reddit. Amazon scored 8 across twenty country domains because a 2KB anti-automation stub slipped past the challenge detector and every check derived from that stub was charged to Amazon. Both are now excluded rather than counted.
Limits that no amount of crawling fixes
- One vantage point
- Everything is measured from GitHub's runners in the US and EU. Origins genuinely serve differently by geography and IP reputation, so this index describes what an agent on that network sees. The vantage is recorded on every record and change detection refuses to compare across a change to it, but the geographic bias is real.
- One page per site
- The probe reads the homepage, robots.txt and three well-known paths. A site with a superb, well-structured article section and a thin homepage scores the thin homepage. Six requests per domain is what makes a nightly crawl of thousands of sites free and polite; fifteen would not be.
- Cloaking detection is conservative
- Only an outright refusal or a response under a quarter the size of the browser response counts. Dynamic pages vary legitimately, so treat an individual result as indicative and the aggregate as sound.
- Fingerprints drift
- Platforms and CDNs change their headers. A wrong label poisons a cross-tab, so the rule is never to guess: an unrecognised stack is null, not a guess, and cohorts under 25 sites are not published at all.
- The corpus is the popular web
- Tranco ranks the most-visited domains. That is the right population for a longitudinal index and the wrong one for a claim about the web as a whole, which is mostly small sites nobody ranks.
Something missing?
Add a domain and it is measured on the next crawl. Check one live without adding it. Or take the dataset and compute your own denominators.
crawl 2026-08-10 04:42 UTCprobe 2.0.0rubric 1.0.0registry 1.0.0vantage gha-ubuntu3,646 of 5,006 reachable
Using these figures
Free to reuse in research, journalism or a product under CC BY 4.0, with credit to Fidget Labs BV. Quote the measurement date so the claim stays checkable as the index moves.
CrawlIndex by Fidget Labs BV. "CrawlIndex coverage and limits." https://crawlindex.org. Licensed CC BY 4.0.