Coverage
What this index does not cover
Every measurement project has a funnel between the domains it tries and the domains it can make a claim about. Most do not publish theirs. Ours loses roughly a third of the corpus before scoring, and you cannot judge any percentage on this site without knowing that.
The funnel
- Domains attempted on the last crawl5,006 (100%)
- Published records3,693 (74%)
- Reachable and observed3,653 (73%)
- Fully scored and comparable2,627 (52%)
Everything on the site that says "of measured sites" uses the third number as its denominator. Everything ranked uses the fourth. Neither uses the first, because counting a domain we could not reach as a failing site would be a claim we have no evidence for.
Why 40 domains could not be reached
A ranking of the most-visited domains contains a surprising amount of rubble: hosts that have moved, parked, expired, or never served a website in the first place. A domain failing three consecutive crawls is demoted out of the published population with the reason recorded.
| Reason | Domains | What it means |
|---|---|---|
| UND_ERR_CONNECT_TIMEOUT | 15 | The host accepted no connection before the timeout. |
| timeout | 9 | A connection was made but no response arrived in time. |
| ETIMEDOUT | 7 | The connection attempt itself timed out. |
| ECONNREFUSED | 2 | Something is at that address and it actively refused the connection. |
| homepage returned HTTP 502 | 1 | A transport failure reported by the runtime. |
| UND_ERR_SOCKET | 1 | A transport failure reported by the runtime. |
| homepage returned HTTP 504 | 1 | A transport failure reported by the runtime. |
| redirect count exceeded | 1 | A redirect loop, or a chain longer than any browser would follow. |
| UNABLE_TO_VERIFY_LEAF_SIGNATURE | 1 | The certificate chain could not be verified. |
| SELF_SIGNED_CERT_IN_CHAIN | 1 | A transport failure reported by the runtime. |
| UNABLE_TO_GET_ISSUER_CERT_LOCALLY | 1 | A transport failure reported by the runtime. |
Why 1,026 more are measured but not ranked
These sites answered, but something stopped us observing part of the rubric honestly. Those checks are removed from the total rather than failed, which leaves a score renormalised over fewer points, which is not comparable with a complete one. So they are published on their own pages and kept out of ranked lists.
- 693A bot wall answered instead of the site
- 155A stub with no readable content was served to our crawler
- 151robots.txt bars every crawler, so no page was fetched
- 17Measured before the current probe, so newer checks were never asked
- 10HTTP 402, access is metered rather than free
This is the rule that keeps the index honest and it has been wrong in both directions. Reddit once scored 3 because our crawler hit a proof-of-work wall, which would have been a lie about Reddit. Amazon scored 8 across twenty country domains because a 2KB anti-automation stub slipped past the challenge detector and every check derived from that stub was charged to Amazon. Both are now excluded rather than counted.
What happens when a crawl goes wrong
Every guard in this project used to protect against the crawl failing. None protected against the crawl succeeding in a changed world. If a network starts refusing our crawler, several hundred origins become unreachable, drop out of the denominator, and the surviving subset gets published as the headline with tonight's date on it. Every process exits zero and nothing says a word. Checking a vendor status page cannot catch that, because the infrastructure is fine and the measurement environment is not.
So each run is now compared against the last one that passed. Reachability, the size of the measured population, the size of the published population, the mean score, and the volume of change records all have to stay within a sane range of the previous night. A run that fails any of those is quarantined.
- The observations are still writtenThey are the evidence needed to work out what went wrong.
- No change records are recordedA change record is a claim about a named site, and a run we do not trust must not make claims about anybody.
- The day is excluded from monthly reportsA sealed report is never rewritten, so this is the one consequence a later re-crawl could not undo.
- The site keeps showing the last day that passedWith a banner naming the reasons, rather than quietly serving old numbers under a new date.
Recovery is a re-run. Nothing is deleted and nothing is hidden, and the quarantine clears the moment a clean crawl replaces the day.
There is one automatic way out, and it exists because of a mistake. Every check compares against the last day that passed, which is right for a fault that persists and wrong for a one-time step. Changing the rubric legitimately moved the mean score five points, and the gate then quarantined every night against a baseline that no longer existed. A gate that can never unstick itself is a gate somebody eventually switches off, which is worse than having none. So a run that reproduces the previous quarantined run within a tight tolerance is accepted, the baseline moves, and the site says so. A fault that is still changing does not agree with itself and keeps tripping, and a run that is broken on its own terms, having measured nothing or produced an impossible volume of change, is never accepted however consistently it repeats.
Limits that no amount of crawling fixes
- One vantage point
- Everything is measured from GitHub's runners in the US and EU. Origins genuinely serve differently by geography and IP reputation, so this index describes what an agent on that network sees. The vantage is recorded on every record and change detection refuses to compare across a change to it, but the geographic bias is real.
- One page per site
- The probe reads the homepage, robots.txt and three well-known paths. A site with a superb, well-structured article section and a thin homepage scores the thin homepage. Six requests per domain is what makes a nightly crawl of thousands of sites free and polite; fifteen would not be.
- Cloaking detection is conservative
- Only an outright refusal or a response under a quarter the size of the browser response counts. Dynamic pages vary legitimately, so treat an individual result as indicative and the aggregate as sound.
- Fingerprints drift
- Platforms and CDNs change their headers. A wrong label poisons a cross-tab, so the rule is never to guess: an unrecognised stack is null, not a guess, and cohorts under 25 sites are not published at all.
- The corpus is the popular web
- Tranco ranks the most-visited domains. That is the right population for a longitudinal index and the wrong one for a claim about the web as a whole, which is mostly small sites nobody ranks.
Something missing?
Add a domain and it is measured on the next crawl. Check one live without adding it. Or take the dataset and compute your own denominators.
crawl 2026-09-24 08:18 UTCprobe 3.0.0rubric 2.0.0registry 1.0.0vantage gha-ubuntu3,653 of 5,006 reachable
Using these figures
Free to reuse in research, journalism or a product under CC BY 4.0, with credit to Fidget Labs BV. Quote the measurement date so the claim stays checkable as the index moves.
CrawlIndex by Fidget Labs BV. "CrawlIndex coverage and limits." https://crawlindex.org. Licensed CC BY 4.0.