Submit
Add a domain to the index
The nightly crawl covers 3,719 domains, drawn from the most-visited sites on the web. Anything outside that can be added here, and it costs nothing.
Opens a prefilled issue on GitHub. It asks for one thing: the hostname.
What happens next
Nobody reviews it. There is no queue and no approval step, which is why this is worth doing rather than being another form that goes nowhere.
- 1You file the issueOne field. GitHub handles the account and the spam filtering, so this site needs neither.
- 2Tonight, the crawler reads itAt 02:30 UTC the nightly run picks up every open submission, validates the domain and adds the valid ones to the corpus.
- 3It gets measured on the same runSame probe, same rubric, same treatment as every other domain. No preferential handling for a site that asked to be here.
- 4The issue closes with a linkEither to the new page, or with the specific reason the domain was refused. Either way you get an answer within a day and it is a public one.
What gets refused
Only two things, and neither is a judgement about the site.
- Anything that is not a valid hostname. No schemes, no paths, no ports.
- Infrastructure rather than a site. CDN hosts, analytics endpoints, ad exchanges, nameservers and deeply nested subdomains. They rank highly because everything embeds them, they have no homepage anyone reads, and scoring them would distort every aggregate on the index. The rule is a published pattern list, not a decision somebody makes case by case.
Notably not on that list: whether the site scores well, whether you own it, or whether the operator would like to be measured. Everything here is read from files those servers already hand to every crawler on the internet.
If you want a domain out
Disallow CrawlIndexBot in the site's robots.txt. It is checked before any other request is made, so opting out costs the server exactly one request, and the domain leaves the published index on the next crawl. No message to anyone, no waiting on a person, and no account.
A blanket User-agent: * / Disallow: / is honoured by fetching no pages, but it is not treated as an opt-out. Deleting the most restrictive operators on the web from an index about restrictiveness would make the index useless, so only naming the token counts.
Measure a domain right now instead . See what the index does and does not cover