crawlindex

Search the index. A domain that is not here can still be measured live on the check page.

Tranco rank 2,781

xnhau.xxx

xnhau.xxx scores 30 out of 100 for agent readiness. It blocks 8 of 11 answer-surface AI crawlers. That is ahead of 1% of measured sites.

Score 30 out of 100, grade F

Measured

Number of measured sites in each ten-point score band
Score bandSites
0 to 90
10 to 193
20 to 2918
30 to 3933
40 to 49251
50 to 59755
60 to 69775
70 to 79451
80 to 89266
90 to 9915
Ahead of 1% of fully measured sites
Last measured
2026-09-22 08:32 UTC
First seen
2026-08-24
Rubric / probe
v2.0.0 / v3.0.0
VantageWhere the request came from. Origins serve differently by geography and IP reputation, so an observation is only comparable with another taken from the same place. Everything here is measured from GitHub runners in the US and EU. More
gha-ubuntu

Policy postureWhether anyone actually decided. Deliberate means robots.txt names AI crawlers by token. Inherited means it names none, so whatever AI policy exists is a side effect of generic rules. Blanket means one rule for everyone. Absent means no robots.txt at all. More

Deliberate

robots.txt names AI crawlers by token. Somebody at this organisation decided what AI systems may do with the site.

Access archetypeThe shape of the policy rather than its size. Open, no training, assistant only, selective, walled, metered, or undeclared. More

Selective

Some answer-surface crawlers are blocked and others are not.

Served through Cloudflare. Compare against everything else on the same stack.

How this score is made up

Points earned in each score band
BandEarnedAvailableNominal maximum
Agent access10.94545
Machine-readable surface92525
Content structure103030

Agent access10.9 / 45

  • Answer-surface crawlers allowed. 8 of 11 blocked: GPTBot, ChatGPT-User, ClaudeBot, Claude-User, Perplexity-User, Google-Extended, Applebot-Extended, meta-externalagent.
  • Secondary crawlers allowed. 8 of 12 blocked: CCBot, Amazonbot, Bytespider, cohere-ai, MistralAI-User, Diffbot, Google-NotebookLM, FirecrawlAgent.
  • Serves crawlers the same content. Requesting as GPTBot returned HTTP 403 and 25 bytes, against 219,417 bytes for a browser.

Machine-readable surface9 / 25

  • robots.txt published. robots.txt is present and parseable.
  • Sitemap declared in robots.txt. robots.txt points crawlers at a sitemap.
  • llms.txt published. No /llms.txt.
  • agents.md published. No /agents.md.
  • Licence terms declared. No RSL License directive and no licence link relation. Reuse terms are undeclared.
  • Granular usage preferences declared. robots.txt carries Content-Signal: search=yes,ai-train=no,use=reference. Preferences are stated per use rather than as a single allow or deny.
  • Agent card published. No agent card at /.well-known/agent-card.json.

Content structure10 / 30

  • Organization schema. No Organization JSON-LD. Agents cannot reliably attribute this site to an entity.
  • WebSite schema. No WebSite JSON-LD.
  • Additional structured data. No further JSON-LD types on the homepage.
  • Readable without JavaScript. 8,730 characters of text in the server response. Most crawlers do not execute JavaScript.
  • Single top-level heading. 1 h1 element found.
  • Semantic landmarks. Uses nav.
  • Dateline declared. No machine-readable date. An agent cannot tell how current this page is.
  • Authorship declared. No declared author. An agent quoting this page has nobody to credit.

How this compares to its peers

A score out of 100 is not information on its own. Against the sites running the same platform and sitting behind the same edge network, it is: this site is behind at least one of its own cohorts, which usually means the gap is its own to close. Cohorts under 25 measured sites are never published, so every comparison here is against a real group.

This site's score against the median of each group it belongs to
GroupMedian scoreSites in groupDifference
This site3010
Sites behind Cloudflare641132-34.0
The whole index612567-31.0

What would move this score

The observable checks that did not earn full marks, heaviest first. Only these; a check we could not observe is not on the list, because it is not a fact about the site.

  1. +21.8Answer-surface crawlers allowed. 8 of 11 blocked: GPTBot, ChatGPT-User, ClaudeBot, Claude-User, Perplexity-User, Google-Extended, Applebot-Extended, meta-externalagent.
  2. +9llms.txt published. No /llms.txt.
  3. +7Serves crawlers the same content. Requesting as GPTBot returned HTTP 403 and 25 bytes, against 219,417 bytes for a browser.
  4. +7Organization schema. No Organization JSON-LD. Agents cannot reliably attribute this site to an entity.
  5. +5.3Secondary crawlers allowed. 8 of 12 blocked: CCBot, Amazonbot, Bytespider, cohere-ai, MistralAI-User, Diffbot, Google-NotebookLM, FirecrawlAgent.
  6. +4agents.md published. No /agents.md.
  7. +3WebSite schema. No WebSite JSON-LD.
  8. +3Additional structured data. No further JSON-LD types on the homepage.

Extraction profile

How cheap this page is for a retrieval pipeline to chunk. Recorded and published, deliberately not scored: it describes a shape rather than a pass or a fail, and compressing it into the grade would destroy the useful part.

Server text
8,730 chars
Text density
4%
Subheadings
6
Lists
18
Tables
0
Feed
yes

Crawler policy

Read from xnhau.xxx/robots.txt. A crawler is listed as blocked when the rules deny it the site root. This operator names 16 AI crawlers explicitly, which means the policy is deliberate rather than inherited from a wildcard rule.

AI crawler access policy for xnhau.xxx
CrawlerOperatorTierNamedStatus
GPTBotOpenAI1yesBlocked
OAI-SearchBotOpenAI1noAllowed
ChatGPT-UserOpenAI1yesBlocked
ClaudeBotAnthropic1yesBlocked
Claude-UserAnthropic1yesBlocked
Claude-SearchBotAnthropic1noAllowed
PerplexityBotPerplexity1noAllowed
Perplexity-UserPerplexity1yesBlocked
Google-ExtendedGoogle1yesBlocked
Applebot-ExtendedApple1yesBlocked
meta-externalagentMeta1yesBlocked
CCBotCommon Crawl2yesBlocked
AmazonbotAmazon2yesBlocked
BytespiderByteDance2yesBlocked
cohere-aiCohere2yesBlocked
MistralAI-UserMistral2yesBlocked
YouBotYou.com2noAllowed
DuckAssistBotDuckDuckGo2noAllowed
kagi-fetcherKagi2noAllowed
DiffbotDiffbot2yesBlocked
Google-NotebookLMGoogle2yesBlocked
TavilyBotTavily2noAllowed
FirecrawlAgentFirecrawl2yesBlocked

Recorded changes

No mark at this score

The embeddable mark starts at grade B, because offering a graphic nobody would put on their own site is a pretence rather than a feature. The list above is the useful version: for most sites the points are concentrated in two or three changes. How the mark works.

CrawlIndex mark for xnhau.xxx, score 30

The neutral mark above exists and is free to use if you want to show that the site is independently measured whatever the number says.

This measurement as JSON

crawl 2026-09-22 08:34 UTCprobe 3.0.0rubric 2.0.0registry 1.0.0vantage gha-ubuntu3,618 of 5,006 reachable

Using these figures

Free to reuse in research, journalism or a product under CC BY 4.0, with credit to Fidget Labs BV. Quote the measurement date so the claim stays checkable as the index moves.

CrawlIndex by Fidget Labs BV. "xnhau.xxx agent readiness." https://crawlindex.org (measured 2026-09-22). Licensed CC BY 4.0.