crawlindex

Search the index. A domain that is not here can still be measured live on the check page.

Tranco rank 5,059

geneanet.org

geneanet.org scores 78 out of 100 for agent readiness. It allows every answer-surface AI crawler.

Score 78 out of 100, grade B, partial assessment

Agent Friendly

Last measured
2026-08-17 03:38 UTC
First seen
2026-08-17
Rubric / probe
v2.0.0 / v3.0.0
VantageWhere the request came from. Origins serve differently by geography and IP reputation, so an observation is only comparable with another taken from the same place. Everything here is measured from GitHub runners in the US and EU. More
gha-ubuntu

Policy postureWhether anyone actually decided. Deliberate means robots.txt names AI crawlers by token. Inherited means it names none, so whatever AI policy exists is a side effect of generic rules. Blanket means one rule for everyone. Absent means no robots.txt at all. More

Absent

No robots.txt. Crawler policy is undeclared, so every crawler applies its own default.

Access archetypeThe shape of the policy rather than its size. Open, no training, assistant only, selective, walled, metered, or undeclared. More

Undeclared

No robots.txt, so nothing is stated either way.

Served through Cloudflare. Compare against everything else on the same stack.

This site's stated policy is not the one being enforcedrobots.txt permits GPTBot and the server refuses GPTBot anyway. The operator published one policy and a different one is being enforced, almost always by an edge rule switched on above them. More

robots.txt permits GPTBot, and a request identifying as GPTBot was refused with HTTP 403. Nothing in robots.txt asked for that, so it is almost certainly an edge rule applied above the operator rather than a decision they made. It is worth knowing about either way, because agents experience the enforcement and not the file.

Partial assessmentSome checks could not be observed, usually because a bot wall answered instead of the site, so those points were removed from the total rather than failed. The remaining points are renormalised to one hundred. A partial score is not comparable with a complete one, which is why partial sites are kept out of ranked lists. More

Our measurement request was met with a bot challenge (HTTP 403 on the control request). Anything we could otherwise read from the page describes that challenge rather than the site, so those checks were excluded and the score renormalised over what remained. The robots.txt findings are unaffected: that file is fetched separately and was served normally.

How this score is made up

Points earned in each score band
BandEarnedAvailableNominal maximum
Agent access383845
Machine-readable surface01125
Content structure0030

Agent access38 / 38

  • Answer-surface crawlers allowed. All 11 answer-surface crawlers are allowed.
  • Secondary crawlers allowed. All 12 secondary crawlers are allowed.
  • Serves crawlers the same content. Not assessed. There is no clean baseline to compare against, because our control request was challenged by a bot wall.

Machine-readable surface0 / 11

  • robots.txt published. No robots.txt. Crawler policy is undeclared.
  • Sitemap declared in robots.txt. robots.txt does not declare a sitemap.
  • llms.txt published. Not assessed. /llms.txt could not be read cleanly, because our control request was challenged by a bot wall.
  • agents.md published. Not assessed. /agents.md could not be read cleanly, because our control request was challenged by a bot wall.
  • Licence terms declared. No RSL License directive and no licence link relation. Reuse terms are undeclared.
  • Granular usage preferences declared. No Content-Signal directive. Policy is expressed only as allow or deny.
  • Agent card published. Not assessed. /.well-known/agent-card.json could not be read cleanly, because our control request was challenged by a bot wall.

Content structurenot assessed

  • Organization schema. Not assessed, because our control request was challenged by a bot wall.
  • WebSite schema. Not assessed, because our control request was challenged by a bot wall.
  • Additional structured data. Not assessed, because our control request was challenged by a bot wall.
  • Readable without JavaScript. Not assessed, because our control request was challenged by a bot wall.
  • Single top-level heading. Not assessed, because our control request was challenged by a bot wall.
  • Semantic landmarks. Not assessed, because our control request was challenged by a bot wall.
  • Dateline declared. Not assessed, because our control request was challenged by a bot wall.
  • Authorship declared. Not assessed, because our control request was challenged by a bot wall.

What would move this score

The observable checks that did not earn full marks, heaviest first. Only these; a check we could not observe is not on the list, because it is not a fact about the site.

  1. +4Sitemap declared in robots.txt. robots.txt does not declare a sitemap.
  2. +3robots.txt published. No robots.txt. Crawler policy is undeclared.
  3. +2Licence terms declared. No RSL License directive and no licence link relation. Reuse terms are undeclared.
  4. +2Granular usage preferences declared. No Content-Signal directive. Policy is expressed only as allow or deny.

Extraction profile

How cheap this page is for a retrieval pipeline to chunk. Recorded and published, deliberately not scored: it describes a shape rather than a pass or a fail, and compressing it into the grade would destroy the useful part.

Server text
16 chars
Text density
0%
Subheadings
0
Lists
0
Tables
0
Feed
no

Crawler policy

Read from geneanet.org/robots.txt. A crawler is listed as blocked when the rules deny it the site root. This operator names no AI crawler explicitly.

AI crawler access policy for geneanet.org
CrawlerOperatorTierNamedStatus
GPTBotOpenAI1noAllowed
OAI-SearchBotOpenAI1noAllowed
ChatGPT-UserOpenAI1noAllowed
ClaudeBotAnthropic1noAllowed
Claude-UserAnthropic1noAllowed
Claude-SearchBotAnthropic1noAllowed
PerplexityBotPerplexity1noAllowed
Perplexity-UserPerplexity1noAllowed
Google-ExtendedGoogle1noAllowed
Applebot-ExtendedApple1noAllowed
meta-externalagentMeta1noAllowed
CCBotCommon Crawl2noAllowed
AmazonbotAmazon2noAllowed
BytespiderByteDance2noAllowed
cohere-aiCohere2noAllowed
MistralAI-UserMistral2noAllowed
YouBotYou.com2noAllowed
DuckAssistBotDuckDuckGo2noAllowed
kagi-fetcherKagi2noAllowed
DiffbotDiffbot2noAllowed
Google-NotebookLMGoogle2noAllowed
TavilyBotTavily2noAllowed
FirecrawlAgentFirecrawl2noAllowed

geneanet.org has earned the Agent Friendly mark

Free to embed, no account and no fee. It regenerates from the nightly crawl, so it stays true, and it links back to this page so anyone can check the working in one click. What the mark means.

Shape

Circular. For a footer or an about page.

Theme

Follows the reader’s own light or dark setting. Note that this tracks the reader, not your page, so a dark-mode visitor sees the dark mark on a light site.

CrawlIndex agent readiness score for geneanet.orgOn a light page
On a dark page

The mark links back to this site's page for your domain, so anyone can check the claim in one click. It regenerates nightly, which means it stays true and it can change.

This measurement as JSON

crawl 2026-08-17 04:00 UTCprobe 3.0.0rubric 2.0.0registry 1.0.0vantage gha-ubuntu3,651 of 5,006 reachable

Using these figures

Free to reuse in research, journalism or a product under CC BY 4.0, with credit to Fidget Labs BV. Quote the measurement date so the claim stays checkable as the index moves.

CrawlIndex by Fidget Labs BV. "geneanet.org agent readiness." https://crawlindex.org (measured 2026-08-17). Licensed CC BY 4.0.