crawlindex

Tranco rank 2,370

congress.gov

congress.gov scores 25 out of 100 for agent readiness. It blocks 10 of 11 answer-surface AI crawlers.

Score 25 out of 100, grade F, partial assessment
Last measured
2026-08-09 18:01 UTC
First seen
2026-08-09
Rubric / probe
v1.0.0 / v2.0.0

Served through Cloudflare. Compare against everything else on the same stack.

Partial assessment

Our measurement request was met with a bot challenge (HTTP 403 on the control request). Anything we could otherwise read from the page describes that challenge rather than the site, so those checks were excluded and the score renormalised over what remained. The robots.txt findings are unaffected: that file is fetched separately and was served normally.

How this score is made up

Agent access3.4 / 38

  • Answer-surface crawlers allowed. 10 of 11 blocked: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Claude-SearchBot, PerplexityBot, Perplexity-User, Google-Extended, meta-externalagent.
  • Secondary crawlers allowed. 11 of 12 blocked: CCBot, Amazonbot, Bytespider, cohere-ai, MistralAI-User, YouBot, DuckAssistBot, Diffbot, Google-NotebookLM, TavilyBot, FirecrawlAgent.
  • Serves crawlers the same content. Not assessed. Our control request was challenged, so there is no clean baseline to compare against.

Machine-readable surface8 / 8

  • robots.txt published. robots.txt is present and parseable.
  • Sitemap declared in robots.txt. robots.txt points crawlers at a sitemap.
  • llms.txt published. Not assessed. /llms.txt could not be fetched cleanly past the bot challenge.
  • agents.md published. Not assessed. /agents.md could not be fetched cleanly past the bot challenge.

Content structurenot assessed

  • Organization schema. Not assessed. Our control request was challenged by a bot wall.
  • WebSite schema. Not assessed. Our control request was challenged by a bot wall.
  • Additional structured data. Not assessed. Our control request was challenged by a bot wall.
  • Readable without JavaScript. Not assessed. Our control request was challenged by a bot wall.
  • Single top-level heading. Not assessed. Our control request was challenged by a bot wall.
  • Semantic landmarks. Not assessed. Our control request was challenged by a bot wall.

Crawler policy

Read from congress.gov/robots.txt. A crawler is listed as blocked when the rules deny it the site root. This operator names 21 AI crawlers explicitly, which means the policy is deliberate rather than inherited from a wildcard rule.

AI crawler access policy for congress.gov
CrawlerOperatorTierNamedStatus
GPTBotOpenAI1yesBlocked
OAI-SearchBotOpenAI1yesBlocked
ChatGPT-UserOpenAI1yesBlocked
ClaudeBotAnthropic1yesBlocked
Claude-UserAnthropic1yesBlocked
Claude-SearchBotAnthropic1yesBlocked
PerplexityBotPerplexity1yesBlocked
Perplexity-UserPerplexity1yesBlocked
Google-ExtendedGoogle1yesBlocked
Applebot-ExtendedApple1noAllowed
meta-externalagentMeta1yesBlocked
CCBotCommon Crawl2yesBlocked
AmazonbotAmazon2yesBlocked
BytespiderByteDance2yesBlocked
cohere-aiCohere2yesBlocked
MistralAI-UserMistral2yesBlocked
YouBotYou.com2yesBlocked
DuckAssistBotDuckDuckGo2yesBlocked
kagi-fetcherKagi2noAllowed
DiffbotDiffbot2yesBlocked
Google-NotebookLMGoogle2yesBlocked
TavilyBotTavily2yesBlocked
FirecrawlAgentFirecrawl2yesBlocked

Show this score

Free to embed. Always reflects the latest measurement, and links back here so anyone can check the working.

CrawlIndex badge for congress.gov, score 25
<a href="https://crawlindex.org/site/congress.gov"><img src="https://crawlindex.org/badge/congress.gov.svg" alt="CrawlIndex agent readiness score for congress.gov" width="196" height="28"></a>

This measurement as JSON

Using these figures

Free to reuse in research, journalism or a product under CC BY 4.0, with credit to Fidget Labs BV. Quote the measurement date so the claim stays checkable as the index moves.

CrawlIndex by Fidget Labs BV. "congress.gov agent readiness." https://crawlindex.org (measured 2026-08-09). Licensed CC BY 4.0.