crawlindex

Tranco rank 154

nytimes.com

nytimes.com scores 38 out of 100 for agent readiness. It blocks 11 of 11 answer-surface AI crawlers.

Score 38 out of 100, grade F
Last measured
2026-08-09 18:14 UTC
First seen
2026-08-09
Rubric / probe
v1.0.0 / v2.0.0

Built on WordPress, served through Fastly. Compare against everything else on the same stack.

How this score is made up

Agent access4 / 45

  • Answer-surface crawlers allowed. 11 of 11 blocked: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Claude-SearchBot, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, meta-externalagent.
  • Secondary crawlers allowed. 6 of 12 blocked: CCBot, Bytespider, cohere-ai, YouBot, DuckAssistBot, Diffbot.
  • Serves crawlers the same content. Requesting as GPTBot returned HTTP 403 and 139,264 bytes, against 1,249,203 bytes for a browser.

Machine-readable surface8 / 25

  • robots.txt published. robots.txt is present and parseable.
  • Sitemap declared in robots.txt. robots.txt points crawlers at a sitemap.
  • llms.txt published. No /llms.txt.
  • agents.md published. No /agents.md.

Content structure26 / 30

  • Organization schema. Organization JSON-LD lets an agent resolve who publishes this site.
  • WebSite schema. WebSite JSON-LD present.
  • Additional structured data. No further JSON-LD types on the homepage.
  • Readable without JavaScript. 6,326 characters of text in the server response. Most crawlers do not execute JavaScript.
  • Single top-level heading. 1 h1 element found.
  • Semantic landmarks. Uses main, nav, header, footer, article.

Crawler policy

Read from nytimes.com/robots.txt. A crawler is listed as blocked when the rules deny it the site root. This operator names 18 AI crawlers explicitly, which means the policy is deliberate rather than inherited from a wildcard rule.

AI crawler access policy for nytimes.com
CrawlerOperatorTierNamedStatus
GPTBotOpenAI1yesBlocked
OAI-SearchBotOpenAI1yesBlocked
ChatGPT-UserOpenAI1yesBlocked
ClaudeBotAnthropic1yesBlocked
Claude-UserAnthropic1yesBlocked
Claude-SearchBotAnthropic1yesBlocked
PerplexityBotPerplexity1yesBlocked
Perplexity-UserPerplexity1yesBlocked
Google-ExtendedGoogle1yesBlocked
Applebot-ExtendedApple1yesBlocked
meta-externalagentMeta1yesBlocked
CCBotCommon Crawl2yesBlocked
AmazonbotAmazon2yesAllowed
BytespiderByteDance2yesBlocked
cohere-aiCohere2yesBlocked
MistralAI-UserMistral2noAllowed
YouBotYou.com2yesBlocked
DuckAssistBotDuckDuckGo2yesBlocked
kagi-fetcherKagi2noAllowed
DiffbotDiffbot2yesBlocked
Google-NotebookLMGoogle2noAllowed
TavilyBotTavily2noAllowed
FirecrawlAgentFirecrawl2noAllowed

Show this score

Free to embed. Always reflects the latest measurement, and links back here so anyone can check the working.

CrawlIndex badge for nytimes.com, score 38
<a href="https://crawlindex.org/site/nytimes.com"><img src="https://crawlindex.org/badge/nytimes.com.svg" alt="CrawlIndex agent readiness score for nytimes.com" width="196" height="28"></a>

This measurement as JSON

Using these figures

Free to reuse in research, journalism or a product under CC BY 4.0, with credit to Fidget Labs BV. Quote the measurement date so the claim stays checkable as the index moves.

CrawlIndex by Fidget Labs BV. "nytimes.com agent readiness." https://crawlindex.org (measured 2026-08-09). Licensed CC BY 4.0.