Tranco rank 1,239
gazzetta.it
gazzetta.it scores 54 out of 100 for agent readiness. It blocks 6 of 11 answer-surface AI crawlers.
Score 54 out of 100, grade D
- Last measured
- 2026-08-09 18:05 UTC
- First seen
- 2026-08-09
- Rubric / probe
- v1.0.0 / v2.0.0
How this score is made up
Agent access20.3 / 45
- Answer-surface crawlers allowed. 6 of 11 blocked: GPTBot, ClaudeBot, PerplexityBot, Google-Extended, Applebot-Extended, meta-externalagent.
- Secondary crawlers allowed. 2 of 12 blocked: CCBot, Amazonbot.
- Serves crawlers the same content. Requesting as GPTBot returned HTTP 403 and 425 bytes, against 457,581 bytes for a browser.
Machine-readable surface8 / 25
- robots.txt published. robots.txt is present and parseable.
- Sitemap declared in robots.txt. robots.txt points crawlers at a sitemap.
- llms.txt published. No /llms.txt.
- agents.md published. No /agents.md.
Content structure26 / 30
- Organization schema. Organization JSON-LD lets an agent resolve who publishes this site.
- WebSite schema. WebSite JSON-LD present.
- Additional structured data. No further JSON-LD types on the homepage.
- Readable without JavaScript. 26,867 characters of text in the server response. Most crawlers do not execute JavaScript.
- Single top-level heading. 1 h1 element found.
- Semantic landmarks. Uses main, nav, header, footer, aside.
Crawler policy
Read from gazzetta.it/robots.txt. A crawler is listed as blocked when the rules deny it the site root. This operator names 8 AI crawlers explicitly, which means the policy is deliberate rather than inherited from a wildcard rule.
| Crawler | Operator | Tier | Named | Status |
|---|---|---|---|---|
| GPTBot | OpenAI | 1 | yes | Blocked |
| OAI-SearchBot | OpenAI | 1 | no | Allowed |
| ChatGPT-User | OpenAI | 1 | no | Allowed |
| ClaudeBot | Anthropic | 1 | yes | Blocked |
| Claude-User | Anthropic | 1 | no | Allowed |
| Claude-SearchBot | Anthropic | 1 | no | Allowed |
| PerplexityBot | Perplexity | 1 | yes | Blocked |
| Perplexity-User | Perplexity | 1 | no | Allowed |
| Google-Extended | 1 | yes | Blocked | |
| Applebot-Extended | Apple | 1 | yes | Blocked |
| meta-externalagent | Meta | 1 | yes | Blocked |
| CCBot | Common Crawl | 2 | yes | Blocked |
| Amazonbot | Amazon | 2 | yes | Blocked |
| Bytespider | ByteDance | 2 | no | Allowed |
| cohere-ai | Cohere | 2 | no | Allowed |
| MistralAI-User | Mistral | 2 | no | Allowed |
| YouBot | You.com | 2 | no | Allowed |
| DuckAssistBot | DuckDuckGo | 2 | no | Allowed |
| kagi-fetcher | Kagi | 2 | no | Allowed |
| Diffbot | Diffbot | 2 | no | Allowed |
| Google-NotebookLM | 2 | no | Allowed | |
| TavilyBot | Tavily | 2 | no | Allowed |
| FirecrawlAgent | Firecrawl | 2 | no | Allowed |
Show this score
Free to embed. Always reflects the latest measurement, and links back here so anyone can check the working.
<a href="https://crawlindex.org/site/gazzetta.it"><img src="https://crawlindex.org/badge/gazzetta.it.svg" alt="CrawlIndex agent readiness score for gazzetta.it" width="196" height="28"></a>Using these figures
Free to reuse in research, journalism or a product under CC BY 4.0, with credit to Fidget Labs BV. Quote the measurement date so the claim stays checkable as the index moves.
CrawlIndex by Fidget Labs BV. "gazzetta.it agent readiness." https://crawlindex.org (measured 2026-08-09). Licensed CC BY 4.0.