Robots.txt Generator & Tester

RFC 9309
Search Engine OptimizationAI Scraper ShieldRFC 9309 Simulator

Robots.txt Generator, Validator & Live Simulator

Generate compliant robots.txt files for Googlebot, Bingbot, and LLM crawlers. Test whether specific URL paths are allowed or blocked in real time using the official IETF RFC 9309 longest-match specificity standard.

Bot Groups
1
Directives
5
Sitemaps
1
Size
672 B

AI Scraper & LLM Training Shield

Disallow AI models from harvesting copyrighted content for generative training while preserving search ranking.

GPTBotAllowed
OpenAI • Training
ChatGPT-UserAllowed
OpenAI • Search
Google-ExtendedAllowed
Google • Training
ClaudeBotAllowed
Anthropic • Training
anthropic-aiAllowed
Anthropic • Training
PerplexityBotAllowed
Perplexity • Search
CCBotAllowed
Common Crawl • Scraper
BytespiderAllowed
ByteDance • Training
Cohere-aiAllowed
Cohere • Training
FacebookBotAllowed
Meta • Training
Applebot-ExtendedAllowed
Apple • Training

User-Agent Directives

*
Quick Select:
Crawl-Delay:
Directives (5)
Quick Insert:

XML Sitemaps & Host

Specify index sitemaps for Googlebot and Bingbot to accelerate search discovery.

https://example.com/sitemap.xml
Main mirror host for search engines like Yandex. Googlebot ignores this directive.
robots.txt
1# ===========================================================================
2# robots.txt - Generated with ToolsAA SEO Robots.txt File Generator & Tester
3# RFC 9309 Compliant | 100% Client-Side Privacy
4# ===========================================================================
5
6User-agent: *
7Allow: /
8Disallow: /admin/
9Disallow: /private/
10Disallow: /api/
11Disallow: /*.json$
12
13# ---------------------------------------------------------------------------
14# Sitemaps
15# ---------------------------------------------------------------------------
16Sitemap: https://example.com/sitemap.xml
17
18# Host directive (supported by select search engines like Yandex)
19Host: example.com
20
RFC 9309 Compliant
20 lines • 672 B
Syntax & Quality Linter
100% Valid
No syntax issues or deprecations detected. Clean and ready for production.

Robots.txt & RFC 9309 Specification Reference Guide

Standard rules, algorithm behavior, and best practices for modern webmasters and software developers.

01. Where to Place robots.txt

Robots.txt MUST be hosted at the exact root of your domain:

https://yourdomain.com/robots.txt

Subdirectories like /site/robots.txt are completely ignored by search engines. It must return HTTP status 200 with MIME type text/plain.

02. Specificity Precedence

Under the official RFC 9309 standard published by the IETF:

  • Longest match wins: The rule with the greatest character length in the path pattern takes priority.
  • Equal length tie-breaker: If an Allow and Disallow directive match the exact same number of characters, the Allow directive wins.

03. AI Bots vs. Search Engines

Traditional bots (Googlebot, Bingbot) index your site for search visibility. AI training crawlers (GPTBot, ClaudeBot, Google-Extended, Common Crawl) scrape content to train generative models.

Blocking AI bots in robots.txt protects your copyrighted articles and code without hurting your regular Google search rankings.