Robots.txt Generator, Validator & Live Simulator
Generate compliant robots.txt files for Googlebot, Bingbot, and LLM crawlers. Test whether specific URL paths are allowed or blocked in real time using the official IETF RFC 9309 longest-match specificity standard.
AI Scraper & LLM Training Shield
Disallow AI models from harvesting copyrighted content for generative training while preserving search ranking.
User-Agent Directives
XML Sitemaps & Host
Specify index sitemaps for Googlebot and Bingbot to accelerate search discovery.
1# ===========================================================================2# robots.txt - Generated with ToolsAA SEO Robots.txt File Generator & Tester3# RFC 9309 Compliant | 100% Client-Side Privacy4# ===========================================================================56User-agent: *7Allow: /8Disallow: /admin/9Disallow: /private/10Disallow: /api/11Disallow: /*.json$1213# ---------------------------------------------------------------------------14# Sitemaps15# ---------------------------------------------------------------------------16Sitemap: https://example.com/sitemap.xml1718# Host directive (supported by select search engines like Yandex)19Host: example.com20
Robots.txt & RFC 9309 Specification Reference Guide
Standard rules, algorithm behavior, and best practices for modern webmasters and software developers.
01. Where to Place robots.txt
Robots.txt MUST be hosted at the exact root of your domain:
Subdirectories like /site/robots.txt are completely ignored by search engines. It must return HTTP status 200 with MIME type text/plain.
02. Specificity Precedence
Under the official RFC 9309 standard published by the IETF:
- Longest match wins: The rule with the greatest character length in the path pattern takes priority.
- Equal length tie-breaker: If an
AllowandDisallowdirective match the exact same number of characters, theAllowdirective wins.
03. AI Bots vs. Search Engines
Traditional bots (Googlebot, Bingbot) index your site for search visibility. AI training crawlers (GPTBot, ClaudeBot, Google-Extended, Common Crawl) scrape content to train generative models.
Blocking AI bots in robots.txt protects your copyrighted articles and code without hurting your regular Google search rankings.