Robots.txt Generator with AI Crawler Rules & Tester

Build a robots.txt from Allow and Disallow lists for all robots or for specific crawlers, block AI bots in one click, add your sitemap, and test any URL against the result.

Loading tool…

About the Robots.txt Generator

robots.txt is a plain text file at the root of a host (https://example.com/robots.txt) that tells crawlers which paths they may fetch, standardised as RFC 9309. It is made of groups: one or more User-agent lines followed by Allow and Disallow rules. A crawler obeys only the most specific group that names it and falls back to User-agent: * otherwise — groups are not combined, so a group for Googlebot must repeat every rule that should apply to it.

Paths are matched from the start of the URL path, case-sensitively. * stands for any sequence of characters and $ anchors the end, so /*.pdf$ blocks every PDF. When several rules match, the longest one wins, and Allow wins a tie; the order of the lines does not matter. The tester applies exactly this logic to the generated file. Crawl-delay is not part of the standard: Bing and Yandex follow it, Google ignores it.

The file is a request, not access control. Well-behaved crawlers — including the AI crawlers GPTBot, ClaudeBot, Google-Extended, PerplexityBot and CCBot — respect it, but scrapers can ignore it, and the file itself is public, so it should not list secret paths. Disallow also does not remove a page from search results: a blocked URL can still be indexed from links to it. To keep a page out of the index, allow crawling and send a noindex meta tag or X-Robots-Tag header instead.

How to use it

  1. Pick a preset or choose the default policy for all robots.
  2. List the paths to disallow and the exceptions to allow, one per line.
  3. Add a group for specific crawlers, such as AI bots, with their own rules.
  4. Enter the URL of your sitemap.
  5. Test a few URLs, then download robots.txt and upload it to the root of your site.

Frequently asked questions

How do I block AI crawlers like GPTBot and ClaudeBot?
Click “Block AI crawlers”: it adds a group with the user-agents of the known AI crawlers and Disallow: /. Google-Extended and Applebot-Extended only opt out of AI training; they do not affect Google Search or Siri.
Does Disallow remove a page from Google?
No. It stops crawling, not indexing: the URL can still appear in results without a description. Use a noindex meta tag or X-Robots-Tag header, and let the page be crawled so that the tag is seen.
Where do I put robots.txt?
At the root of each host and scheme: https://example.com/robots.txt. A file in a subdirectory is ignored, and every subdomain needs its own file.
What is the difference between an empty Disallow and Disallow: /?
“Disallow:” with no value blocks nothing, so everything may be crawled. “Disallow: /” blocks every URL of the site.

Related tools