About the Robots.txt Generator
robots.txt is a plain text file at the root of a host (https://example.com/robots.txt) that tells crawlers which paths they may fetch, standardised as RFC 9309. It is made of groups: one or more User-agent lines followed by Allow and Disallow rules. A crawler obeys only the most specific group that names it and falls back to User-agent: * otherwise — groups are not combined, so a group for Googlebot must repeat every rule that should apply to it.
Paths are matched from the start of the URL path, case-sensitively. * stands for any sequence of characters and $ anchors the end, so /*.pdf$ blocks every PDF. When several rules match, the longest one wins, and Allow wins a tie; the order of the lines does not matter. The tester applies exactly this logic to the generated file. Crawl-delay is not part of the standard: Bing and Yandex follow it, Google ignores it.
The file is a request, not access control. Well-behaved crawlers — including the AI crawlers GPTBot, ClaudeBot, Google-Extended, PerplexityBot and CCBot — respect it, but scrapers can ignore it, and the file itself is public, so it should not list secret paths. Disallow also does not remove a page from search results: a blocked URL can still be indexed from links to it. To keep a page out of the index, allow crawling and send a noindex meta tag or X-Robots-Tag header instead.
How to use it
- Pick a preset or choose the default policy for all robots.
- List the paths to disallow and the exceptions to allow, one per line.
- Add a group for specific crawlers, such as AI bots, with their own rules.
- Enter the URL of your sitemap.
- Test a few URLs, then download robots.txt and upload it to the root of your site.