Robots.txt Generator and Tester
Build a robots.txt file by picking a preset and editing the rules, then check any address against it before you upload. Free, and it runs entirely in your browser.
Test a URL against these rules
🔒 Everything you type stays in your browser. Nothing is sent to a server.
By Shah Rukh, software developer · Last updated:
How to make a robots.txt file
- Pick a preset that is close to what you want, or start with “Allow everything”.
- Edit the groups. Each group names one or more crawlers (the
User-agentlines) and the paths they may or may not fetch. Use*for “every crawler”. - Add your sitemap address if you have one. Need one? Use the sitemap generator.
- Test a few addresses in the tester under the output to check the rules do what you expect.
- Download robots.txt and upload it to the top folder of your site, so it opens at
https://your-site.com/robots.txt. It only works at that exact address, and each subdomain needs its own file.
What each line means
| Line | Meaning |
|---|---|
User-agent: * | The rules below apply to every crawler that has no group of its own. |
Disallow: /private/ | Do not fetch any address that starts with /private/. |
Allow: /private/help.html | An exception inside a blocked folder. |
Disallow: (empty) | Nothing is blocked. |
Sitemap: https://… | Where your sitemap lives. It must be a full URL and applies to all crawlers. |
Crawl-delay: 10 | Asks a crawler to wait that many seconds between requests. Google ignores this line; some other crawlers, such as Bing’s, respect it. |
Wildcards and which rule wins
Two special characters are understood by the major search engines and are part of the robots.txt standard (RFC 9309): * matches any run of characters, and $ at the end of a path means “the address ends here”. So Disallow: /*.pdf$ blocks every address ending in .pdf.
When several rules match the same address, the longest matching path wins. If an Allow and a Disallow rule are equally long, Allow wins. The order of the lines does not matter. A crawler follows only one group: the one that names it most specifically, falling back to *.
A worked example
User-agent: * Disallow: /wp-admin/ Allow: /wp-admin/admin-ajax.php Sitemap: https://example.com/sitemap.xml
Here /wp-admin/options.php is blocked, because it starts with /wp-admin/. But /wp-admin/admin-ajax.php is allowed: both rules match, and the Allow path is longer. Type each path into the tester above to see the deciding rule.
Blocking AI training crawlers
The “Block AI training crawlers” preset adds a group for GPTBot (OpenAI), CCBot (Common Crawl), ClaudeBot (Anthropic) and Google-Extended. Google-Extended is not a separate crawler; it is a name Google reads to decide whether your pages may be used for its AI models, and it does not change how your site appears in Google Search. The list is a starting point, not a complete register: new crawlers appear often, so check each company’s documentation for its current names.
Limits you should know
- It is a request, not a lock. Well-behaved crawlers follow robots.txt; bad ones can ignore it. Never use it to hide private files. Use a password instead.
- Blocking is not the same as removing from search. A blocked page can still appear in results if other sites link to it. To keep a page out of search, leave it crawlable and add a
noindextag with the meta tag generator. - Don’t block CSS or JavaScript that your pages need, or search engines may not see them the way visitors do.
- The tester follows the published standard. Individual crawlers can differ in small ways, so treat it as a careful check, not a guarantee.
Paths are case-sensitive: /Photos/ and /photos/ are different. If you need to show a robots.txt snippet on a web page, the HTML entity encoder will make it safe to paste.
Frequently asked questions
Where do I put my robots.txt file?
In the top-level folder of your site, so it opens at https://your-site.com/robots.txt. Crawlers do not look for it anywhere else, and each subdomain needs its own file.
Does robots.txt remove a page from Google?
No. It only asks crawlers not to fetch the page. A blocked page can still be listed if other sites link to it. To keep a page out of results, let it be crawled and add a noindex robots meta tag.
Does Google follow Crawl-delay?
No. Google ignores the Crawl-delay line. Some other crawlers, such as Bing’s, do read it.
What do * and $ mean in a robots.txt path?
* matches any run of characters and $ at the end of a path means the address must end there. For example, Disallow: /*.pdf$ blocks every address that ends in .pdf.
Which rule wins if Allow and Disallow both match?
The rule with the longest matching path wins. If both are the same length, Allow wins. The order of the lines makes no difference.
Will blocking AI crawlers stop all AI companies using my content?
No. It only works for crawlers that choose to follow robots.txt and whose names are in your file. The preset covers a few well-known names and is not a complete list.
Related tools
Telegram Bot Maker
Build a Telegram bot without coding. Download ready-to-run Python.
Chrome Extension Maker
Make a site blocker, new tab page or link popup extension. No code.
Discord Embed Builder
Design Discord embeds with a live preview, JSON and discord.py code.
WhatsApp Link Generator
Make a wa.me chat link with a message, QR code and button.
Privacy Policy Generator
Answer a few questions and get a privacy policy template.
Terms & Conditions Generator
Tick the sections you need and get a terms and conditions template.