Robots.txt generator: build the file and test every rule
Pick your site type, adjust the paths and take the finished robots.txt to your domain root. Right below it, a tester tells you what happens to a specific URL: which group decided it, which rule won and why. Free, no signup, and nothing leaves your browser.
Your robots.txt
Test a URL against this file
-
Paste the path or the full URL. The verdict follows the protocol: most specific group first, then longest matching pattern.
The file in numbers
- - groups
- - rules
- - blocks
- - exceptions
- - sitemaps
- - size
Checks
Everything runs in your browser: neither the paths nor the file leave this page, and no URL is fetched. The tester applies the protocol rules to the generated text, the same way a crawler does. Remember that robots.txt controls CRAWLING, not indexing: a URL blocked here can still show up in search if someone links to it.
Who this page is for: anyone who needs to control which paths of a site may be crawled, whether to keep catalog filters, internal search and account areas out of a crawler's way, or to close a staging environment. If your question is the other one, which AI bots may train on, cite or fetch your content, the right place is the AI bots robots.txt generator, which handles the same file from the angle of access policy. Both pages write a complete robots.txt because there is only one such file; what changes is the decision each one helps you make. For the rest of the page head, use the meta tag generator.
The rule that decides is not the first one, it is the longest
Almost every robots.txt generator hands you a block of text and disappears. The trouble is that the file is not read top to bottom like a script: it is a set of patterns competing with each other, and people who write it by hand find that out late, when an entire section of the site has vanished from the index and nobody knows which line did it. That is why this tool has two halves. The top one builds the file; the bottom one answers the question that actually matters, which is what happens to a specific URL.
Three rules explain nearly every surprise. First: a crawler obeys one group only, the most specific one that matches its name, and ignores the others completely. Second: inside that group, the longest pattern wins, with Allow taking the tie. Third: absence of a rule is permission, so whatever nobody forbade is allowed. Together they explain why moving lines around never fixes anything and why opening a group for one bot can unlock exactly what you meant to close.
How to use it
- Pick a site type. The preset fills the list with that platform's draft, and it is a starting point rather than truth: every install has its own paths.
- Adjust the blocked paths, one per line. You can paste a full URL and the path is extracted for you, which avoids the classic mistake of writing a rule that starts with https.
- Use the exceptions to open one address inside a closed folder. They beat the block because the pattern is longer, not because they sit on another line.
- Declare the sitemap with an absolute URL. That line belongs to no group and applies to every crawler.
- Test in the box below: pick a crawler, paste a path and read the verdict. Repeat with the URLs you cannot afford to get wrong, such as a product page, the sitemap and a CSS file from your theme.
How it works: pattern matching, term by term
The tester applies the protocol's algorithm to the text sitting in the output box, not to the form fields. That is deliberate: you test the file you are about to publish, with the same ambiguities a crawler will meet.
Pattern length is literal, counted character by character. Disallow: /wp-admin/ is 10 characters; Allow: /wp-admin/admin-ajax.php is 24. Since 24 beats 10, that specific file stays open inside a closed folder, and it holds even though the Allow is written afterwards. Keep that number in mind: it is what you see in the tester when two rules compete.
Worked example (reproduces the default output)
The tool opens with the WordPress preset and one declared sitemap. Checking the numbers panel, item by item:
- 1 group and 5 rules: only the asterisk User-agent exists, with four blocks and one exception.
- 4 blocks: the admin folder, the login screen, internal search and the comment reply parameter, which is the quiet multiplier of duplicate URLs on WordPress.
- 1 exception and 1 sitemap. No check is lit, because nothing here is in conflict.
Now look at the tester, which already carries Googlebot and the path /wp-admin/admin-ajax.php. The verdict is Allowed, decided by Allow: /wp-admin/admin-ajax.php, and the line below reports that 2 rules matched. Those two are the folder Disallow and the file Allow: the longer one won. That is the whole mechanism in a single line.
Swap the path for /wp-admin/options.php and change nothing else. The verdict becomes Blocked, decided by Disallow: /wp-admin/, and now only one rule matched, because the exception is specific to one file. Swap again for /blog/new-post/ and the verdict goes back to Allowed, this time with no rule involved at all: the box says nothing matched, which is the third rule of the protocol in action.
Finally, add the line /*.css to the blocked list, a common attempt at saving crawl budget. The numbers move to 6 rules and 5 blocks, and a red check lights up naming two sample resources that just became unreachable: the theme stylesheet and the build CSS. It is the most useful warning on the page, because that mistake produces no error anywhere: it only changes what the search engine sees.
How to read each number
Groups is how many User-agent blocks the file has. While it stays at 1, every crawler follows the same rules and the file is easy to maintain. From 2 upward, check agent by agent in the tester, because each new group switches the general one off for whoever falls into it.
Blocks against exceptions describes the shape of the file. Many exceptions usually signal a block that was too broad higher up: instead of closing a whole folder and reopening five addresses, it is almost always simpler to close the five paths that actually get in the way.
Size rarely matters, but there is a ceiling: Google reads the first 500 KB and ignores the rest. A file anywhere near that has become a script generated exclusion list, and that logic belongs in your URL patterns, not in robots.txt.
The checks subtract no score, because there is no score here. Blocking checkout is neither better nor worse than allowing it: it depends on the site. What the tool verifies is coherence, meaning the cases where the file does something different from what you meant.
Known limits
The tool checks logic, not existence. It does not visit your site, does not know whether /cart/ exists and does not know whether that path already carries a noindex. After publishing, confirm in the Search Console page indexing report and the live URL test, which is where real crawling shows up.
It is also worth knowing what the file does not do. It removes nothing from the index, protects no private area and is not mandatory: whoever wants to ignore it will, because obedience is voluntary and only serious crawlers commit to it. And it is case sensitive on paths, so /Blog/ and /blog/ are different things to a rule, even if your server treats both as the same page.
Finally, the pattern matching applied here is Google's, which is also the one published as RFC 9309. Bing and Yandex follow the same design in essence, with differences on secondary points, and Crawl-delay is exactly one of them. If your traffic comes from a search engine outside that list, test in its own official tool before trusting this result.
From the published file to the live test
Controlling crawling is the start. What decides outcomes is what happens to the people who arrive:
- Decide your AI training and citation policy in the AI bots robots.txt generator, the same file from the other angle.
- Hand models a curated map of your site with the llms.txt generator.
- If the site has versions in more than one language, build the cluster in the hreflang tag generator.
- See how your result looks in search, with the exact truncation point, in the SERP preview tool.
- Before swapping one page for another because you think it converts better, check how much traffic the test needs in the sample size calculator.
Frequently asked questions
- Does blocking a page in robots.txt remove it from Google?
- Not necessarily, and this is the most expensive misunderstanding about the file. Robots.txt controls crawling, not indexing: it asks a crawler not to fetch that URL. If other pages link to it, the address can still show up in search, with no description and a note saying the content could not be read. To take a page out of the index the right instrument is a meta robots noindex, or the X-Robots-Tag header. And there is a trap in combining the two: if you block the URL in robots.txt, the crawler will never fetch the page and will never see the noindex you put on it. To remove a page, allow crawling and apply noindex; only after the removal happens does blocking the path make sense.
- Does the order of the lines change the result?
- It does not. Inside a group, the rule with the longest pattern wins, not the one that appears first. An exact tie between an Allow and a Disallow of the same length is resolved in favor of Allow. That is why Allow: /wp-admin/admin-ajax.php opens that file even though it is written after Disallow: /wp-admin/, and it is also why moving lines around to fix a behavior never works. What changes the result is pattern length. The tester on this page shows which rule won and how many competed, precisely so you stop guessing.
- What happens if I open a group for Googlebot?
- It starts obeying that group only and ignores the general group entirely. This is the rule that breaks files most often in practice: someone wants to grant Googlebot one extra permission, writes a two line group, and without noticing opens to it the whole private area that was closed under User-agent asterisk. If you really need a specific group, repeat inside it every block that should still apply. In the tester, pick the agent and read the group line: it tells you which group decided the URL.
- Can I use regular expressions in the paths?
- No. There are exactly two wildcards: the asterisk, which matches any sequence of characters, and the dollar sign, which anchors the end of the URL. Everything else is literal, including dots, question marks, parentheses and brackets. That means /*.pdf is not the same as /*.pdf$: the first one also matches /manual.pdf.html, because without the anchor the pattern only has to appear somewhere in the path. An asterisk at the end of a pattern is harmless and useless, since matching was always by prefix. The tool flags both cases.
- Does blocking CSS and JavaScript save crawl budget?
- It saves bandwidth and costs rankings, which rarely pays off. A search engine renders your page like a browser to know what it actually shows: with no stylesheet and no scripts it sees a skeleton that may be missing the main content, may look broken in the mobile test and may lose entire blocks loaded on the client. The effect is not immediate and never appears as an error; it arrives weeks later, as a drop with no visible cause. This tool tests a sample of typical theme, asset folder and build paths, and warns you when a rule catches one of them.
- Does Crawl-delay work on Google?
- No. Google ignores the directive and adjusts its crawl rate on its own, watching how your server responds; when you really need to slow it down, the place to do that is Search Console. Bing and Yandex honor the line. It is also worth remembering that Crawl-delay is the wrong cure for a slow server: it spaces out the bot that obeys and does nothing about real traffic, which is usually the actual cause. If the problem is load, the place to fix it is caching and infrastructure.
- Is robots.txt a way to protect a private area?
- It is not, and using it that way makes things worse. The file is public, sits at a fixed address and is the first thing any curious person opens on your site. Listing /secret-admin/ there publishes a map of the places you want hidden, and malicious crawlers simply ignore the request, because obedience is voluntary. An area that must not be seen is protected with authentication, not with a line of text politely asking people to stay out.
- Do I still need the Sitemap line if I submitted it in Search Console?
- You do, and it is cheap. The Sitemap line belongs to no group: it applies to every crawler, including the ones with no dashboard where you could submit anything, such as Bing, DuckDuckGo and the AI agents that now send traffic. Submitting in Search Console only solves it for Google. The line requires an absolute URL, with protocol and domain, and it is the only item in the file that actively helps whoever crawls properly instead of merely restricting.
- Is anything I type here stored somewhere?
- No. Everything runs in JavaScript in your browser: the file is built, parsed back and tested on your machine, with no network call at all, and nothing is saved or logged. The tool also does not visit your site, which means it checks the logic of the rules you wrote and not the existence of the paths. That matters when your list reveals the internal structure of a project that is not live yet. You can confirm it by opening the network tab while you type: no request goes out.
Keep reading
To understand what changes between optimizing for search and optimizing for conversion, read CRO vs SEO, what changes in each, see how to survive answers without clicks in zero click search optimization and learn to write quotable blocks in FAQ content for AI citation.