Robots.txt tester

Enter an address and pick a crawler. We fetch the site’s robots.txt, apply it the way the standard says to, and tell you whether the address is open or blocked and which line decides.

The token field is read only when “Another crawler” is selected. Rules are matched against the path and query string, so give the exact address, not only the domain.

Or try en.wikipedia.org/w/index.php?title=Special:Search (Googlebot), www.nytimes.com/section/world (GPTBot), github.com/search?q=robots (Bingbot)

A notice on the door, read before every visit

Before a well-behaved crawler requests pages from a host, it asks for one file, /robots.txt. The file lists the paths the site owner would rather not have fetched, sorted by which crawler is asking. For decades it ran on habit alone. It got a written standard in 2022, RFC 9309, and this tester follows that text.

The file is made of groups. A group starts with one or more User-agent lines and continues with Allow and Disallow rules:

User-agent: *
Disallow: /cart/
Disallow: /*?sort=

User-agent: GPTBot
Disallow: /

Sitemap: https://www.example.com/sitemap.xml

Each host and each scheme has its own file. The rules of https://www.example.com say nothing about https://shop.example.com.

How the verdict is reached

It takes two choices, in this order.

  1. Which group. The crawler looks for a group that names its own token, ignoring case. If it finds one, it obeys that group and nothing else. It only falls back on User-agent: * when no group names it. That catches people out all the time: add a short group for Googlebot and you’ve freed Googlebot from every rule written under *. Groups that name the same token are read as one. A few crawlers have a documented second choice, which the tester applies. Googlebot-Image and Applebot follow the Googlebot group when they have none of their own, and YandexBot follows Yandex.
  2. Which rule. Inside the group, every rule whose pattern matches the start of the path is a candidate. The candidate with the longest pattern wins, wherever it sits in the file. If an Allow and a Disallow of equal length both match, Allow wins. If nothing matches, the address is open.

Patterns are compared with the path and query string, character by character and case-sensitive. Two characters are special: * stands for any run of characters, and $ at the very end means “the path stops here”.

PatternMatchesDoes not match
/fish/fish, /fish.html, /fishing/rods/Fish, /catfish
/fish//fish/, /fish/salmon/fish, /fish.html
/*.pdf$/guide.pdf, /docs/a.pdf/guide.pdf?v=2
/*?sort=/shoes?sort=price/shoes?page=2&sort=price

Percent-encoding is evened out before comparing, so /caf%C3%A9 and /café are the same path.

Reading the result

The summary gives the verdict, the number of the deciding line and the group that applied. Under it, the file is printed with line numbers. The lines of the applied group carry a bar, and the deciding line is highlighted. When several rules match, a small table lists them with their pattern length, and that length is all it takes to see why one beats another. A last table gives the verdict for the other crawlers on the list, because one file often treats them differently.

The status of the file matters as much as what’s in it:

  • 404 or any other 4xx: there are no rules, every address is open. That includes 403. Forbidding the file doesn’t forbid crawling.
  • 5xx, 429 or no answer: the crawler can’t know the rules and has to stay out of the whole site until the file comes back.
  • 200 with a web page inside: no valid line, so nothing is blocked. This happens when a site answers every unknown address with its home page.

The tester also reports lines a crawler will skip: rules placed before any User-agent line, unknown or misspelled directive names, patterns that don’t start with /, a file over 500 KiB, Noindex: lines (Google dropped them in 2019) and Crawl-delay, which Googlebot ignores.

What a robots.txt can’t do

It’s a request. Search engines and the large AI companies say they honor it. Scrapers and security scanners don’t, so never list a secret path there. Besides, the file is public and anyone can read it.

Blocking a page doesn’t remove it. A blocked page isn’t fetched, but its address can still show up in results when other pages link to it. To keep a page out of an index, leave it crawlable and answer with noindex in a meta tag or an X-Robots-Tag header.

This page reads the file once, from our server. A crawler may hold a copy up to a day old, and a site may serve a different file to different visitors. Each crawler also has its own parser. Where the standard is silent, on misspelled directives for instance, we read the line as Google’s published parser does and tell you so.

Questions people ask

How do I block one crawler but allow all the others?

Give that crawler its own group with Disallow: / and leave the others alone:

User-agent: GPTBot
Disallow: /

Crawlers that find no group with their name use the * group, or have no restriction if there isn’t one.

Why is my page blocked when there is an Allow rule for it?

There are two possibilities. A longer Disallow pattern also matches, and the longest pattern wins. Or the Allow rule sits in a group the crawler doesn’t read, since a crawler with its own group ignores the * group entirely. The tester shows which group applied and every rule that matched.

Does Disallow remove a page from Google?

No. It stops Google from fetching the page, not from listing its address. Use noindex on the page itself, and don’t block that page in robots.txt, or the crawler never gets to see the instruction.

Is robots.txt case-sensitive?

Paths are: Disallow: /Private doesn’t cover /private. Directive names and crawler names aren’t: user-agent: googlebot and User-Agent: Googlebot mean the same thing.

What does an empty Disallow line mean?

Disallow: with nothing after it blocks nothing. A group with only that line says “this crawler may fetch everything”. Disallow: /, with the slash, blocks the whole site.

How long does a change to robots.txt take to apply?

Crawlers keep a copy of the file, and the standard asks them not to use it for more than 24 hours. Expect a new rule to be followed within a day, sometimes sooner. A new Disallow doesn’t remove pages that are already in an index.