Sitemap check
Enter a site and we go looking for its sitemap the way a crawler does, read it, and list the entries a search engine would drop. If you already know where the sitemap lives, enter its address to test that file directly.
A sitemap is a list, not a command
A search engine finds pages by following links. A sitemap spares it the trouble. It’s one file where you write the address of every page you want crawled, so nothing depends on a link being found. The format is small. A <urlset> holds <url> entries, and each entry needs only a <loc> with the full address:
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://www.example.com/pricing</loc>
<lastmod>2026-03-14</lastmod>
</url>
</urlset>One file may hold 50,000 addresses and weigh 50 MB before compression. Bigger sites split the list into several files and publish a <sitemapindex>, which is a sitemap whose entries are other sitemaps. An index can’t point to another index.
Listing a page only asks for a visit. Whether the page gets indexed afterward is up to the page.
How this check finds and reads the file
Give us a bare domain and our server requests /robots.txt, then looks for Sitemap: lines. If there are several, it reads the first and lists the others so you can test them one by one. If there aren’t any, it tries /sitemap.xml, /sitemap_index.xml and /wp-sitemap.xml, in that order, and keeps the first one that contains a sitemap.
The file is then read with a streaming XML parser, up to 5 MB of download. Gzip files are decompressed first. For an index, the first five child sitemaps are fetched and read the same way, and the report says how many were left unread. Last, five addresses spread across the list are requested once each, without following redirects, to see what they answer.
What each result means
| Check | Why a crawler cares |
|---|---|
| Status, Content-Type, size | The file has to answer 200. A redirect is followed, but it changes the address the entries are judged against. An HTML answer usually means the file is missing and the site sent its “not found” page with status 200. |
| Well-formed XML, namespace | An XML parser stops at the first syntax error, so one bad character hides every entry after it. The xmlns value has to be exactly http://www.sitemaps.org/schemas/sitemap/0.9. |
| Absolute URLs, same host and scheme | A sitemap may only list pages of the host it’s served from, in the same scheme. example.com and www.example.com are two hosts. |
| Folder rule | A sitemap stored in /blog/ may only list addresses under /blog/. Declaring the file in robots.txt lifts that limit, and the check then skips it. |
| Duplicates, long or unencoded URLs | The file still works, but each one is a sign that the generator isn’t listing the canonical addresses. |
| lastmod | Has to be a W3C date: 2026-03-14 or 2026-03-14T09:30:00+00:00. Dates in the future get flagged, and so does the same date on every entry of a file with ten or more. |
| Live sample | A 404 means the list is stale. A 301 or 302 means the list gives an address that isn’t the final one. |
<priority> and <changefreq> are reported for information only. Google has said it ignores both.
Repairs for the usual faults
- The parser stops in the middle of the file
- Look for a raw
&in an address. Inside XML it’s written&, so?a=1&b=2becomes?a=1&b=2. A blank line or a PHP warning printed before<?xmlbreaks the file from its very first byte. - Addresses on the wrong host or in
http:// - The generator builds addresses from the site address in its settings. Set that to the form visitors end up on after redirects, then regenerate.
- Every lastmod is identical
- The generator is stamping the time it ran. Configure it to use the date each page was last edited, or leave the tag out. A wrong date is worse than no date: once search engines see it doesn’t match real changes, they stop trusting the tag for the whole site.
- Dead or redirected pages in the sample
- Regenerate the sitemap after deleting or moving pages, and list only addresses that answer 200 and name themselves as canonical. For the links inside your pages, there’s the dead link finder.
- Not declared in robots.txt
- Add one line with the full address anywhere in the file:
Sitemap: https://www.example.com/sitemap.xml.
What this check can’t tell you
It reads at most 5 MB per file and five children of an index, so on a very large site the counts describe the part that was read. Duplicates are looked for inside each file, not across files. Image, video, news and hreflang extensions are counted but not validated.
Five pages are a sample. They can all answer 200 while hundreds of others don’t, and a page that answers 200 may still carry a noindex tag or a canonical that points somewhere else. The on-page SEO readout shows both for a given page, and the search consoles report them for the whole sitemap once you’ve submitted it.
Questions people ask
Where is the sitemap of a website?
Open /robots.txt on the site and look for a line that starts with Sitemap:. If there isn’t one, try /sitemap.xml, then /sitemap_index.xml (Yoast and Rank Math on WordPress) and /wp-sitemap.xml (WordPress itself since version 5.5). This check goes through those steps for you.
Does a small site need a sitemap?
Not strictly. If every page can be reached from the menu in a few clicks, crawlers will find them all. A sitemap helps when the site is new and few other sites link to it, when it’s large, or when some pages can only be reached through a search form.
How many URLs can a sitemap contain?
50,000 per file, and 50 MB once uncompressed. Past either limit, split the list into several files and list them in a sitemap index, which can itself name up to 50,000 sitemaps.
Does Google use priority and changefreq?
No. Google ignores both and reads only <loc> and, once it has proved reliable on the site, <lastmod>. Bing also favors lastmod. Leaving the two other tags in does no harm.
What date format does lastmod need?
The W3C profile of ISO 8601. A day by itself is enough: 2026-03-14. With a time, put a T between date and time and end with a time zone: 2026-03-14T09:30:00+00:00 or 2026-03-14T09:30:00Z.
My sitemap is valid. Why are the pages not indexed?
A sitemap only tells search engines the pages exist. They can still decide not to index a page that is thin, duplicated, blocked by robots.txt, marked noindex, or whose canonical tag names another address. Check a few of the missing pages one by one.