What this XML sitemap validator checks
Enter a domain and the validator reads your robots.txt for Sitemap: lines, then tries /sitemap.xml and /sitemap_index.xml. Enter a sitemap URL and it starts there. If the file is a sitemap index, it reads up to 50 of the child sitemaps it lists, compressed or not, and draws them as a tree.
Every file is checked against the protocol and what Google and Bing document:
- Fetching. An HTTP 200, an XML content type, and no HTML page where the sitemap should be.
- Structure. Well-formed XML, a
<urlset>or<sitemapindex>root, and the exact namespacehttp://www.sitemaps.org/schemas/sitemap/0.9. - Every URL. Absolute, http or https, properly escaped, under 2,048 characters, on the same host as the sitemap, with no duplicates and no mix of http and https or www and non-www.
- Dates and hints.
<lastmod>in W3C Datetime, not in the future, and not the same on every URL.<changefreq>and<priority>are noted, with the fact that Google and Bing ignore them. - Limits and nesting. No file over 50,000 entries or 50 MB uncompressed, and no sitemap index inside another.
- robots.txt. Whether the sitemap is declared, whether Googlebot may fetch it, and whether any listed URL is disallowed for Googlebot, Bingbot, GPTBot, ClaudeBot and PerplexityBot.
Then it does the part a pure XML check cannot. It requests 20 URLs from across your sitemaps, once each and without following redirects, and reports which answer 200, which redirect, which fail, which carry a noindex, and which name a different canonical URL. Every finding comes with a plain-English fix, and where it helps, a corrected snippet you can copy.
Sources: the sitemap protocol[1]; Google[3].
The sitemap protocol rules, in one place
The format is defined at sitemaps.org, and Google, Bing and other engines follow it[1]. A sitemap is a UTF-8 XML file whose root element is <urlset>. Each page is a <url> entry with one required child, <loc>, and three optional ones: <lastmod>, <changefreq> and <priority>.
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://www.example.com/</loc>
<lastmod>2026-09-30</lastmod>
</url>
<url>
<loc>https://www.example.com/pricing?plan=team&billing=annual</loc>
<lastmod>2026-10-02T09:15:00+00:00</lastmod>
</url>
</urlset>The rules that trip people up:
- Full URLs only. Each
<loc>must start with the protocol and be less than 2,048 characters[1]. Google asks for fully qualified, absolute URLs[2]. - Escape for XML. An
&in a URL must be written&, and the same goes for<,>, and quotes[1]. One raw ampersand makes the whole file unparseable, and Search Console names an unescaped character in a URL as the usual cause of a parsing error[5]. - One host, one protocol. All URLs in a sitemap must use the same protocol and sit on the same host as the sitemap. A sitemap at
/catalog/sitemap.xmlmay only list URLs under/catalog/[1]. Google applies the folder rule unless you submit the sitemap in Search Console[2]. - Size. At most 50,000 URLs and 50 MB uncompressed per file. You may gzip it to save bandwidth, but the limit counts the uncompressed size[1][2].
- Dates.
<lastmod>uses W3C Datetime: a date such as2026-09-30, or a date and time with a time zone such as2026-09-30T09:15:00+00:00[8]. A time without a zone, or a space instead of theT, is invalid, and Search Console reports it as an invalid date[5].
The protocol also allows a plain text file with one URL per line, in UTF-8[1]. It works, and this validator reads it, but it has nowhere to put a <lastmod>, so it gives search engines no freshness signal.
What Google and Bing actually use
Of the three optional tags, only one matters. Google says it ignores <priority> and <changefreq>, and uses <lastmod> if it is consistently and verifiably accurate[2]. Bing said the same about the two hints in July 2025: they are ignored and do not influence how content is crawled or ranked. It called <lastmod> a key signal for deciding which URLs to recrawl, or to skip because nothing changed[7].
Sources: Google[2]; sitemaps.org[1]; Yoast[11].
That makes the identical-date pattern worth catching. When every URL carries the same <lastmod>, the generator is almost always writing the time it ran rather than the time each page changed. The dates then say nothing true, and an engine that checks them learns to ignore them. This validator flags it when every dated URL, or 95% of a large set, share one value.
A sitemap is also not a requirement. Google says a site of about 500 pages or fewer, well linked from its home page, may not need one, and that one helps most on large sites, new sites with few links, and sites with a lot of media[4]. Submitting it is a hint, not a guarantee of crawling[2].
Sitemap index files
When a site outgrows one file, it splits the URLs into several sitemaps and lists them in a sitemap index. The index has the same shape, with <sitemapindex> as the root and a <sitemap> entry for each child. It may hold up to 50,000 child sitemaps, and Search Console accepts up to 500 index files per site[3].
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<sitemap>
<loc>https://www.example.com/sitemaps/posts.xml</loc>
<lastmod>2026-10-01</lastmod>
</sitemap>
<sitemap>
<loc>https://www.example.com/sitemaps/products.xml.gz</loc>
</sitemap>
</sitemapindex>Three rules apply to the index:
- No nesting. “A sitemap index file can’t list other sitemap index files, only sitemap files”, in Search Console’s words[5]. The validator reads the children of your index and flags any child that is itself an index.
- Same site, same folder or lower. Google requires child sitemaps to be hosted on the same site as the index, and in the same directory or lower[3].
- Dates are about the child file. In an index,
<lastmod>is when that sitemap changed, in the same W3C format[3].
Bing gives the scale in round numbers: 50,000 URLs per file and 50,000 files per index make up to 2.5 billion URLs per index[7].
Common sitemap errors and how to fix them
Most broken sitemaps fail in one of a handful of ways. The names below match the ones Search Console uses in its sitemaps report[5], so you can match a validator finding to a Search Console error.
| Error | Usual cause | Fix |
|---|---|---|
| Parsing error | A raw & in a URL, or an HTML entity like that XML does not define | Escape & as &. Use numeric references for other characters. |
| Incorrect namespace | https:// in the namespace, a trailing slash, or an old Google namespace | Use http://www.sitemaps.org/schemas/sitemap/0.9 exactly. |
| Invalid date | 2026-09-30 10:00:00, a missing time zone, or day-first dates | Write 2026-09-30 or 2026-09-30T10:00:00+00:00. |
| URL not allowed | URLs on www while the sitemap is on the bare domain, or the reverse | Serve the sitemap from the host its URLs use. |
| Sitemap is HTML | The URL serves a page: a moved sitemap, or a 404 answered with 200 | Point to the real XML file. Make missing URLs return 404. |
| Nested indexing | An index lists another index, often from two plugins | List the inner index's children directly. |
| Too many URLs | More than 50,000 entries in one file | Split the file and list the parts in an index. |
The quieter problems are the ones a schema check passes. A sitemap can be perfectly valid and still list URLs that redirect, URLs that return 404 after a page was deleted, pages carrying noindex, and duplicates whose canonical points elsewhere. Google asks you to list the canonical URLs you want in search results[2], so each of those is a sitemap entry that should not be there. The sample table in the result shows them, and the redirect checker traces any chain hop by hop.
Disallow rule that matches it in robots.txt says “do not crawl this”. Search engines get both messages and you get neither result. The validator checks every listed URL against your robots.txt with the RFC 9309 matching rules, where the longest matching rule wins[9], and shows each crawler separately. Test a single path in the robots.txt tester.Fixing the sitemap on your platform
Most sites do not write their sitemap by hand. The fix is usually a setting in the tool that generates it, so start by finding out which one does.
/wp-sitemap.xml, adds it to the robots.txt WordPress generates, and lists only URLs, with no <lastmod> by default. SEO plugins such as Yoast switch the core sitemap off and serve their own[10]. If you see both /wp-sitemap.xml and a plugin’s sitemap, one of them is stale.// Only if another plugin serves your sitemap. add_filter( 'wp_sitemaps_enabled', '__return_false' );
/sitemap_index.xml, and that is the URL to submit[12]. Post types set to noindex are left out automatically, and Yoast dropped <priority> because Google does not use it[11]. If the index will not load, Yoast asks you to remove other sitemap plugins and any physical sitemap files first[11]. Entries per file can be lowered with the wpseo_sitemap_entries_per_page filter./sitemap_index.xml, defaults to 200 links per sitemap file, and adds the sitemap to robots.txt for you[13]. If the validator says robots.txt does not declare it, check the plugin’s robots.txt settings and any robots.txt file uploaded to the server by hand./sitemap.xml automatically, as an index linking to separate sitemaps for products, collections, blogs and pages, and updates it as you add content[14]. Because it is generated, the fixes are upstream: if the sample shows listed URLs that redirect, 404 or carry noindex, change those products, collections or pages in the admin. A store in password-protected (private) mode cannot be read by search engines at all[14].app/sitemap.ts returns an array and Next.js serves it as /sitemap.xml. For large sites, generateSitemaps splits it into files served at /.../sitemap/[id].xml[15]. Pass each page’s real modification date. new Date() stamps every URL with the build time, which is exactly the identical-lastmod pattern. More on what crawlers get from Next.js in the Next.js SEO guide.import type { MetadataRoute } from "next";
import { getPosts } from "@/lib/posts";
export default async function sitemap(): Promise<MetadataRoute.Sitemap> {
const posts = await getPosts();
return [
{ url: "https://www.example.com/", lastModified: "2026-09-30" },
...posts.map((p) => ({
url: `https://www.example.com/blog/${p.slug}`,
lastModified: p.updatedAt, // the post's own date, not new Date()
})),
];
}Declare it in robots.txt
Search Console and Bing Webmaster Tools are where you submit a sitemap to Google and Bing. Every other crawler has to find it on its own, and the one place they all look is robots.txt. The protocol defines a Sitemap: line that takes the sitemap’s full URL[1], Google accepts it as a way to submit[2], and Bing recommends it alongside Bing Webmaster Tools[7]. RFC 9309 leaves such lines outside the user-agent groups, so put it anywhere in the file[9].
User-agent: * Disallow: Sitemap: https://www.example.com/sitemap_index.xml
Use the full URL, protocol included. A bare path such as Sitemap: /sitemap.xml is not what the protocol describes, and the validator flags it. To build a file from scratch, use the robots.txt generator.
Why sitemaps matter for AI search too
Start with what is documented. Bing’s July 2025 post on sitemaps in AI powered search says freshness signals directly influence how quickly updates are reflected in its search results and AI generated answers, and that <lastmod> remains a key signal for what it recrawls[7]. Google says a page must be indexed and eligible for a snippet in Google Search to appear as a supporting link in AI Overviews or AI Mode, with no additional technical requirements[6]. A sitemap that helps a page get indexed therefore helps it become eligible for both.
What is not documented: OpenAI, Anthropic and Perplexity do not publish sitemap guidance for GPTBot, ClaudeBot or PerplexityBot that we could find. Nobody outside those companies can say how much weight their crawlers give a sitemap, so be wary of anyone who claims to. Two things are within your control. Declare the sitemap in robots.txt, where any crawler can find it. And make sure robots.txt does not block the URLs you list, which is why this validator checks GPTBot, ClaudeBot and PerplexityBot next to Googlebot and Bingbot. Blocking a training crawler can be a deliberate choice, so the validator marks AI blocks as warnings, not errors. The AI crawler checker shows what your file says to every major AI bot, and server log analysis shows which ones actually visit.
XML sitemap checklist
Run through this after a redesign, a migration or a change of SEO plugin. Your ticks are saved in this browser only.
Check every page the sitemap lists
This validator requests 20 of your URLs. A sitemap with thousands of entries can hide a whole section of redirects or noindexed pages that a sample of twenty misses. The GEO Audit starts from the same discovery, robots.txt then the sitemap, then crawls the pages and runs 31 checks on each, including whether it answers a clean 200 and whether AI crawlers are allowed in.
The free account audits 10 pages with an email address and no card. Paid plans audit more and re-run on a schedule; every limit is on the pricing page.
Questions people ask
What does an XML sitemap validator check?
That the file is valid XML, that its root is <urlset> or <sitemapindex> with the exact sitemaps.org namespace, that every <loc> is an absolute, escaped URL on the sitemap's own host, that <lastmod> dates use W3C Datetime, and that the file stays within 50,000 URLs and 50 MB uncompressed. This one also follows sitemap indexes, reads .gz files, checks robots.txt, and requests 20 listed URLs to see whether they answer 200.
How do I find my sitemap?
Look in your robots.txt for a line starting with Sitemap:. If there is none, try /sitemap.xml and /sitemap_index.xml on your domain. WordPress core serves /wp-sitemap.xml, Yoast SEO and Rank Math serve /sitemap_index.xml, and Shopify serves /sitemap.xml. Enter just your domain above and this validator tries all of them.
Does Google use changefreq and priority?
No. Google's documentation says it ignores <priority> and <changefreq>, and Bing said in July 2025 that it ignores both too. Google does use <lastmod> when it is consistently and verifiably accurate, and Bing calls lastmod a key signal for deciding what to recrawl.
What is the maximum size of a sitemap?
50,000 URLs or 50 MB (52,428,800 bytes) uncompressed, whichever comes first. Gzip saves bandwidth but does not raise the limit, which applies to the uncompressed file. A sitemap index has the same limits, counted in sitemaps rather than URLs.
Can a sitemap index contain another sitemap index?
Not for Google. Search Console's help says a sitemap index file can't list other sitemap index files, only sitemap files, and reports it as nested indexing. List every child sitemap directly in one index, or submit each index separately.
Should my sitemap list URLs that redirect?
No. Google asks for the canonical URLs you want to appear in search results, and a URL that redirects is not one. Replace each redirecting URL with its destination, and remove URLs that return 404, carry noindex or name a different canonical. The sample table above shows which of yours do.
Do AI crawlers use sitemaps?
Bing says sitemaps and accurate lastmod dates directly influence how quickly updates reach its search results and AI generated answers. Google's AI features only use pages that are indexed in Google Search, which is what sitemaps help with. OpenAI, Anthropic and Perplexity do not publish sitemap guidance for their crawlers that we could find, so the safe position is to declare your sitemap in robots.txt where any crawler can see it.
