What this canonical checker reads
Enter a URL and the checker requests it the way a browser would, following redirects one hop at a time so you see the chain. On the page it lands on, it reads every canonical signal, not only the first tag it finds:
- Every
<link rel="canonical">, with where it sits: in the<head>, in the<body>, or after an element that ends the head early. Relative hrefs are resolved, against<base href>when the page sets one. - The HTTP
Linkheader, the other way to declare a canonical, and whether it agrees with the HTML. - og:url and the final URL, so you can see at a glance which URL each system is being told about.
- noindex in a robots or googlebot meta tag, or in an
X-Robots-Tagheader. - hreflang alternates, from the head and the Link header.
Then it follows the canonical. It fetches the canonical URL once and reports its status, whether it carries noindex, and whether it names yet another URL as its own canonical. In single mode it also fetches up to 5 hreflang alternates to check each is self-canonical and links back, and requests the page again with GPTBot’s and ClaudeBot’s user-agent strings to compare the canonical they are served. Bulk mode checks up to 5 URLs at once, without the crawler and hreflang steps. Every finding comes with a fix and a copy-ready tag.
What a canonical URL is
Google defines canonicalization as the process of selecting the representative, or canonical, URL of a piece of content[1]. When several URLs show the same or very similar main content, Google clusters them and picks one to crawl most regularly and to show in results; the others are crawled less often[1].
Duplicates are normal. Google lists the usual causes: region variants, mobile and desktop versions, http and https, filtering and sorting parameters, and accidental copies such as a staging site left public[1]. Bing’s list adds uppercase and lowercase URLs, trailing slashes and printer-friendly pages[11]. Google says some duplicate content on a site is normal and not a violation of its spam policies[1], and Bing says duplicate content does not trigger a penalty on its own, but dilutes authority and slows how updates are picked up[11].
A canonical tag is how you state your preference among those URLs. Google gives four reasons to state one: to choose which URL people see in search results, to consolidate signals such as links into one URL, to simplify tracking metrics, and to avoid spending crawl time on duplicates[2].
<head> <title>Trail running shoes</title> <link rel="canonical" href="https://www.example.com/shoes/trail/" /> </head>
Sources: Google[2][4][8]; RFC 6596[10].
How Google chooses the canonical
Your canonical tag is one input among several. Google names the factors: whether the page is served over HTTP or HTTPS, redirects, presence in a sitemap, and rel="canonical" annotations[1]. It then says plainly that “indicating a canonical preference is a hint, not a rule”, and that it may choose a different page for various reasons[1].
The methods do not carry equal weight. In Google’s own words:
- Redirects are “a strong signal that the target of the redirect should become canonical”[2].
- rel="canonical", as a link tag or an HTTP header, is also a strong signal[2].
- Sitemap inclusion is a weak signal[2].
- Other signals: Google prefers HTTPS pages over equivalent HTTP pages, except when there are issues or conflicting signals, and prefers URLs that are part of hreflang clusters[2].
The methods stack: “when you use two or more of the methods, that will increase the chance of your preferred canonical URL appearing in search results”[2]. The reverse also holds. Google asks you not to name one URL in the sitemap and another in the canonical tag, and to link internally to the canonical URL rather than a duplicate[2]. When internal links, sitemap and canonical all agree, the hint is hard to overrule. When they disagree, you are asking Google to guess.
Sources: Google[1][2][3]; Bing[11].
The HTTP header form (RFC 6596)
The canonical link relation is an IETF standard. RFC 6596, written by Maile Ohye of Google and Joachim Kupke in April 2012, defines “canonical” as the preferred IRI from a set of resources with duplicated content, and says “the target (canonical) IRI MUST identify content that is either duplicative or a superset of the content at the context (referring) IRI”[10]. Because it is a link relation, it can be sent as an HTTP header as well as a tag:
Link: <https://www.example.com/downloads/white-paper.pdf>; rel="canonical"
That is the only way to give a canonical to a PDF or any other file that has no <head>, and Google documents it for exactly that, noting that it supports the header for web search results only[2]. On HTML pages the header usually appears by accident: a CDN rule, a proxy or a server module adds one that nobody remembers. If it names a different URL from the tag in the HTML, you have two canonicals for one page, and this checker reports it as an error.
The RFC also lists the mistakes to avoid, and they map one to one onto the checks above. The canonical target should not be the source of a permanent redirect, should not return an error, should not itself name a different canonical (a chain), and you should specify only one canonical link relation for a resource[10].
Common canonical mistakes and how to fix them
Google’s troubleshooting guide notes that some content management systems and plugins use canonicalization techniques incorrectly and point to undesired URLs[3]. These are the patterns this checker looks for, and what to do about each.
| Mistake | What happens | Fix |
|---|---|---|
| Canonical in the <body> | Google only accepts rel=canonical in the head, so the tag is ignored. | Move it into the <head>. |
| <img> or <iframe> above it in the head | Google treats the first invalid element as the end of the head and stops reading. | Put the canonical first, or move the invalid element into the body. |
| Relative URL | It resolves against whatever host served the page, including staging and mirrors. | Write the full https:// URL. |
| Two different canonicals | Usually two plugins, or a theme plus a plugin. The signals cancel out. | Switch the canonical off in all but one system. |
| Header and HTML disagree | Two techniques name two URLs for one page. | Remove one, usually the server or CDN header. |
| Canonical redirects | You name a URL that itself says "not me". | Name the redirect's destination directly. |
| Canonical answers 404 or a server error | Signals are sent to a page that cannot be indexed. | Point at the live page, usually this one. |
| Canonical is noindex | "Index that one" meets "do not index me". | Remove the noindex, or point elsewhere. |
| Tracking parameters kept | Every campaign link names a different canonical. | Build the canonical from the route, not the request. |
| Page 2 names page 1 | Items listed only on later pages lose their path into the index. | Make each paginated page self-canonical. |
Three of these deserve a closer look. First, the head. Google accepts only title, meta, link, script, style, base, noscript, template in the <head>, and “once Google detects one of these invalid elements, it assumes the end of the <head> element and stops reading any further elements”[4]. A tracking pixel <img> pasted near the top of the head is enough to hide every tag below it, canonical included. The checker names the element that ends your head early.
Second, pagination. Google says: “Don’t use the first page of a paginated sequence as the canonical page. Instead, give each page its own canonical URL”[7]. RFC 6596 lists the same mistake[10].
Third, JavaScript. Google lets you set the canonical with JavaScript, but says not to use JavaScript to change it to something other than the URL in the original HTML, and that the best way to set it is in HTML[5]. This checker reads the HTML your server sends, before any script runs, which is also what crawlers that do not run JavaScript receive. If it finds no canonical where your browser’s inspector shows one, a script is adding it.
Fixing the canonical on your platform
Most canonicals are printed by a plugin, a theme or a framework, so the fix is usually a setting, not a hand-written tag.
wpseo_canonical filter[13].within filter, which produces URLs such as /collections/sale/products/item. Shopify’s own docs warn that because the standard product page and the collection version “have the same content on separate URLs, you should consider the SEO implications”[15]. Run one of those collection URLs through the checker: on a healthy store the canonical names the plain /products/ URL. Since Google asks for internal links to the canonical[2], linking product cards straight to /products/ reinforces the tag instead of contradicting it.alternates.canonical in a page’s metadata or generateMetadata, and Next.js prints <link rel="canonical"> in the head. With metadataBase set in the root layout, a relative path is composed into a full URL, and an absolute URL ignores metadataBase[16]. More on what crawlers get from Next.js in the Next.js SEO guide.// app/layout.tsx
export const metadata = { metadataBase: new URL("https://www.example.com") };
// app/shoes/trail/page.tsx
export const metadata = {
alternates: { canonical: "/shoes/trail/" },
};
// prints <link rel="canonical" href="https://www.example.com/shoes/trail/" />Canonical vs redirect vs noindex
The three tools solve different problems, and mixing them up is behind most canonical errors. A redirect removes the duplicate: Google says to use one when you want to get rid of existing duplicate pages[2]. A canonical keeps both URLs live for visitors and asks search engines to consolidate them. A noindex keeps a page out of search entirely, and Google does not recommend it for choosing a canonical within one site[2].
| Canonical | 301 redirect | noindex | |
|---|---|---|---|
| Visitors see the duplicate | Yes | No, they land on the target | Yes |
| Signals consolidate | Into the canonical, if Google agrees | Into the target | No consolidation |
| Use it for | Sort, filter and tracking variants, syndicated copies | Moved or merged pages, http to https, www | Pages that should not be in search at all |
| Strength | Strong hint | Strong signal | Removes the page from Search |
Reading the Search Console canonical statuses
Search Console’s page indexing report uses three canonical statuses, and they mean different things[8]:
- Alternate page with proper canonical tag. The page correctly points to an indexed canonical, “so there is nothing you need to do”.
- Duplicate without user-selected canonical. The page is a duplicate, declares no preference, and Google chose another URL as canonical, so it will not serve this page in Search. Add a canonical to make the choice yours.
- Duplicate, Google chose different canonical than user. Your canonical was overruled. Google has indexed the URL it considers canonical instead. Run both URLs through this checker and look for signals that disagree.
To see Google’s choice for any URL, use the URL Inspection tool, and after fixing a conflict, use Request Indexing to ask Google to re-evaluate the page[3].
Canonicals and hreflang
International sites are where canonicals go wrong most expensively. Google says that if you use hreflang, you should specify a canonical page in the same language, or the best possible substitute if none exists[2], and that localized versions are only considered duplicates if the main content is untranslated[6]. Pointing the German page’s canonical at the English page asks Google to drop the German page.
Google also requires each language version to list itself as well as every other version, and warns that if two pages do not both point to each other, the tags will be ignored[6]. When a page has hreflang links, the checker confirms the page lists itself, then fetches up to 5 alternates and checks that each answers 200, names itself as canonical and links back. It does not validate language codes or fetch every alternate on a large set, so treat it as a spot check.
What Bing says about canonicals
Bing’s December 2025 post on duplicate content speaks to canonicals and AI answers directly. It recommends 301 redirects to consolidate variants, canonical tags “on variations that do not represent a distinct search intent”, and says that with clear canonical tags, consistent metadata and IndexNow “you can reinforce which version matters and help search engines and AI systems surface the correct page”[11]. Our Bing Webmaster Tools guide covers IndexNow and the rest of the setup.
Canonicals in AI search: what is known
Start with what is documented. Bing says that “LLMs group near-duplicate URLs into a single cluster and then choose one page to represent the set”, and that when the differences between pages are minimal, “the model may select a version that is outdated”[11]. Google says a page must be indexed and eligible to be shown with a snippet to appear as a supporting link in AI Overviews or AI Mode, with no additional technical requirements[9]. Since a Google result usually points to the canonical page[1], the canonical you get right for Search is the URL that can be linked from Google’s AI features.
What is not documented: OpenAI’s crawler documentation and Perplexity’s do not mention canonical tags at all[17][18]. Nobody outside those companies can tell you how GPTBot, OAI-SearchBot or PerplexityBot weigh a rel="canonical", so be wary of anyone who claims to. What you can check is whether those crawlers receive the same page and the same canonical as everyone else, which is why this checker fetches the page as GPTBot and ClaudeBot. It sends their user-agent strings from our server, not from their IP addresses, so a site that verifies crawlers by IP can treat the real bots differently. A difference is a strong hint, not proof. To see which AI crawlers your robots.txt lets in, use the AI crawler checker.
Canonical tag checklist
Run through this after a redesign, a migration, or a change of theme or SEO plugin. Your ticks are saved in this browser only.
Check the canonical on every page
This checker tests one URL, or 5 at a time. Canonical problems rarely stay on one page: a theme or plugin prints the same wrong tag on every product, post or archive page. The GEO Audit crawls your site and runs 31 checks on each page, including whether it declares a canonical URL, whether it answers a clean 200 and whether AI crawlers are allowed in.
The free account audits 10 pages with an email address and no card. Paid plans audit more and re-run on a schedule; every limit is on the pricing page.
Questions people ask
What does a canonical tag checker check?
It reads the <link rel="canonical"> tag on a page and tells you which URL it names. This one also reads the HTTP Link header, og:url and the final URL after redirects, flags a tag in the <body> or after an element that ends the head early, then fetches the canonical URL to see whether it answers 200, redirects, returns 404 or carries noindex. In single mode it checks up to 5 hreflang alternates and fetches the page again as GPTBot and ClaudeBot.
Is a canonical tag a directive?
No. Google calls a canonical preference a hint, not a rule, and may choose a different URL as canonical. It treats rel=canonical as a strong signal, a redirect as a strong signal and a sitemap as a weak one, and says the methods stack when they agree. Search Console's URL Inspection tool shows which URL Google actually chose.
Should every page have a self-referencing canonical?
Google recommends it: include a rel=canonical link on the canonical page itself. It costs one tag, and it means that tracking-parameter, http, www or trailing-slash variants of the URL all point back to the version you want, even when you did not anticipate them.
Can I use noindex and a canonical on the same page?
Not with a canonical that points somewhere else. Google's John Mueller called that pair "very contradictory pieces of information" and said Google will generally pick the rel=canonical. Use a canonical for duplicates you want consolidated, noindex for pages that should stay out of search, and a self-referencing canonical if a noindexed page needs one at all.
Should a canonical URL be absolute or relative?
Absolute. Google asks for absolute URLs in rel=canonical. A relative canonical resolves against whatever host served the page, so a staging copy, an http variant or a scraped mirror ends up naming itself as canonical.
What does "Duplicate, Google chose different canonical than user" mean?
Search Console shows it when your page is marked as canonical for a set of pages but Google thinks another URL makes a better canonical, and has indexed that one instead. Check which URL Google picked with URL Inspection, then make your signals agree: internal links, redirects, sitemap entries and the canonical tag should all name the same URL.
Do AI search engines use canonical tags?
Bing says clear canonical tags help search engines and AI systems surface the correct page, and that AI systems group near-duplicate URLs and pick one to represent the set. Google's AI Overviews and AI Mode only link to indexed pages, and Google usually indexes the canonical. OpenAI and Perplexity do not mention canonical tags in their crawler documentation, so how their crawlers treat them is not documented.
