Canonical tags and the version Google picks
What a canonical tag actually tells Google
A rel=canonical tag sits in the head of a page and points to the URL you want treated as the “real” version when duplicate or near-duplicate content exists. You put it on print versions, filtered category pages, session-ID URLs, http and https pairs, www and non-www pairs, syndicated content, and pretty much anywhere the same content is reachable through more than one URL.
The tag itself is one line: <link rel="canonical" href="https://example.com/product" />. Simple to write, easy to get wrong, and it does not do what most people assume it does.
I run a portfolio of content sites and I’ve had canonical tags quietly redirect crawl budget and indexing to the wrong URL more times than I want to admit. The tag is a suggestion. Google treats it as one signal among several, and it will override your suggestion when other signals disagree with it strongly enough.
Rel=canonical is a hint, not a rule
This is the part that trips people up. A canonical tag is not a redirect and it is not a directive. It’s a request. Google calls the URLs on your canonical tags “user-declared canonicals,” and it runs them through its own duplicate-content clustering process before deciding what actually shows up in the index. The output of that process is what Google calls the “selected canonical,” and it can be a different URL than the one you declared.
This matters because a lot of SEO advice talks about canonical tags like a light switch: set it and the duplicate problem is solved. It isn’t a switch. It’s an input into a decision Google makes on its own, based on signals that have nothing to do with the tag itself.
The signals that override your tag
I don’t have access to Google’s internal weighting and I’m not going to pretend I do. What’s documented and observable from working on real sites is that Google looks at more than the tag when clustering duplicate URLs and choosing which one to show:
Redirects. If URL A 301-redirects to URL B, that’s a much stronger signal than a canonical tag pointing the same direction, because a redirect changes what’s actually served. If you canonicalize A to B but users and bots can still load A directly with a 200 status, you’ve sent a weaker signal than a redirect would.
Internal linking patterns. If your site links to the non-canonical version from navigation, sitemaps, and internal content far more than it links to the canonical version, Google is looking at a site that behaves as if the non-canonical URL is the real one. The tag says one thing, your link graph says another.
XML sitemap inclusion. Sitemaps are supposed to list canonical URLs only. If your sitemap has the duplicate URL in it and the canonical tag points somewhere else, that’s a contradiction Google has to resolve, and it doesn’t always resolve it the way you want.
HTTPS and www consistency. These are supposed to be settled by redirects at the server level, not by canonical tags. If you’re relying on a canonical tag to tell Google that the https version is the “real” one while the http version still serves 200 responses, fix that with a redirect instead. Canonical tags are not a substitute for that.
Content similarity. Google’s clustering is based on comparing actual page content, not just reading your tag. If two pages are similar enough to cluster but not identical, and you point the canonical the “wrong” way relative to which page has more unique content, more internal links, or more external signals, Google can pick the other one anyway.
None of this means the tag is useless. It means it’s one vote in a process that also counts redirects, links, and sitemaps as votes, and those other votes can outweigh it.
How to see which version Google actually picked
You don’t have to guess. Search Console’s URL Inspection tool shows both the “user-declared canonical” (what your tag says) and the “Google-selected canonical” (what Google actually chose to index). When these two rows disagree, that’s your signal that something in your setup is fighting the tag.
This is the single most useful check for canonical problems and it’s underused. People set a tag, assume it worked, and move on. I’ve found pages on my own sites where the declared and selected canonical had been different for months because a sitemap generator was still listing the old URL, or an internal linking widget was linking to a parameter version out of habit.
Check this on any URL where duplication is a real risk: paginated series, faceted or filtered category pages, anything with tracking parameters, and any page that exists at more than one path. If the tool says Google picked something other than what you declared, go look at your redirects, your sitemap, and your internal links before you touch the tag again, because the tag probably isn’t the problem.
Where this breaks in practice
A few patterns show up constantly on real sites:
Canonicalizing to a noindexed or blocked page. If page A canonicalizes to page B, and page B is noindexed or blocked by robots.txt, you’ve told Google to treat A as a duplicate of a page it’s not allowed to index. The practical result is often that neither page gets indexed cleanly. Canonical targets need to be indexable pages, full stop.
Cross-domain canonicals on syndicated content. If you publish the same article on your own site and let a partner site run it too, a cross-domain canonical tag pointing back to your original is the correct move, and it’s respected the same way as a same-domain canonical: as a hint, weighed against other signals. It doesn’t guarantee your version outranks the syndicated copy. It just tells Google which one you consider the source.
Self-referencing canonicals as a default. Most CMS platforms now output a self-referencing canonical on every page by default, meaning the page points to itself. That’s the right baseline for pages that aren’t duplicates of anything. The problem shows up when a templated system applies that same self-referencing logic to parameter URLs or paginated pages that should be pointing elsewhere, and nobody overrides the default.
Pagination. Page 2 of a paginated series should not canonicalize to page 1. They’re different content. Google has said as much for years, but I still see site owners do this because they read somewhere that canonical tags “consolidate ranking signals,” and they want page 1 to get the benefit. It doesn’t work that way, and it risks page 2’s actual unique content getting dropped from the index entirely because you told Google it’s a duplicate of something it isn’t.
Canonical chains and other self-inflicted wounds
A canonical chain happens when page A canonicalizes to page B, and page B canonicalizes to page C. This usually happens by accident, often after a site migration where old canonical tags never got updated to point directly at the final URL. Google can follow these chains, but it adds ambiguity to a process that already involves clustering and comparison, and I’d rather not add ambiguity to something I don’t fully control. Every canonical tag should point directly at the final destination URL, not at another URL that itself canonicalizes elsewhere.
The other common self-inflicted wound is parameter handling done inconsistently across a site. If your ecommerce platform generates ?color=red and ?sort=price variants of the same category page, and half of those pages have a correct canonical tag while the other half don’t because a template was updated in one place but not another, you end up with a mixed signal at scale. This is worth auditing with a crawler across the whole site rather than spot-checking a handful of URLs, because the inconsistency itself is the problem, not any single tag.
What to actually do about it
Treat the canonical tag as one part of a consistent story you’re telling Google, not the whole story. Fix redirects at the server level first. Make sure your sitemap only lists the URLs you actually want indexed. Check that internal links point at the canonical version, not the duplicate. Then set the tag, and use URL Inspection to confirm Google agrees with you rather than assuming it does.
None of this guarantees a specific outcome. Google’s clustering process isn’t something outsiders can fully see into, and anyone telling you they know exactly how it weighs each signal is guessing same as the rest of us. What you can control is whether your redirects, sitemap, internal links, and canonical tags all point the same direction. When they do, you’ve given Google a clean signal to work with. When they don’t, the tag alone won’t fix it.
If you want more breakdowns like this on how technical SEO actually behaves in practice, head back to The SEO Desk for the rest of our guides.