Canonicalization is the process of designating the “preferred” URL when multiple URLs serve equivalent or similar content. The <link rel="canonical"> tag is the primary mechanism — placed in the <head> of a page, it tells Google: “this is the definitive URL for this content; consolidate any ranking signals here.”
Canonicalization affects indexation, PageRank flow, and which URL appears in search results. Getting it wrong doesn’t produce error messages — it produces quietly suboptimal rankings.
Why Canonicalization Exists
The same content can be accessible at multiple URLs for many reasons:
https://example.com/page/andhttps://example.com/page(with and without trailing slash)https://www.example.com/page/andhttps://example.com/page/(www vs. non-www)https://example.com/page/andhttp://example.com/page/(HTTPS vs. HTTP)https://example.com/page/?utm_source=newsletter(tracking parameters)https://example.com/page/?ref=sidebar(referral parameters)https://example.com/page/?sort=price(sort/filter parameters)
Without canonicalization, Google may:
- Split link equity across the URLs (weakening all of them)
- Index the “wrong” version (the one without tracking parameters showing up in search results is fine, but the HTTP version showing instead of HTTPS is not)
- Crawl wastefully, spending budget on duplicate URL variants
The Canonical Tag Syntax
The self-referential canonical (every page should include this):
<link rel="canonical" href="https://www.example.com/page/" />
Best practice: every page should have a canonical tag. Even if a page has no known duplicates, a self-canonical makes it explicit which URL is authoritative.
Absolute URLs, not relative: Use https://www.example.com/page/ not /page/. Relative canonical URLs can cause problems when the canonical is inherited across different URL contexts.
Canonical must point to a 200 response: A canonical pointing to a redirected URL or a 404 is invalid. Google will follow redirects from canonicals in some cases but it’s unreliable — point canonicals to the final, live canonical URL.
Canonical must match the preferred domain version: If your preferred URL includes www, your canonicals must include www. If your preferred URL uses HTTPS, your canonicals must use HTTPS.
Self-Canonical vs. Cross-Page Canonical
Self-canonical: The canonical tag points to the page’s own URL. Signals: “this is the canonical version of itself; there are no other preferred versions.” Every page should have this as default.
Cross-page canonical: The canonical tag points to a different URL. Signals: “the content here is a duplicate or near-duplicate; the definitive version is at [other URL].” Used for pagination, parameter variants, syndicated content, and URL consolidation.
Canonical for Filtered and Parameterized URLs
URL parameters are the most common canonicalization use case. For every parameter variant that generates a different URL:
<!-- On https://example.com/products/?sort=price&color=blue -->
<link rel="canonical" href="https://example.com/products/" />
This signals that the sorted/filtered view isn’t independently canonical — the unfiltered URL is the preferred version. Google typically respects this and indexes only the canonical.
Important: if a parameter-based URL serves genuinely different content that warrants independent ranking (e.g., a category page for a product subcategory accessed via parameter), it should be canonical to itself, not to the parent. Canonical is a quality signal, not just a duplicate suppressor.
Canonical in HTTP Headers
For non-HTML resources (PDFs, XML files) or situations where editing the HTML <head> isn’t possible, canonical can be specified in HTTP response headers:
Link: <https://example.com/document.pdf>; rel="canonical"
This is functionally equivalent to the <head> tag but accessible via the HTTP response.
Common Canonical Errors
Canonical pointing to a non-indexable page: If the canonical target has noindex, you’ve created a conflict — you’re saying “this is the canonical” but also “don’t index it.” Google will usually honor the noindex and not index either version.
Canonical chain: Page A canonicals to Page B, which canonicals to Page C. Google will typically follow the chain to Page C, but chains are fragile and a single broken link invalidates the entire chain. Point canonicals directly to the final canonical URL.
Conflicting canonical and robots.txt block: A page with canonical pointing to an uncrawlable URL (blocked by robots.txt). Google can’t crawl the canonical to confirm it’s the right target.
Site migration canonical errors: After a domain migration, old domain URLs may still serve canonical tags pointing to old domain URLs. All canonicals should be updated to new domain URLs immediately on migration.
CMS auto-generating incorrect canonicals: Some CMS platforms generate canonical tags based on templating logic that doesn’t account for URL parameter variants. Verify your CMS’s canonical tag output by crawling parameterized URLs and checking the canonical in the rendered HTML.
Canonical vs. Redirect: When to Use Which
Use canonical when: The alternate URL has user-facing purpose (users arrive there legitimately via links or navigation) but you want search equity consolidated to the preferred URL.
Use redirect when: The alternate URL should not be accessed at all — there’s no user-facing reason for it to exist. Redirect combines the SEO signal with the user experience fix.