How to find and fix keyword cannibalisation
Keyword cannibalisation is when two or more pages on your own site compete for the same search query, and Google keeps swapping which one it shows. The symptom I see most often is a page that ranks at position 6 one week, drops to 14 the next, and is replaced by a different URL from the same site.
This is for anyone running a site with more than a few dozen pages, especially sites that have published steadily for a year or more. Affiliate sites, blogs that cover one niche from many angles, and ecommerce sites with overlapping category pages get hit hardest. I run several content sites out of Singapore and I find this problem on nearly every one after about 100 pages.
By the end you will have a list of queries where your own pages collide, a decision for each (merge, redirect, differentiate or leave alone), and a way to check the fix worked.
What you need
- google search console: access to the property, verified. free.
- a spreadsheet tool: google sheets or excel is fine. for bigger exports, python with pandas.
- a crawler: screaming frog seo spider is free up to 500 URLs, which covers most small sites. any crawler that exports titles and h1s works.
- 16 months of search console data if the property has it. that is all it keeps, so start exporting now.
- edit access to your cms and your redirect config (htaccess, nginx, cloudflare rules or a redirect plugin).
- optional: a rank tracker such as ahrefs, semrush or a self-hosted one, if you want daily position data instead of search console’s averages.
Step by step
1. Export query and page data from search console
Open Search Console, go to Performance, then Search results. Set the date range to the last 3 months. Turn on clicks, impressions and position. You need query and page together, and the web interface will not give you that as a single table by default, so for a whole-site view use the API.
Google documents the Search Analytics query method, which accepts dimensions: ["query", "page"] and returns up to 25,000 rows per request. The Looker Studio connector or a Google Sheets add-on also works, as long as you get one row per query and page pair.
Expected output: a CSV with columns query, page, clicks, impressions, ctr, position. On a few-hundred-page site this is usually tens of thousands of rows.
If it breaks: the export is truncated at 1,000 rows. That is the interface limit. Use the API or the Looker Studio connector, or export per-page and stitch the files.
2. Find queries with more than one ranking URL
Now group by query and count distinct pages. Anything with two or more pages that each have real impressions is a candidate. Here is the script I use.
import pandas as pd
df = pd.read_csv("gsc_query_page.csv")
# ignore tiny impression counts, they are noise
df = df[df["impressions"] >= 20]
grp = df.groupby("query").agg(
pages=("page", "nunique"),
total_impr=("impressions", "sum"),
total_clicks=("clicks", "sum"),
)
cand = grp[grp["pages"] >= 2].sort_values("total_impr", ascending=False)
cand.to_csv("cannibal_candidates.csv")
print(cand.head(30))
Expected output: a ranked list of queries with the number of competing pages and total impressions.
If it breaks: you get thousands of candidates. That is normal. Raise the impression threshold to 100 and only review the top 50 queries. You are not trying to fix everything, just the collisions that cost you traffic.
3. Remove the false positives
Not every multi-page query is cannibalisation. Three common cases to drop:
- branded queries where your homepage and an about page both appear. harmless.
- queries where one page has 95 percent of impressions and the second has a handful. that is a stray, not a fight.
- sitelinks. Google sometimes shows several of your URLs under one result, and search console can count each as a separate page for the same query.
For the rest, add a column for each page’s share of the query’s impressions. A genuine conflict usually looks like a 60/40 or 50/50 split, or a split that changes over time. To see the change over time, run the same export for two different months and compare which page led.
Expected output: a shortlist of 10 to 30 queries where two pages each hold a meaningful share.
If it breaks: you cannot tell whether a split is real. Search the query yourself in an incognito window with the location set to your target country, and see which of your URLs shows up.
4. Read both pages and diagnose intent
Open each competing pair side by side. Ask what the searcher wants, and what each page is actually for. I sort pairs into four types:
- same intent, same content: two articles that say roughly the same thing. typical after years of publishing “best X” and “top X” posts.
- same intent, one is thin: a short old post and a newer long one. the thin one is leaking rankings.
- different intent, same keyword in the title: a guide and a product page, for example. Google is confused because both titles use the phrase.
- a hub and its child: a category or pillar page competing with a detailed article. often fine, sometimes needs a rewrite.
Pair types 1 and 2 need merging. Type 3 needs differentiating. Type 4 needs a look at internal links. I wrote about how clusters should be structured in topic clusters on a site with twelve articles, and the same logic applies when you have a hundred.
Expected output: a label next to every pair in your sheet.
If it breaks: you cannot tell whether the intent is the same. Look at what ranks on page one for the query. If Google shows all guides, the intent is informational and a product page has no business competing.
5. Pick a winner per pair
For merges, choose the URL to keep. I pick by, in order: most backlinks from other sites, most clicks over the last 3 months, and the best existing URL (short, clean, no dates in it). If an old page has real links pointing at it, keep that one even if the newer article is better written, and move the better content into it.
Check backlinks in Search Console under Links, or in Ahrefs or Semrush if you have them. The mechanics are covered in what a redirect does to a link.
Expected output: a column marking keep and retire for every merge pair.
If it breaks: both pages have strong, different backlink profiles. Keep the one with more referring domains and accept a small loss on the other. A 301 passes most of the value.
6. Merge the content and redirect the loser
Take anything useful from the retiring page (a section, a table, a clean example) and fold it into the keeper. Do not paste the whole thing in. The merged page should be better than either original, not longer for its own sake. If the keeper is also dated, this is a good moment to run a proper refresh, which I cover in how to do an SEO content refresh that actually works.
Then 301 the retired URL to the keeper. On nginx:
location = /blog/old-cannibal-post {
return 301 /blog/keeper-post;
}
Google explains how it treats permanent redirects in its site move documentation. Also update internal links so they point straight at the keeper, not through the redirect.
Expected output: the old URL returns a 301 and lands on the keeper with a 200.
If it breaks: you see a redirect chain or loop. Test with curl -I -L https://yoursite.com/blog/old-cannibal-post and count the hops. You want one 301 then a 200. Fix any earlier redirect rules that point at the old URL.
7. Differentiate pages that should both exist
For type 3 and 4 pairs, both pages stay. The fix is to make each one clearly about something different.
- rewrite the title tag and h1 so only one page targets the head phrase. the other gets a longer, more specific phrase.
- change the internal anchor text. if ten articles link to the wrong page with the head keyword as anchor, move those to the right page.
- trim overlapping sections on the secondary page and link to the primary one instead of repeating it.
- where two near-identical pages must remain (a print version, a tracking variant), use a canonical tag. canonical tags explained covers when that works and when Google ignores it. Google’s own guidance on consolidating duplicate URLs is the source I trust on the details.
Expected output: each page has a distinct primary query, and internal anchors agree with that.
If it breaks: you rewrote the titles and nothing moves after a month. Check the internal links and the body copy. Titles alone rarely settle it if the body and the anchors still point at the same phrase.
8. Request reindexing and measure
After the changes go live, open URL Inspection in Search Console for the keeper and click Request indexing. Write down the date and the baseline numbers for the query (clicks, impressions, average position).
Check again at 4 weeks and 8 weeks. Filter performance to the query and look at the page tab. You want one dominant URL, a steadier position line, and ideally a higher click count. Other things change at the same time, so I treat the result as a direction, not proof. I cannot promise a traffic lift on any given site, and on some queries nothing changes.
Expected output: the query’s impressions concentrate on the keeper, and the retired URL fades out of the report.
If it breaks: the retired URL still shows impressions after 8 weeks. Confirm the 301 is live, that your sitemap no longer lists it, and that no internal links still point to it.
Common pitfalls
- treating every multi-page query as a problem. plenty of queries legitimately show two of your pages and that is fine, even good.
- deleting the loser without a redirect. you throw away whatever links and history it had and get a 404 in return.
- merging by word count. pasting two 1,500 word articles into one 3,000 word page does not help. cut duplicated sections.
- fixing titles only. if internal anchors and body copy still push the same phrase, the conflict stays.
- judging the result at week one. Google needs time to recrawl, and average position in Search Console swings while it does.
Scaling this
At 10x, meaning a site of a few hundred pages, everything above works in a spreadsheet. You will find 10 to 30 real conflicts and can fix them in a day or two.
At 100x, around a few thousand pages, the manual reading stops scaling. Automate the candidate list with the pandas script, add a similarity score on titles and h1s from your crawl, and only read the pairs that score high on both. Keep the export on a schedule, monthly is enough, and append to a history file so you are not limited by the 16 month window. Programmatic sites should prevent the problem in templates, by giving each page type a single target query pattern.
At 1000x, tens of thousands of URLs, you are working in a database. Load the Search Console API output into BigQuery or Postgres, run the grouping in SQL, and set thresholds that open a ticket only when a query crosses a click loss or a split ratio. Humans review only the top few dozen conflicts by lost clicks. For the scripting side of this, the tooling roundups on aitoolgazette.com are a useful place to look at what AI assistance can do with clustering titles, though I would still check any merge decision by hand.
Where to go next
- canonical tags explained: for pages that must stay live but should not compete
- how to do an SEO content refresh that actually works: to make the merged keeper worth ranking
- what a redirect does to a link: before you retire anything with backlinks
The full article index is at /blog/.
Written by Xavier Fok
disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-10-01.