Faceted Navigation and the Pages That Multiply
A catalog of 412 that crawled like millions
Last year I looked at a shop with 412 products. Its crawl report came back with just over 2.3 million URLs. The person running the site assumed the crawler was broken, which was a reasonable first guess, because nothing about the site looked like it contained two million of anything.
The crawler was fine. The site had six filters on its category pages, each one appending a parameter to the URL, and every combination of those parameters was a link a crawler could follow. Colour, size, brand, price band, material, in stock. 412 products had quietly turned into millions of addresses, and the crawler had been working through them patiently for months.
This is worth understanding properly, because it’s one of the few technical SEO problems that gets worse on its own while you do nothing, and the fix involves a decision most people skip straight past.
What a facet actually is
A facet is a filter that narrows a listing. The important part is what happens to the URL when someone clicks one.
On most platforms, the filter appends a query string. /category/shoes becomes /category/shoes?colour=blue. Click a second filter and you get ?colour=blue&size=9. Each of those is a distinct address as far as anything crawling your site is concerned, and if the filter is rendered as an anchor tag, it’s also a link that gets followed.
That last detail is the whole problem in one sentence. A filter that’s a link is an instruction to go and look at something. A filter that changes what’s on screen without changing the address is not.
The arithmetic nobody does
Take six filters. Colour has 12 options, size has 10, brand has 20, price is split into 6 bands, material into 8, and there’s a stock toggle with 2 states. If any of those can be selected alongside any other, the number of distinct URLs on a single category page is all of those numbers multiplied together, and the result runs well past a million addresses from one category on a shop where a human would tell you there are a couple of hundred things to look at.
Add a sort dropdown with three options and you triple everything, because every one of those addresses can now exist in three sequences. Add pagination and you multiply again. The shop with 412 products wasn’t unusual, and nobody involved had done anything stupid. Six ordinary filters and a sort dropdown is genuinely all it takes.
Why it costs you something real
I want to be careful here, because there’s a lot of confident writing about how search engines allocate crawling internally, and I don’t have that information and neither does anyone selling you a course on it.
What I can tell you is what’s observable from the outside, in log files and in Search Console. A crawler arrives with some finite amount of effort per visit. When a large share of that effort goes into fetching pages that differ by one colour value, the pages you actually care about get fetched less often. On the shop above, product pages were being recrawled roughly every five weeks. After the filter URLs stopped being followable, that dropped to about four days.
The second cost shows up in the index coverage report: thousands of URLs sitting in “discovered, currently not indexed” or “crawled, currently not indexed,” which makes the report nearly useless for spotting real problems, because the real problems are buried under filter noise.
The third cost is subtler. When several near-identical pages all target the same query, whatever signals point at that query get spread across them instead of landing on one page you actually chose.
The question that decides everything
Before you touch a single directive, work out which filter combinations deserve to exist as a page at all. This is the step people skip, and skipping it is why so many of these projects either block too much or too little.
A facet URL earns its place if three things are true. Someone actually searches for that combination as a phrase. The resulting page is meaningfully different from its parent, not the same dozen items in a different order. And it could stand on its own as a landing page if someone arrived from search knowing nothing about the site.
Blue running shoes passes all three: it’s a real search with real volume, the page is a genuinely different set of products, and a stranger landing on it gets what they expected. Blue running shoes in size 9 under 80 pounds sorted by rating passes none of them.
What to do with the filters that pass
Promote them. Give the combination a clean static URL rather than a parameter string, something like /shoes/blue instead of /shoes?colour=blue.
Then treat it like the landing page it now is: its own title, its own short piece of copy above or below the grid, and its own place in the internal linking so it isn’t only reachable by clicking a filter. Most sites end up with somewhere between 5 and 40 of these, which is a very manageable number of real pages.
The tools that don’t work the way you’d think
Everything else, which is nearly all of them, gets handled with one of a few mechanisms, and each has a catch worth knowing before you pick one.
Robots.txt disallow stops the crawling. That’s the point of it. The catch is that a blocked URL can still end up in the index as a bare address with no description, because blocking crawling isn’t the same as forbidding indexing, and a crawler that can’t fetch a page also can’t see any noindex tag on it.
Noindex does forbid indexing, and it works reliably. The catch is that the page has to be fetched for the tag to be read, so it does nothing for the crawling problem itself. It’s the right tool for cleaning up URLs that are already in the index, and the wrong tool for stopping millions from being discovered in the first place.
Canonical tags are treated as a strong hint rather than an instruction, and they’re most likely to be respected when the two pages really are close to identical. A filtered page showing 11 of 200 products isn’t close to identical to its parent, so a canonical pointing back is asking for something the search engine may well decline.
Nofollow on filter links reduces the chance of a URL being discovered through that specific link. It does nothing once the address exists somewhere else, and addresses leak into sitemaps, internal search, and other people’s links.
The mechanism that actually solves it
Stop making the filters links.
If a filter updates the listing without producing a followable anchor tag pointing at a new address, there’s nothing for a crawler to find. That’s usually a small front-end change rather than an SEO project, and it removes the problem at the source instead of managing it afterward.
On most modern platforms this means the filter fires a request and updates the grid, and the address either doesn’t change or changes in a way that isn’t a link anyone can follow. The combinations you promoted earlier stay as real links, because you want those crawled.
The order to work in is: turn the valuable combinations into proper pages, make everything else stop being a link, use noindex to clean up whatever’s already indexed, and only reach for robots.txt if you have an active crawling emergency and need to stop the bleeding today.
Sort order and tracking parameters
Sort order never deserves its own page. Same items, same query, different sequence. Every sort parameter should never appear in a followable link, and there’s no interesting judgement call to make there.
Tracking parameters ride along with this and are worth checking while you’re in there. Campaign tags appended to internal links create a second address for a page that already exists, and internal links should never carry them. Use tracking parameters on inbound links from outside the site, and strip them everywhere internal.
Pagination is different and gets treated differently. Page two of a category is a real, distinct set of products and generally should be crawlable. It doesn’t belong in the same bucket as filters.
How you know it worked
Give it time and watch three numbers.
Server logs tell you what proportion of crawler requests are hitting URLs with parameters in them. That’s the number that should fall, and it’s the most direct evidence you have. Search Console’s crawl stats should show total fetches per day staying roughly level while the composition shifts toward pages you care about. And in the index coverage report, “discovered but not indexed” should shrink over weeks.
It isn’t fast. The addresses already exist and they’ll keep being revisited for a while after you stop linking to them. On a large site, it can take two to three months for the logs to settle. That’s normal and not a sign you got it wrong.
The mistake that makes things worse
The tempting move when you first see two million URLs is to block every parameter in robots.txt on a Friday afternoon.
The problem is that some of those filter pages have been quietly earning traffic for years, and nobody knows which ones until they’re gone. I’ve watched a site lose a meaningful chunk of revenue this way and spend six weeks working out what it had switched off.
Pull the data first. Export every page that got any impressions at all in the last year, find the ones with parameters, and check them against the three-question test before anything gets blocked. That export takes ten minutes, and it’s the difference between a clean project and an incident.
The honest limit
Fixing this doesn’t lift your rankings by itself. It isn’t a growth lever, and I’d be lying if I said the traffic line jumps afterward.
What it does is stop you wasting crawl effort on pages that were never going to earn anything, make your coverage reports readable again, and remove a source of duplication that was working against you quietly. On a site with a few hundred pages, none of this matters much. On a shop with a large catalog and filters on every category, it’s often the single biggest technical fix available.
If your crawl report looks wrong
Count the URLs and compare that to how many things you actually sell, because the gap is the whole diagnosis. Do the multiplication on your own filters so you know what you’re dealing with. Export a year of impressions before you block anything. Pick the handful of combinations people genuinely search for and turn those into real static pages. Make every other filter stop producing a followable link. Then clean up the existing index with noindex and leave it alone for a couple of months.
The shop I opened with is down to about 900 crawlable URLs now, and roughly 40 of those are filter combinations that earn their keep.
If you want more of this kind of walkthrough, the technical breakdowns and the tool tests from the sites we actually run are on The SEO Desk.