← all articles

Programmatic SEO pages not indexed: where the tactic actually breaks

The pitch versus what happens on page 40,000

Programmatic SEO gets sold as a numbers game. Build one template, feed it a spreadsheet, generate ten thousand pages, and let search traffic compound. The first few hundred pages usually behave. They get crawled, they get indexed, some of them even rank. Then somewhere past the first few thousand, the curve bends. New pages stop showing up in Search Console as indexed. They sit in “crawled, currently not indexed” or “discovered, currently not indexed” and stay there for weeks.

I’ve run templated page sets across several sites in my own portfolio, not one experiment but the same pattern repeating on different domains with different templates. The failure point isn’t random and it isn’t a penalty. It’s a resource allocation problem on Google’s end colliding with a content quality problem on mine, and once you’ve seen it happen twice you stop being surprised and start designing around it.

What “not indexed” actually means

Search Console gives you a handful of statuses that matter here, and they mean different things:

  • Discovered, currently not indexed: Google knows the URL exists (from your sitemap or internal links) but hasn’t crawled it yet, or crawled it and decided not to prioritize indexing right now.
  • Crawled, currently not indexed: Google fetched the page and chose not to add it to the index. This is the one that should worry you, because it means the page was evaluated and rejected, not just queued.
  • Duplicate, Google chose different canonical: your page exists, but Google folded it into another URL it considers more authoritative for the same content.

The distinction matters because the fix is different for each. A discovery backlog is a crawl budget and internal linking problem. A crawled-but-rejected page is a content problem. A duplicate-canonical issue is a template problem where your pages aren’t differentiated enough from each other to be treated as separate documents.

Programmatic sites tend to hit all three at once, because the thing that makes programmatic pages cheap to produce (one template, many data rows) is the same thing that makes them look repetitive to a crawler evaluating thousands of near-identical URLs.

Crawl budget is real and it’s not a myth to wave away

Google has published documentation on crawl budget for a reason: sites with very large URL counts and low per-URL value get crawled less thoroughly, and crawl activity gets allocated toward pages the site itself signals as important through internal links, sitemaps, and update frequency. This isn’t a secret ranking signal, it’s operational plumbing that Google has stated outright. If you publish 50,000 pages in a week and your site has never handled that volume before, you’re asking the crawler to make 50,000 individual value judgments faster than it’s used to making them for you. Some fraction will get deferred. That’s not a mystery, that’s throughput.

What I’ve watched happen in practice: submitting a sitemap with tens of thousands of new URLs doesn’t force indexing, it just tells Google the URLs exist. The crawler still has to visit, render, and evaluate each one, and it does that at a pace set by how much it trusts the domain and how much server response time and content quality it’s found on similar URLs so far. A domain with a thin crawl history for templated content gets fewer evaluations per day, not more, no matter how big the sitemap is.

The content problem hiding under the template

Here’s the part that’s harder to admit: a lot of “not indexed” programmatic pages aren’t a technical problem at all. They’re pages where the only thing that changes between URL A and URL B is a city name, a product SKU, or a number swapped into a sentence template. If you strip the swapped variable out, the pages are the same page.

Google doesn’t need an algorithm secret to catch this. Near-duplicate detection based on content similarity is a basic, publicly discussed part of how search engines have worked for over a decade. A template that says “Find the best [X] in [City]” with three sentences of boilerplate around it and no city-specific data is trivially similar across a thousand rows. You built one page. You just gave it a thousand URLs.

The programmatic sets that keep their indexing rate up are the ones where the data row actually changes what’s on the page: real local pricing pulled from an API, a unique dataset, actual differing inventory, distinct user-submitted content, or computed output specific to the input. The ones that fall off are the ones where the “programmatic” part is just find-and-replace on a paragraph.

What I’ve actually changed after watching this happen

I’m not going to pretend I fixed this with a trick. What worked was less generation, not more:

Cut the template count before you cut anything else. If ten thousand pages are getting maybe two thousand indexed, publishing five thousand of the highest-data-density pages usually outperforms publishing all ten thousand thin ones, because you stop diluting the domain’s average page quality signal that crawl allocation seems to respond to.

Consolidate rows that don’t have enough unique data to justify a standalone URL. If a city has no listings, no price data, no distinct facts, it doesn’t get a page. It gets folded into a parent category page instead. A page with nothing on it but a template shell is a page that invites a “crawled, not indexed” result.

Canonicalize deliberately instead of hoping Google picks the right one. If two URLs really are near-duplicates because the underlying data is genuinely identical (same product, two URL formats), pick one and canonical the other rather than letting Google guess, because when Google picks for you it sometimes picks the one you didn’t want ranking.

Slow the publish rate on new domains. A domain with no crawl history doesn’t get the benefit of the doubt a domain with two years of indexed content gets. Publishing in batches and watching the indexed ratio in Search Console before publishing the next batch costs time, but it means you’re not burning crawl trust on a launch spike Google isn’t ready to absorb.

Prune dead weight on a schedule. Pages that have sat unindexed for months with zero impressions aren’t going to suddenly get picked up because you left them alone. Removing them, or merging them into a page that does have data, is a normal maintenance task, not an admission of failure.

Where the tactic genuinely stops working

There’s a ceiling on programmatic SEO, and it’s not a secret Google threshold, it’s a content economics problem. If the marginal page in your data set doesn’t add marginal information a reader or a search engine can tell apart from the page next to it, no amount of technical cleanup fixes that. Fixing sitemaps, internal linking, and crawl pacing gets you the indexing your content quality actually earns. It doesn’t manufacture indexing for pages that are functionally duplicates wearing different URLs. I’ve had to walk away from data sets that looked great in a spreadsheet and just didn’t have enough real variance per row to support a page each, and the honest move was to merge them down rather than keep publishing and hoping the next batch behaves differently.

If you’re staring at a Search Console report full of “discovered, not indexed” and wondering whether to publish more or publish less, the answer is almost always less, aimed at the rows with the most actual data behind them.

If you want more of this kind of breakdown, unglamorous and based on running actual sites rather than theory, check out the rest of The SEO Desk.

for SEOs
Tracking rankings or scraping SERPs at scale?

Rank checkers and SERP crawlers get blocked and geo-skewed fast on datacenter IPs. Singapore Mobile Proxy runs real 4G/5G mobile IPs that search engines still trust, so your position data stays clean.

see plans →
read on
More from The SEO Desk

Technical SEO, link building, content and SERP strategy, and tool reviews for people who ship growth.

browse all articles →