How to fix crawl budget problems on a big site
I once watched Googlebot spend 40% of its crawl requests on a single site hitting ?sort=price&color=&size= combinations that produced zero unique content. Meanwhile the actual product pages, the ones that made money, went weeks between visits. That’s a crawl budget problem, and if you’re running a site past 50,000 URLs, you’ve probably got one even if nothing looks obviously broken.
This is for operators running big sites: ecommerce catalogs, marketplaces, programmatic content sites, anything where the URL count runs into the tens or hundreds of thousands. If your site has 500 pages, skip this one. Google is explicit that crawl budget mostly matters once you’re past roughly a million URLs, or you’re publishing fast enough that freshness matters more than usual, per Google’s guide on managing crawl budget for large sites. Below that scale, indexing problems are almost never a crawl budget issue, they’re a content or internal linking issue.
Fix this properly and the outcome is straightforward. Googlebot spends its visits on pages you actually want ranked, instead of parameter soup, redirect chains, and dead URLs. I’ve seen indexed counts for the pages that matter jump 20-30% within two months of a proper cleanup, with zero new content published. Nothing else about the site changed except where the crawler was allowed to go.
what you need
- Search Console verified at the domain property level, not just URL-prefix, so you get the full picture instead of one subdomain
- raw server or CDN access logs, minimum 30 days, ideally 90. Cloudflare Logpush, Fastly real-time log streaming, or your host’s raw logs all work
- a log analyzer. I use Screaming Frog’s Log File Analyser, $99/year for one license. The free tier caps at 1,000 log lines, which is useless past a blog-sized site
- Screaming Frog SEO Spider ($259/year) or a comparable crawler, for finding redirect chains and soft 404s
- admin access to robots.txt and whatever generates your XML sitemaps
- someone who can ship a redirect rule or a code change, since a chunk of these fixes aren’t things you can do from Search Console alone
step by step
1. pull your crawl stats baseline
Open Search Console, go to Settings, then Crawl stats. Look at total crawl requests over 90 days, average response time, and the breakdown by response code and file type.
Expected output: a chart of daily crawl requests plus tables breaking down what’s actually being hit. Search Console’s own crawl data beats any third-party estimate here, since it’s Google telling you exactly what Google did. I’ve gone through the tradeoffs between Search Console and outside rank trackers before in Search Console vs third-party rank trackers if you want the fuller comparison.
If it breaks: the crawl stats report is sparse or empty. That usually means you verified a URL-prefix property instead of the domain property. Re-verify at the domain level via a DNS TXT record and wait 24-48 hours for data to populate.
2. get the real crawl data from your logs
Search Console samples and aggregates. Your logs don’t. Pull the last 30 days of raw access logs and filter for Googlebot:
grep -i "Googlebot" access.log | awk '{print $7}' | sort | uniq -c | sort -rn | head -50
Verify the hits are really Googlebot with a reverse DNS lookup, not just the user agent string, since that’s trivially spoofable. Load the filtered log into your analyzer.
Expected output: a ranked list of every URL Googlebot actually hit, how many times, and what status code came back.
If it breaks: your log retention is shorter than 30 days and you’ve already lost the window. Set up log push to cold storage today and come back to this step in a month. There’s no shortcut, you need real data.
3. find the crawl traps
Sort the log data by URL pattern and look for parameter combinations multiplying out of control: faceted navigation, session IDs in the URL, infinite date-based calendar pages, internal search results. This is where programmatic and ecommerce sites bleed the most crawl budget, and it’s the exact failure mode I wrote about in programmatic pages and where they stop working: a page template that works fine at 500 URLs generates garbage at 500,000.
Expected output: you’ll usually find 30-60% of crawl requests going to a handful of patterns resolving to duplicate or near-empty content.
If it breaks: you can’t tell which parameters are load-bearing. Cross-check against the “duplicate without user-selected canonical” bucket in Search Console’s page indexing report before you touch anything.
4. fix redirect chains and soft 404s
Crawl the site with a spider tool and filter for redirect chains of three or more hops, and soft 404s, pages returning a 200 while the content says “not found” or “out of stock, nothing here.”
Expected output: a list of URLs to repoint directly to their final destination, and a list of dead pages that should return a real 404 or 410 instead of a fake 200.
If it breaks: the redirects trace back to a plugin or middleware layer nobody remembers installing. Fix it at the source, the CMS routing table or the CDN edge rule, rather than stacking another redirect on top. Otherwise you’re back here in six months with a four-hop chain instead of three.
5. block or noindex the low value patterns
For pure crawl-trap parameters with zero unique value, session IDs, sort orders that don’t change content, disallow them in robots.txt:
User-agent: *
Disallow: /*?sort=
Disallow: /*?sessionid=
For parameters that create thin duplicates but still serve a real purpose, color filters, pagination, use canonical tags pointing to the clean URL instead of blocking outright.
Expected output: crawl requests to those patterns drop within one to two weeks, once Googlebot re-reads robots.txt on its normal schedule, per Google’s own robots.txt documentation.
If it breaks: you disallow a pattern that also happens to serve canonical, indexable content, and now those pages can’t be recrawled at all. Check URL Inspection for a sample of affected URLs before pushing the robots.txt change live, not after.
6. rebuild the XML sitemaps
Split sitemaps by section, keep each file under 50,000 URLs and 50MB uncompressed per the sitemaps.org protocol, and only include canonical, indexable, 200-status URLs. Strip out anything you just blocked or noindexed in step 5.
Expected output: a sitemap index referencing clean sub-sitemaps, resubmitted through Search Console, where “discovered URLs” roughly matches your real live count instead of running 3x over it.
If it breaks: the “last read” date in Search Console doesn’t update right after resubmission. That’s normal, it can take several days. Resubmitting repeatedly doesn’t speed it up.
7. speed up server response time
Check average response time in the crawl stats report. Google ties crawl rate directly to how fast your server responds. Consistently over 500ms-1s throttles how much Google is willing to crawl, independent of anything else you fix. Add edge caching for anything that isn’t truly dynamic.
Expected output: response times drop, and heavier crawl activity typically follows within a few weeks.
If it breaks: the crawl volume you unblocked in step 5 hits a server now getting crawled harder than before, and response times get worse instead of better. Ramp the changes gradually rather than unblocking everything and rebuilding the sitemap in the same week.
8. re-run the log comparison
Repeat step 2 for the 30 days after your changes and compare crawl share before and after.
Expected output: crawl requests moving away from the junk patterns and toward the URLs that actually earn money.
If it breaks: the numbers look flat because your CDN is now caching crawler requests at the edge before they hit origin, undercounting total crawl activity in your origin logs. Pull the CDN’s own log export instead.
9. monitor for six to eight weeks
Check the crawl stats report weekly and watch the indexed count for pages that matter, not total indexed count, which can go down as junk URLs finally drop out.
If it breaks: nothing moves after a month. Go back through steps 3 and 5 and check you didn’t accidentally block something load-bearing, then check the “why pages aren’t indexed” report for a pattern you missed.
common pitfalls
- treating every indexing problem as a crawl budget problem. If the site is under a few hundred thousand URLs and isn’t publishing constantly, a stuck page almost always traces back to thin content or missing internal links, not crawl budget. Chasing crawl budget here wastes weeks.
- blocking in robots.txt something that should have been noindexed instead. A blocked page can still get discovered and shown with no description, and once it’s blocked, Google can’t even see a noindex tag on it to remove it properly.
- fixing the sitemap and stopping there. The sitemap is a hint, not a directive. Internal linking is what actually tells Google what you consider important, so a clean sitemap sitting on top of a messy nav does very little.
- confusing crawled with indexed and reporting the wrong number to whoever’s asking. A URL can get crawled constantly and still sit in “crawled, not indexed” for months because of a quality judgment, not a crawl budget one.
- running the whole cleanup in the same week as a big content push. I did this once on a client site: changed robots.txt, resubmitted sitemaps, and launched 4,000 new pages inside the same ten days. When rankings moved I had no idea which change caused what. If you’re on a publishing cadence that ships new content weekly, stagger the crawl budget work around it, not on top of it.
scaling this
At 10x the baseline here, say 100,000 to 500,000 URLs, most of this stays manageable by hand. You can eyeball log samples, rebuild a sitemap manually once a quarter, and one person owns the whole process on a Friday afternoon.
At 100x, past a million URLs, manual review stops working. You need log ingestion into something queryable, BigQuery, Athena, or a local DuckDB setup, because grepping a day’s logs by hand takes longer than the fix. Sitemap generation has to be programmatic and tied to the CMS directly, not a static file someone remembers to update. This is also where a good chunk of the “crawl budget problem” turns out to be your own programmatic page templates outrunning what your internal linking can support. There’s a handful of AI-flavored log analyzers now claiming to auto-categorize crawler intent at this scale. I haven’t found one worth paying for yet, but aitoolgazette.com/blog/ has been tracking that space if you want to see what’s out there.
At 1000x, tens of millions of URLs, crawl budget stops being a cleanup task and becomes an architecture decision: which templates get priority in internal link structure, which sections get pruned outright, whether the render layer holds up under sustained crawler load. The fix isn’t “clean up the crawl traps” anymore, it’s “cut the URL space down to what people actually convert on.” Either way, the fix stops being technical.
where to go next
If the crawl trap turned out to be your own page templates rather than a technical misconfiguration, read what is programmatic SEO and when to use it before you build the next batch. If you’re doing this cleanup on a site that’s still early, the calculus is different, see the first ninety days of a new site for what matters before you ever hit this scale. For the tools mentioned above plus a few more, the best technical SEO audit tools in 2026 has the fuller list. The rest of these are on the blog.
Written by Xavier Fok
disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-09-14.