← all articles

How to audit and fix thin content at scale

Every site I have run past a few hundred URLs ends up with the same junk drawer. Old tag pages, product pages with two lines of copy, city pages that differ by one word, blog posts written to hit a publishing schedule. None of it is dramatic on its own. Together it dilutes the site, wastes crawl time, and makes your better pages harder to rank.

This tutorial is for operators who own or maintain a site with somewhere between 50 and a few thousand pages and want a repeatable way to find the weak ones and decide what to do with each.

The outcome is a spreadsheet where every URL has a verdict (keep, expand, merge, noindex or delete) and a batch of fixes shipped in order of value. I won’t promise a traffic number, but you will know which pages earn their place.

what you need

  • a crawler: Screaming Frog SEO Spider is what I use. The free version crawls 500 URLs, the paid licence is roughly £200 a year, check the current price on their site
  • google search console access for the property, with at least 6 months of data
  • google analytics or any analytics export, optional but useful for sessions and engagement
  • python 3 with pandas installed (pip install pandas), or a spreadsheet tool if you prefer pivot tables
  • a redirect mechanism: access to your server config, CDN rules, or CMS redirect plugin
  • a spreadsheet for the final decision log
  • about half a day for a 500 page site, more if you have to rewrite a lot

step by step

Step 1: crawl the whole site and export the basics

Run a full crawl from the homepage and also feed in your XML sitemap so orphan pages show up. In Screaming Frog, go to Configuration, Spider, and tick “Crawl all subdomains” only if you want them included. Then export the Internal tab as CSV.

Expected output: a file with one row per URL, including address, status code, indexability, word count and inlinks.

If it breaks: if the crawl stops early, check whether the site blocks your crawler user agent or rate limits. Lower the crawl speed to 2 or 3 threads. If JavaScript renders the main content, switch to JavaScript rendering, otherwise word counts will read as near zero for pages that are actually fine.

Step 2: export search performance per URL

In Search Console, open Performance, set the date range to the last 6 months (or 12 if your content is seasonal), switch to the Pages tab and export. You want clicks and impressions for every URL.

Expected output: a CSV with columns like Top pages, Clicks, Impressions, CTR and Position.

If it breaks: the UI export caps at 1,000 rows. For bigger sites, pull the data through the Search Console API or a Looker Studio connector.

Step 3: merge the two files

Join crawl data to search data by URL. Normalise trailing slashes and protocol first, or half your rows will fail to match.

import pandas as pd

crawl = pd.read_csv("internal_all.csv")
gsc = pd.read_csv("gsc_pages.csv")

crawl["url"] = crawl["Address"].str.rstrip("/").str.lower()
gsc["url"] = gsc["Top pages"].str.rstrip("/").str.lower()

df = crawl.merge(gsc[["url", "Clicks", "Impressions"]], on="url", how="left")
df[["Clicks", "Impressions"]] = df[["Clicks", "Impressions"]].fillna(0)

df = df[df["Status Code"] == 200]
df = df[df["Indexability"] == "Indexable"]
df.to_csv("audit_base.csv", index=False)
print(len(df), "indexable 200 pages")

Expected output: one clean table of indexable pages with word count, inlinks, clicks and impressions side by side.

If it breaks: if the merge returns almost no matches, print five URLs from each file and compare them by eye. It is nearly always www versus non-www, or http versus https.

Step 4: flag candidates with a simple score

Word count alone is a bad definition of thin. A 150 word page that answers “what time does the store close” is fine. A 900 word page that says nothing is not. So I flag on two signals together: low word count for the page type, and low demand.

df["thin_words"] = df["Word Count"] < 300
df["no_demand"] = (df["Clicks"] == 0) & (df["Impressions"] < 50)
df["few_links"] = df["Inlinks"] < 3

df["flag_score"] = df[["thin_words", "no_demand", "few_links"]].sum(axis=1)
candidates = df[df["flag_score"] >= 2].sort_values("Impressions", ascending=False)
candidates.to_csv("thin_candidates.csv", index=False)

The thresholds are mine, not a standard. Adjust them to your site.

Expected output: a shortlist, usually 10 to 30 percent of the site on a neglected property.

If it breaks: if the shortlist is 80 percent of your site, your thresholds are wrong for your page types. Split by template (product, category, post) and score each group on its own.

Step 5: read a sample before you decide anything

Open 20 pages from the shortlist and read them. Actually read them. Google’s own guidance on creating helpful, reliable, people-first content asks whether a page provides original information, substantial value compared to other results, and enough that a reader would not need to go elsewhere. Use those questions as your test, not a word count.

Expected output: a rough sense of the failure types on your site. Usually you will see three or four: empty stubs, near duplicates, outdated posts, and pages that are thin on purpose (contact, legal, thank you).

If it breaks: if you cannot tell whether a page is thin, check the query it gets impressions for in Search Console and ask if the page honestly answers that query. If not, it is thin for that purpose.

Step 6: assign a verdict to every candidate

Add a column called action and use only five values.

  • keep: the page is short but complete and has a job (contact, legal, a tight FAQ)
  • expand: it has impressions and a real query, but the content is weak. Rewrite it
  • merge: two or more pages cover the same thing. Fold them into the strongest one
  • noindex: useful to visitors, useless in search (filtered views, internal search results, thank you pages)
  • delete: no demand, no links, no reason to exist

For merge decisions, find the pages competing for one query first. I wrote about how to spot that in two pages competing for the same query. And pick the surviving page by clicks and backlinks, not by which one you like more.

Expected output: every candidate row has one verdict. Nothing is left as “maybe”.

If it breaks: if you have too many “expand” rows to ever finish, that is normal. Rank them by impressions and only commit to the top 20 percent this month. The rest become “noindex” or “delete” until you have capacity.

Before anything is removed, look at which candidate URLs have external links pointing at them. Ahrefs, Semrush or the free Search Console Links report will show it. A thin page with three real links is not a delete, it is a merge with a redirect.

This is also where I check whether the links are worth keeping at all. My notes on what a link is worth on a page with no traffic cover how I judge that.

Expected output: a has_links column, and any “delete” row with links flipped to “merge”.

If it breaks: if your backlink tool is out of credits, use the Search Console Links report instead. It is less complete but free.

Step 8: ship the fixes in batches

Do it in this order, smallest risk first.

  1. noindex the pages that should stay live but out of search. Add <meta name="robots" content="noindex, follow"> in the head, or send an X-Robots-Tag header
  2. merge pages, and 301 redirect each old URL to the surviving one
  3. delete pages with no value and no links, and return a 410 or a 301 to the closest relevant page
  4. rewrite the “expand” pages, and update the internal links pointing at them

For Apache the redirect map looks like this:

Redirect 301 /old-thin-page/ https://example.com/main-guide/
Redirect 301 /another-thin-page/ https://example.com/main-guide/

Redirect to a page that actually matches the topic. Sending everything to the homepage is treated as a soft 404 in practice, and I explain the mechanics in what a redirect does to a link.

Expected output: old URLs return 301 or 410, noindexed URLs return 200 with the directive, and nothing in your sitemap points at a removed page.

If it breaks: run a fresh crawl after each batch and filter for redirect chains and 404s. Remove dead URLs from the sitemap and fix internal links that still point at old addresses.

Step 9: tighten canonicals on what remains

Merged content often leaves duplicate variants behind, such as parameter URLs, print views and pagination. Make sure every surviving page has a self-referencing canonical and that variants point to the main version. I covered the rules in canonical tags explained.

Expected output: the crawl’s Canonicals tab shows no missing canonicals and no canonical pointing at a redirected or noindexed URL.

If it breaks: if Google picks a different canonical than the one you declared, check the Page indexing report in Search Console using Google’s documentation for that report and compare the declared and Google-selected canonical for the URL.

Step 10: monitor for 8 to 12 weeks

Log the date each batch shipped. Watch indexed page count, impressions and clicks for the surviving URLs. Expect pages to drop out of the index over several weeks, not overnight. Do not judge the result after ten days.

Expected output: fewer indexed pages, and equal or better clicks on the pages you kept or merged into.

If it breaks: if a merged page loses clicks, check that the redirect resolves cleanly and that the surviving page really covers what the old one did. Restore any content you cut that people were actually searching for.

common pitfalls

  • treating word count as the verdict: short pages can be excellent and long pages can be empty. Always pair it with demand and a read-through
  • deleting pages that carry links: you throw away the one asset the page had. Check backlinks first, then merge and redirect
  • rewriting with a generic AI pass: if you mass-produce filler to lift word counts, you have made the same problem longer. Google’s spam policies call out scaled content with little added value. If you use a model to draft, add real data, real examples and a human edit. The AI Tool Gazette blog covers which drafting tools are worth the time
  • redirecting everything to the homepage: the signal is usually discarded. Match topics one to one
  • doing it all in one deploy: when traffic moves you will not know which change did it. Ship in batches and log dates

scaling this

At 10x, meaning around 500 pages, everything above works in a spreadsheet with an afternoon of reading.

At 100x, around 5,000 pages, you stop reading everything. Group by template and sample 20 pages per group. If a template fails the sample, act on the template (noindex the whole pattern, or fix the generator) rather than page by page. Pull Search Console data through the API, and store results in a database or a proper BigQuery table instead of CSVs.

At 1000x, past 50,000 pages, the audit becomes a pipeline. You need log file analysis to see what Googlebot actually crawls, rules that auto-flag new thin pages at publish time, and templates that refuse to publish below a minimum content bar. Manual verdicts only happen for exceptions. The risk changes too. A wrong rule can noindex thousands of good pages at once, so test every rule on a 1 percent sample and keep a rollback list.

Whatever your size, keep the decision log. It is the only way to explain a traffic change six months later.

where to go next

Written by Xavier Fok

disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-09-29.

for SEOs
Tracking rankings or scraping SERPs at scale?

Rank checkers and SERP crawlers get blocked and geo-skewed fast on datacenter IPs. Singapore Mobile Proxy runs real 4G/5G mobile IPs that search engines still trust, so your position data stays clean.

see plans →
read on
More from The SEO Desk

Technical SEO, link building, content and SERP strategy, and tool reviews for people who ship growth.

browse all articles →