← all articles

Reading a crawl report without panicking

The first time a crawl report wrecks your afternoon

The first time I ran a full site crawl on a property I actually cared about, the tool spat out 4,200 issues. I remember staring at that number and mentally rewriting my entire week around it. I didn’t touch anything else that day. I just clicked through tabs, opening tickets, assigning myself work that, in hindsight, didn’t need doing.

Most of that list was noise. I didn’t yet know how to read what it was telling me. A crawl report for SEO is a list of what a bot found when it walked your site the way a search engine bot might. It is not a list of what’s currently hurting your rankings. Those are two different things, and mixing them up is the fastest way to burn a week on nothing.

What a crawl report is actually doing

A crawler (Screaming Frog, Sitebulb, the crawl feature in Ahrefs or Semrush, whatever you use) starts at a URL, follows every link it finds, and records what happens at each one: status code, response time, title tag, canonical tag, whether it’s indexable, how deep it sits in your link structure. It does this mechanically. It doesn’t know your site’s intent, your CMS quirks, or which URLs you actually want indexed.

That matters because a lot of what shows up as an “issue” is the crawler faithfully reporting something you did on purpose. A faceted filter URL with no meta description. A tag archive that’s noindexed by design. A redirect chain from an old URL structure you migrated off two years ago and never fully cleaned up. None of these are wrong in the sense that something is broken. They’re just visible now, because you’re looking.

The report is a mirror, not a verdict. Your job is to work out which reflections matter.

Start with the count, not the list

Before you open a single URL, look at the totals by issue type, and more importantly, what fraction of your site each one touches. 4,200 issues sounds bad until you find out 3,800 of them are the exact same missing meta description pattern on a paginated blog archive that Google was never going to reward for having a hand written description anyway.

I sort every crawl report by issue count, largest first, and I ask one question per row: is this one thing, repeated many times, or is it many different things? A single templating mistake that hits every product page is one fix that clears thousands of rows. A scattered mix of one-off errors across different page types is a slower, manual slog. Knowing which one you’re looking at changes how you plan your afternoon.

Errors vs warnings vs notices (they are not the same emergency)

Every crawler I’ve used buckets findings into some version of errors, warnings, and notices (or issues, opportunities, and info, depending on the tool’s marketing team). Treat that hierarchy as real, because it roughly maps to how much a page’s ability to be crawled and indexed is actually affected.

Errors usually mean the page can’t be reached properly: 5xx server errors, broken internal links pointing to 404s, redirect loops that never resolve. These stop a bot cold on that URL. They’re worth fixing on a real timeline.

Warnings are things that degrade the page without breaking it: duplicate title tags, missing H1s, thin content flags, slow response times. These are worth reviewing but rarely worth an emergency fix session.

Notices are mostly informational: URLs with query parameters, pages more than three clicks from the homepage, missing hreflang on a site that only serves one language anyway. Read them, learn from them, and move on unless a pattern jumps out.

I’ve seen people spend a whole day on notices while a batch of 500 pages sat on a 5xx error untouched in the same report. The report doesn’t prioritize for you. You have to.

The five issues worth stopping for

Not everything deserves equal attention, so here’s what actually earns a stop-what-you’re-doing response, based on how crawlers and indexing behave:

Server errors (5xx) on pages that used to work. If a page was fine last crawl and is throwing 500s now, something changed on your server or CMS, not on Google’s end. That’s infrastructure, not SEO, and it compounds the longer it sits.

Noindex tags on pages you didn’t mean to noindex. This happens constantly after CMS migrations, staging-to-production pushes, or a plugin update that reset a default. A page you want ranked, tagged noindex, is invisible no matter how good the content is.

Canonical tags pointing somewhere other than themselves, unintentionally. Self-referencing canonicals are normal and fine. A canonical pointing to a different URL because of a templating bug means you’re telling search engines “index that one instead,” even when you meant this one.

Redirect chains longer than one hop. Each extra hop adds a small amount of crawl budget waste and a small amount of signal dilution. On a big site this adds up. On a small site it’s rarely urgent, but it’s cheap to fix once you see it.

A sudden spike in blocked-by-robots.txt on URLs that were previously crawlable. Someone edited the robots file, often without realizing the blast radius. This is worth checking within the day, because it can silently cut off entire sections.

Everything else in a typical report is worth logging, batching, and working through over weeks, not hours.

Issues that look bad and usually are not

A few report items look alarming on first read but rarely deserve the reaction they get.

“Thin content” flags on utility pages like a login screen, a cart page, or a search results template. These pages aren’t supposed to rank on their own, and a word-count-based thin content warning doesn’t know that.

Duplicate title tags across paginated series, like page 2 and page 3 of the same category. This is a structural reality of pagination, not a mistake, and search engines have handled paginated series for a long time without needing every page to have a unique title.

“Orphan pages” that are, on inspection, old promotional landing pages you deliberately unlinked after a campaign ended. Orphaned isn’t automatically bad. It only matters if the page is something you still want found.

High click depth on pages that are genuinely deep in your catalog on purpose, like a niche product variant. Click depth is a signal worth watching in aggregate, but one deep page isn’t a crisis.

I don’t ignore these categories entirely. I scan them for the exception, the one row where the pattern actually is a real problem, and then I move on.

Cross check before you touch anything

A crawl report tells you what a bot found while crawling your site the way you configured the crawl (respecting or ignoring robots.txt, following or not following nofollow, rendering JavaScript or not). It does not tell you what’s actually indexed, and it doesn’t tell you how a page is performing in search. Before I fix anything based on a crawl finding alone, I check it against Search Console’s coverage and page indexing reports for the same URLs.

I’ve caught crawl-tool false alarms this way more than once: a page flagged as noindex in the crawl because of a caching layer serving a stale header, while Search Console showed it indexed and getting impressions the whole time. Fixing that “issue” would have meant chasing a ghost. The crawler was accurate about what it saw. It just saw a snapshot that didn’t match reality anymore.

A workflow that keeps you out of the ditch

What I actually do now, on every site I run: sort by issue count, separate errors from warnings from notices, fix the five categories above first if they’re present, cross check anything before acting on it, and batch the rest into a backlog I work through weekly instead of all at once. That’s it. It’s not clever. It’s just the difference between a report telling you what to look at and a report telling you what to panic about, and those are not the same list.

If you want more of this kind of unglamorous, no-hype breakdown of how technical SEO actually works day to day, The SEO Desk covers it on the blog and the channel.

for SEOs
Tracking rankings or scraping SERPs at scale?

Rank checkers and SERP crawlers get blocked and geo-skewed fast on datacenter IPs. Singapore Mobile Proxy runs real 4G/5G mobile IPs that search engines still trust, so your position data stays clean.

see plans →
read on
More from The SEO Desk

Technical SEO, link building, content and SERP strategy, and tool reviews for people who ship growth.

browse all articles →