← all articles

Log file analysis: what your server logs show that a crawl never will

Two very different pictures of your site

Run Screaming Frog against your site and you get a simulation. The crawler starts at a URL, follows links the way it’s configured to, and builds a map of what it found. It’s useful, but it’s a guess about how a search engine might move through your site, generated by a tool that isn’t a search engine.

Your server logs are not a guess. Every time anything, a browser, a bot, a script, requests a page, image, or file from your server, that request gets written to a log line: the IP, the timestamp, the URL requested, the user agent, the status code returned, and usually the bytes served. If Googlebot hit a URL on your site last Tuesday at 3am, it’s in there. If it didn’t hit a URL at all for six months, that’s in there too, as an absence.

That’s the core difference. A crawl tells you what’s crawlable. A log tells you what actually got crawled, by whom, and how often. Those two things overlap less than most people expect, especially on anything bigger than a small brochure site.

Where to actually get the logs

This is the part that kills log analysis before it starts for a lot of people. You need raw access to server or CDN logs, and a lot of agencies working on client sites simply don’t have it. If you’re on shared hosting, check the control panel, cPanel usually has raw access logs sitting there unused. If you’re behind Cloudflare, you need Logpush or Enterprise-level log access, the free tier doesn’t give you this. If you’re on Nginx or Apache directly, the logs are just files on the server, usually in /var/log/nginx/ or similar, and you can pull them with SSH.

If nobody at the company can hand you a raw log file, log analysis is off the table until that changes. There’s no workaround, no third-party tool that reconstructs it after the fact. You either capture the requests as they happen or you don’t have the data.

The user agent string lies, reverse DNS doesn’t

Anyone can send a request with the user agent string “Googlebot” in it. Scrapers do this constantly to get past basic bot filters. So step one of any real analysis is throwing out the user agent as proof of anything and verifying the IP instead.

Google publishes documentation on how to verify Googlebot through reverse DNS lookup: take the IP that made the request, do a reverse DNS lookup, confirm it resolves to a googlebot.com or google.com domain, then do a forward lookup on that hostname and confirm it matches the original IP. Tools like Screaming Frog’s Log File Analyser do this verification automatically when you feed it a raw log. If you’re doing it by hand with a script, it’s a bit of work but not complicated. Skip this step and your entire analysis is built on requests that might not be from Google at all.

What shows up that a crawl can’t tell you

Which pages Google actually bothers with. A crawl shows you every URL that’s technically reachable. Logs show you which of those URLs Googlebot has visited, and how recently. I’ve pulled logs on sites where a page sat in the main navigation, fully crawlable, internally linked from the homepage, and Googlebot hadn’t touched it in over four months. A crawler would never flag that as a problem because structurally there’s nothing wrong with the page. The log is the only thing that shows the bot isn’t interested.

Crawl budget going somewhere you didn’t intend. On larger sites, log analysis regularly turns up bots spending a disproportionate share of requests on filtered category URLs, session ID variants, or old paginated archives nobody links to anymore, while pages you actually want indexed get crawled rarely. A crawl tool won’t show you this distribution because it doesn’t know how Google is allocating its own attention. It only knows the site’s link graph.

Status codes bots actually received, not what’s live now. If a page 404’d for two weeks after a bad deploy and has since been fixed, a fresh crawl today shows a clean 200. The log from that two-week window shows Googlebot hitting a 404, which matters because repeated errors on a URL can affect how much Google trusts and revisits it, independent of whether it’s fixed now.

Orphaned pages still getting hit. Pages with no current internal links can still get crawled if Google remembers them from before, or found them through an external backlink or an old sitemap. A crawl starting from your homepage will never find these because there’s no path to them anymore. The log doesn’t care about your current link structure, it just shows the request came in.

The gap between your sitemap and what’s actually being crawled. Cross-referencing log data against your XML sitemap shows you sitemap URLs Google is ignoring, and non-sitemap URLs Google is crawling anyway. Both are worth knowing and neither shows up in a standard crawl audit.

Mobile vs desktop bot behavior, and crawl frequency trends over time. Since the move to mobile-first indexing, most sites should see the mobile Googlebot doing the majority of the crawling. If your logs show desktop-Googlebot dominating, that’s worth investigating on its own. Trend the volume of verified Googlebot hits week over week and you can usually spot when a migration, a robots.txt change, or a big content push changed how much attention the site is getting, days or weeks before it would show up in rankings or Search Console reporting.

What it does not tell you

Log analysis will not tell you why a page isn’t ranking. It tells you whether the page is being fetched, not what Google’s ranking systems decided to do with it afterward. Those are separate systems and conflating them leads to wrong conclusions. A page can be crawled constantly and still rank poorly for reasons that have nothing to do with crawling: thin content, weak relevance, no supporting links, better pages already occupying the results.

It also isn’t going to be revelatory on a small site. If you’re running a 40-page local business site, a crawl plus Search Console coverage reports will tell you almost everything log analysis would, with a fraction of the setup. Log analysis earns its keep on sites with thousands of URLs, faceted navigation, frequent publishing, or a history of crawl and indexing problems that Search Console’s sampled data doesn’t explain well enough. Below that scale, it’s real work for a marginal answer.

A basic workflow that actually works

Pull at least two to four weeks of raw logs, more if you can get it. Filter to requests with a bot-like user agent, then verify each IP with reverse DNS and keep only the confirmed Googlebot hits. Group by URL and count hits, and note the most recent hit date per URL. Pull your full URL list from your CMS or a crawl, and your sitemap separately. Now you can compare three sets against each other: URLs that exist, URLs in the sitemap, and URLs Googlebot actually requested. The mismatches between those three lists are where the useful findings live, not in any single list on its own.

Do this once and you’ll have a baseline. Do it quarterly on a site that’s actively growing and you’ll start to see whether crawl attention is moving toward the pages you’re actually trying to rank or away from them.

If you want a walkthrough of pulling and filtering log files without a paid crawler license, or a second pair of eyes on what your own logs are showing, The SEO Desk covers technical SEO groundwork like this alongside the rest of our keyword research and link building work.

for SEOs
Tracking rankings or scraping SERPs at scale?

Rank checkers and SERP crawlers get blocked and geo-skewed fast on datacenter IPs. Singapore Mobile Proxy runs real 4G/5G mobile IPs that search engines still trust, so your position data stays clean.

see plans →
read on
More from The SEO Desk

Technical SEO, link building, content and SERP strategy, and tool reviews for people who ship growth.

browse all articles →