← all articles

What log files tell you that analytics cannot

A log file is a plain text record your web server keeps of every request it receives. Analytics tools like Google Analytics 4 record what happens inside a visitor’s browser. Log files record what happens at the server. Those are two different views of the same site, and the gap between them is bigger than most people expect.

I run sites out of Singapore, and the first time I opened a raw access log I found things my analytics had never shown me: bots hammering URLs I did not know existed, and pages Googlebot ignored for weeks. If you care about how search engines actually treat your site, logs are the closest thing to a ground truth you can get. This article covers what they are, how they work, and where they beat analytics.

What it is

A server access log is a file where your web server writes one line per request. Every time anything asks your server for a page, an image, a script or a stylesheet, the server notes it. That “anything” includes real people, Googlebot, Bingbot, AI crawlers, uptime monitors, scrapers and vulnerability scanners.

A typical line contains:

  • the IP address of the requester
  • a timestamp
  • the request itself, for example GET /blog/some-post HTTP/1.1
  • the HTTP status code the server returned, such as 200, 301, 404 or 500
  • the size of the response in bytes
  • the referrer, if there was one
  • the user agent string, which is the name the requester gives itself

Analytics works differently. A tracking script (a snippet of JavaScript) loads in the visitor’s browser, runs, and sends a message to the analytics vendor. No script execution means no record. That one design choice explains almost everything in this article.

Apache documents its log formats in the Apache HTTP Server logging guide, and nginx covers the same ground in its logging documentation. If you use a CDN or managed host, the logs may live in a dashboard rather than a file, but the contents are much the same.

How it works

Here is the path of a single request.

  1. A client (a browser or a bot) asks your server for /pricing.
  2. The server decides what to send back and attaches a status code. The meanings of those codes are defined in the HTTP standard, RFC 9110.
  3. The server writes one line to the access log. This happens whether or not the client is a human, and whether or not it runs JavaScript.
  4. If the client is a browser that runs JavaScript, the analytics script then loads and sends its own hit to the analytics vendor. A bot that does not run scripts stops after step 3.

So the log is written first, by the server, with no cooperation from the visitor. Analytics is written second, by the visitor’s browser, and only if everything goes right.

Things that stop analytics from recording a visit:

  • the visitor uses an ad blocker or tracking protection
  • the visitor declines a consent banner, so the script never fires
  • the visitor leaves before the script loads
  • the client is a crawler that does not execute JavaScript
  • the script is broken, misplaced or blocked by a content security policy

None of those stop the server from writing a log line.

One practical catch: user agent strings can be faked. Anyone can call their scraper “Googlebot”. Google documents how to confirm a real Googlebot visit with a reverse DNS lookup in its verifying Googlebot guide. Check that before you draw conclusions from any bot line.

To read logs at scale, people export them and filter them in a spreadsheet, load them into a database, or use a dedicated tool. Screaming Frog Log File Analyser is a common desktop option, and some hosts and CDNs offer built-in log views. For a small site, a command line and a week of logs are enough to start.

Why it matters

You see what Googlebot actually does

Analytics tells you about people. It says almost nothing about crawlers. Logs show which URLs Googlebot requested, how often, and what status code it got back. Google lists its crawlers and their user agents in its overview of Google crawlers, which is the reference to match against when you filter.

This matters because a page can only rank if it is crawled. If your logs show Googlebot visiting your top pages once a month while spending thousands of requests on parameter URLs, you have found a real problem that no analytics report would reveal. I go deeper on this in log file analysis and what it shows that a crawl cannot.

You find wasted crawl effort

Sites with filters, sorting options and tracking parameters can generate huge numbers of near-duplicate URLs. A crawler tool like Screaming Frog tells you those URLs exist. The log tells you whether Googlebot is really spending time on them. If you run a store or a listings site, read faceted navigation and the pages that multiply alongside your logs, because that is the usual source of the waste. The related problem of too many low value pages in the index is covered in index bloat explained.

You catch errors people never report

Visitors who hit a 404 or a 500 rarely tell you. Many just leave. The log records every one of those responses, including the ones triggered by bots following old links. A spike of 5xx errors in the log around a deployment is often the explanation when a page suddenly drops out of search. If you are chasing that kind of mystery, when a page vanishes from the index is a good companion read.

You see traffic analytics hides

Analytics undercounts real visitors who block scripts, and it misses bots entirely. The log catches both. That helps in a few ways:

  • spotting scrapers and aggressive bots that eat server resources
  • seeing how much AI crawler traffic you get, by user agent
  • confirming that a traffic dip in analytics is a measurement change (for example a consent banner update) rather than a real loss of visitors

One note on privacy. Logs contain IP addresses, and depending on where your visitors are, IP addresses can count as personal data. Check how long you keep logs and who can read them. This is not legal advice, so talk to someone qualified if you handle visitors in regulated regions. For broader reading on privacy practice, our sister site The Privacy Wire covers that side in more depth.

Common misconceptions

“Analytics already shows me all my traffic”

It does not. Analytics shows visits where a script ran and was allowed to report. Everything else is invisible to it. The size of the gap depends on your audience. A technical audience with ad blockers will show a much bigger gap than a general consumer one. I will not quote a figure, because it varies by site and I would be guessing.

“Logs replace analytics”

They do not. Logs are poor at telling you about people. They cannot say whether a visitor scrolled, converted, or came back next week as the same person. A shared office IP can look like one visitor, and one visitor on mobile data can look like many. Analytics is better for behaviour and goals. Logs are better for server-side truth and crawler behaviour. I use both.

“Only big sites need log analysis”

Size is not the deciding factor. A 200 page site rarely needs it. But a small site with a messy URL structure, a recent migration, or a sudden indexing problem can learn a lot from one afternoon in the logs. The test is whether you have a question about how crawlers treat the site that a crawl report cannot answer. If you are still getting used to crawl tools, start with reading a crawl report without panicking.

“A crawl tool shows the same thing”

A crawl tool starts at your homepage and follows links, so it shows what is reachable and how deep pages sit. That is useful, and what is crawl depth and why it matters explains it well. But a crawl is a simulation. The log is the record of what real bots did. Orphan pages, old redirected URLs and parameter junk that Googlebot still remembers all show up in logs and often not in a crawl.

Where to go from here

If this was new to you, a sensible order is:

My own routine is simple. I pull a week of access logs, filter to verified Googlebot, group requests by URL pattern and status code, and look for the top ten patterns. Nine times out of ten the answer to “why is Google ignoring my good pages” is sitting in that list.

Written by Xavier Fok

disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-10-10.

for SEOs
Tracking rankings or scraping SERPs at scale?

Rank checkers and SERP crawlers get blocked and geo-skewed fast on datacenter IPs. Singapore Mobile Proxy runs real 4G/5G mobile IPs that search engines still trust, so your position data stays clean.

see plans →
read on
More from The SEO Desk

Technical SEO, link building, content and SERP strategy, and tool reviews for people who ship growth.

browse all articles →