← all articles

The technical SEO that actually moves something

technical-seo seo site-audit static-sites

Two of my WordPress properties went down on the same day, from the same cause underneath, and stayed down while I worked out what had happened.

Nobody writes SEO posts about that. It reads as an ops story. But every ranking factor anyone has argued about on a podcast sits downstream of the page being reachable, and for those days mine were not.

Which is roughly the frame I have ended up with. There is a short list of things that decide whether a page can work at all, and a very long list of things an audit tool can detect. The two lists barely overlap.

One caveat before the rest of it. I cannot show you rank or traffic numbers to support any of this. My reporting is in a state I do not trust, so anything I said about positions or sessions would be decoration. What I can tell you is what broke, what I changed, and what I could see afterwards.

The five that matter

Does the url return the right status code. Is the page allowed into the index at all. Is the content inside the response the server sends, or fetched afterwards by the browser. Is it fast on a real phone on a real network. Does one address serve each piece of content.

That order is deliberate. Failing any one of them makes the ones below it irrelevant.

Reachable, and honest about it

A 404 ends the argument. So does a 500 arriving at the same moment your visitors do.

One of my properties lost a large batch of articles this way. Published, live for years, then returning 404 for months before anybody checked. No penalty, no manual action, no event to point at. The site quietly got smaller.

The sneaky variant is the soft 404. The content is gone, the server returns 200, and a polite “we could not find that” page renders in its place. Uptime monitors score it healthy. You spend months asserting a page exists.

Crawl your own sitemap. Crawl your own internal links. Read the status codes off the wire instead of out of the CMS admin, because the admin reports intentions and the wire reports what happened.

While you are in there, flatten redirect chains. Three hops to reach the current url is two extra requests and two extra places for a future migration to lose the thread.

Indexable is a separate question

A noindex tag that escaped from staging removes a page completely, and the page looks perfect the entire time. Loads. Fast. Correct in a browser. Invisible.

I look for this after any migration or rebuild, because that is when it turns up. View source, search for noindex, then read robots.txt as it currently stands rather than as you remember writing it.

Get the distinction right, because plenty of people have it inverted. robots.txt controls crawling. The noindex directive controls indexing. Blocking a url in robots.txt does not remove it from the index, it prevents anything from reading the instruction that would have removed it. If you want a page out, allow the crawl and serve the noindex.

What the server actually sent

Open view source, not the element inspector. The inspector shows the document after every script has run, which is a different artefact from what arrived.

If your article body is in that raw HTML, you are finished with this question.

If the raw response is a shell div plus a large javascript bundle that fetches the words afterwards, you have asked every crawler to run your application before it can read a sentence. Google does render javascript. It is a queue with a budget attached, and joining that queue for a 900 word article with two images buys you nothing.

Quickest possible check: curl the url and grep the output for a sentence from the middle of the piece. Present, move on. Missing, you have a rendering dependency that nobody chose on purpose.

This is where the static properties I run are simply in a better position. The HTML is assembled once at deploy and every request afterwards receives the finished document.

Fast on a phone, not fast in a lab

A lab score is a simulation, run on a machine in a data centre with a throttling profile applied. It is repeatable and cheap, which is how it became the industry’s shared vocabulary. It measures the simulation.

I can check the other version cheaply because of the day job. I run mobile proxy lines on real carrier SIMs, so putting a page through an actual mobile network on an actual handset costs me the walk to the rack.

Pages of mine scoring in the nineties have taken six or seven seconds to become readable on a real line with weak signal. No score mentioned that.

What eats those seconds is dull and almost identical across every site I have looked at. An uncompressed hero image. Four font files. An analytics tag, a heatmap tag, a consent banner and a chat widget, every one of them ahead of the text.

Deleting a chat widget nobody staffs beats a week of build configuration.

The weekend I have nothing to show for

Here is technical work I did that produced no visible result.

I took one property from a lab score in the high seventies to 96 over a weekend. Deferred scripts, inlined critical CSS, preloaded fonts, the whole ritual. The score moved. Nothing else I can point at did, and that site was already serving prebuilt HTML from a CDN, so what I optimised was the fastest component in the stack.

Same site, another evening: structured data across several hundred pages. Article markup, breadcrumbs, validating clean. No rich result has ever appeared.

Structured data may be earning something I cannot see. I would rather say I have no evidence than write it up as a win.

One address, or four

The most common serious problem I find presents as nothing at all.

An article on one of my sites answered at four addresses. http and https, with www and without, each returning 200, none of them redirecting to the others. On the static properties there is a fifth, because a folder path and the same path with index.html both resolve, which is what serving files off a file system means.

There is no error state here. Every version is correct and quick. A human clicking around would never see it.

Until something decides otherwise, though, those are separate documents. Inbound links split between them. One gets crawled more, another gets indexed, and the version you promote is not necessarily the version representing you.

On WordPress this is usually drifted settings plus a plugin rewriting urls. On static hosting it is the file system being honest about what it holds.

So pick one of everything. One protocol, one hostname, one slash convention. Everything else 301s to it, and every canonical tag names it. Then check your own internal links agree, because I had pages linking to their own site without the www while the canonical insisted on it.

Reading an audit report backwards

Point a crawler at a site of any size and several hundred issues come back. That list looks like a work queue. It is not one.

Audit tools rank by detectability. A missing meta description is trivial to spot, so it appears 400 times in red. Two hostnames serving your entire site is an architectural judgement, so it appears once, low down, phrased as a suggestion.

Ease of detection and cost of ignoring have no relationship to each other, and the report is sorted by the first.

Sort it yourself against the five questions. Reachability, then indexability, then duplicate addresses. Past that you are in preference territory.

Alt text is worth writing because blind people use your site. It is an accessibility obligation with an SEO footnote attached. Treating it as a ranking emergency because a tool coloured the row red is how a weekend disappears.

The static trade, priced honestly

I moved several properties to static builds on a CDN and I would do it again. The ledger looks like this.

Removed: PHP versions, plugin updates that break overnight, a database, cache plugins, an origin server that can fall over. The outage I opened with cannot happen on a static site, because there is nothing running to break.

Added: a build step. Edit, build, deploy. No fixing a typo from your phone in a taxi. A broken build stops publishing entirely, which is a worse day than a slow page.

I take that trade because the problems it removes are the ones that have actually taken my sites offline, and the problem it adds is one I can watch happening in front of me.

What none of this does

It does not create demand.

A perfectly delivered article about something nobody wants is a perfectly delivered article nobody reads. I have sites in decent technical condition with very few readers, because the topics were chosen badly, and no status code was ever going to fix that.

Technical work removes obstacles. If the writing is good and something in the delivery layer is stopping it being served, fixing that releases something that already existed. If the writing is weak, you have paved a fast road to somewhere nobody was going.

The checks I run, in the order I run them, are here.

for SEOs
Tracking rankings or scraping SERPs at scale?

Rank checkers and SERP crawlers get blocked and geo-skewed fast on datacenter IPs. Singapore Mobile Proxy runs real 4G/5G mobile IPs that search engines still trust, so your position data stays clean.

see plans →
read on
More from The SEO Desk

Technical SEO, link building, content and SERP strategy, and tool reviews for people who ship growth.

browse all articles →