← all articles

Robots.txt vs noindex: when each one backfires

I’ve broken this on my own sites more than once, so let’s start with the mistake instead of the theory. A few years back I disallowed a whole staging subdomain in robots.txt because I didn’t want it competing with the live site. Six months later it was still showing up in search, as a bare URL with “no information is available for this page” under it. No title, no description, just a naked link sitting in the results. That’s what happens when you use the wrong tool for the job, and it’s the exact confusion this article is about.

Robots.txt and noindex sound like they do the same thing. They don’t. One controls whether Googlebot is allowed to fetch a page. The other controls whether a page that’s already been fetched gets stored in the index. Mixing them up doesn’t just fail quietly, it actively backfires, and the failure mode is different depending on which one you get wrong.

Robots.txt only controls crawling

Robots.txt is a set of instructions that live at yourdomain.com/robots.txt telling crawlers which paths they’re allowed to request. That’s it. It’s a crawling gate, not an indexing gate. Google has said plainly, and my own logs confirm it, that a URL disallowed in robots.txt can still end up indexed if something else links to it. Google never fetches the page, so it never sees the content, but it knows the URL exists because of the inbound link, and it can list that bare URL in search results anyway. That’s the “no information is available for this page” result. It’s not a bug. It’s the system working as designed, just not the way people expect.

This matters because a lot of site owners reach for robots.txt when what they actually want is “keep this out of Google.” Robots.txt can’t promise that. It can only promise “don’t request this file.” If your goal is exclusion from the index, robots.txt is the wrong lever, full stop.

Where robots.txt earns its keep is crawl efficiency, not secrecy. Internal search results pages, infinite faceted navigation on ecommerce category filters, calendar archives that generate a URL for every day going back a decade, login-gated account pages, these are places where Googlebot can burn its crawl budget on near-duplicate junk instead of your actual content. Blocking them in robots.txt tells the crawler “don’t waste time here,” and on a large site that has a real, measurable effect on how fast new or updated pages get picked up. On a small site with a few hundred URLs, crawl budget is rarely the bottleneck, so this whole category of problem barely applies.

Noindex only works if the page gets crawled

Noindex is a meta tag in the page’s head, or an X-Robots-Tag HTTP header, that tells Google “you can look at this page, but don’t put it in the index.” The catch, and it’s the catch that causes most of the damage, is that Google has to crawl the page to read that tag. If the page is blocked in robots.txt, the crawler never gets far enough to see the noindex directive at all. The two signals conflict, and robots.txt wins by default because it stops the process before noindex ever gets a chance to run.

This is exactly what happened with my staging subdomain. I’d added a noindex tag to those pages at some point too, thinking I was being thorough. It did nothing, because robots.txt was already blocking the fetch. The noindex tag was sitting there, correct and useless, on pages Google would never read again.

The backfire that traps pages in the index

Here’s the scenario that costs people the most time. A page gets indexed normally, no blocks in place. Later someone decides the page is low value and wants it gone, so they noindex it. That works fine, as long as robots.txt lets the crawler back in to re-read the page and notice the new tag. But if someone else, often on the same project, adds a robots.txt disallow rule for that same path around the same time, thinking they’re reinforcing the removal, the opposite happens. Google can no longer recrawl the page to confirm the noindex. Whatever was in the index at the last successful crawl just sits there. I’ve watched a URL stay indexed for months this way, with a noindex tag that was never wrong, just never read.

The fix, if you’re in this position, is almost always to remove the robots.txt block first, let Google recrawl and process the noindex, and only add a disallow rule afterward if you still want it, once the URL is confirmed out of the index. Order matters here in a way that isn’t obvious from reading either directive on its own.

The other backfire: blocking pages you never wanted hidden

The mirror image of that mistake is blocking a section in robots.txt for crawl budget reasons, then being surprised when those pages start collecting inbound links from other sites and showing up bare in search anyway. This is common with parameter-based URLs, print versions, or tracking-tagged links that get shared and linked externally more than anyone expects. If you actually care what shows up under those URLs in search, robots.txt won’t give you control over that. Noindex, with the page still crawlable, will.

There’s also a quieter version of this that trips up JavaScript-heavy sites. If a noindex tag gets injected client-side, after the page loads, and Google’s renderer either doesn’t get to that step or the resources needed to run the JS are themselves blocked in robots.txt, the directive never fires. I’ve seen this on a site where the CMS added noindex tags through a tag manager script, and the script’s own path was disallowed for an unrelated reason. The tag was never wrong. It just never ran.

How I actually decide between them

I use a rough rule now, learned the expensive way. If the goal is “don’t let Google spend time crawling this,” that’s robots.txt. Internal search, filtered URLs with no unique content, admin paths. If the goal is “let Google see this but don’t put it in search results,” that’s noindex, and the page has to stay crawlable for the tag to do anything. If I want a page to disappear from an index it’s already in, noindex first, confirm removal, then decide if a robots.txt block is even still needed on top of that.

One more thing worth saying plainly: neither of these tools is a security measure. Robots.txt is a public file that lists exactly the paths you didn’t want people looking at, which is its own small irony. If a page genuinely needs to stay private, that’s what authentication is for, not a crawler directive that well-behaved bots respect and nothing else has to.

None of this is exotic. It’s the kind of detail that only bites you once you’re managing enough pages, or enough sites, that a wrong assumption compounds instead of staying a one-off. If you want a second set of eyes on how your directives are actually interacting, or you’re trying to sort out an indexing mess that robots.txt and noindex made together, that’s the kind of technical SEO work we do at The SEO Desk.

for SEOs
Tracking rankings or scraping SERPs at scale?

Rank checkers and SERP crawlers get blocked and geo-skewed fast on datacenter IPs. Singapore Mobile Proxy runs real 4G/5G mobile IPs that search engines still trust, so your position data stays clean.

see plans →
read on
More from The SEO Desk

Technical SEO, link building, content and SERP strategy, and tool reviews for people who ship growth.

browse all articles →