← all articles

How to track rankings at scale with proxies

Last month one of my keyword trackers logged four keywords against the wrong city for eleven days before anyone noticed. The proxy pool had quietly failed over from a Singapore exit node to a Jakarta one after a modem dropped, and nobody was watching the location field.

Nothing crashed. No errors. Just wrong data sitting in a spreadsheet, looking exactly as confident as the right data would have looked.

That’s the actual risk with rank tracking at scale. Not getting blocked. Getting silently wrong.

This is for anyone pulling more than a handful of keyword positions a day and doing it themselves rather than paying for a rank tracker. Maybe you’re tracking rankings across fifty client sites, building an SEO tool with a checker built in, or you just don’t want a recurring bill for something you can assemble with a proxy pool and a cron job. I run mobile proxy infrastructure for a living, LTE modems, SIM rotation, carrier IPs, so this is written from the infrastructure side of the problem more than the SEO side. By the end you’ll have a rotating proxy setup, a query script that pulls positions without tripping a captcha every third request, somewhere to store the results, and a rough sense of where DIY stops being worth it.

what you need

  • a keyword list with location and device tagged per keyword. national rank and “Singapore, mobile” rank are different numbers, don’t conflate them
  • a proxy provider, residential or mobile, not datacenter, for anything pointed at a search engine. i run my own LTE modems for my proxy business, but for a first setup something like IPRoyal, Smartproxy, or Oxylabs will get you moving without buying hardware
  • python 3.10 or newer, the requests library, and either beautifulsoup4 for parsing rendered HTML or a proper SERP API (SerpApi, DataForSEO) if you’d rather not parse Google’s markup yourself
  • somewhere to store results. postgres if you’re doing this properly, a CSV if you’re just testing the idea
  • budget: expect somewhere in the range of $50-150/month in proxy traffic for a few hundred keywords tracked daily, scaling with volume and how conservatively you rotate
  • a cron-capable box or scheduler. a $5/month VPS is plenty to start

step by step

1. define exactly what you’re tracking

Write down the keyword, the search engine, the location (country, and city if it matters), and the device for every row. Google shows different results by city for anything with local intent, and mobile SERPs increasingly diverge from desktop ones with different pack placements and featured snippets.

Expected output: a table with one row per keyword-location-device combination, not per keyword.

If it breaks: if your numbers don’t match an incognito browser search from the same city, the mismatch is almost always here. Wrong location tag, not a proxy problem.

2. pick your proxy type

Datacenter proxies are cheap and get flagged fast, search engines fingerprint their IP ranges directly. Residential and mobile proxies route through real ISP or carrier connections and are much harder to blacklist wholesale, because blocking them means blocking real users too. For search engine querying specifically, skip datacenter IPs and pay the premium.

Expected output: an account with API or credential-based access to a rotating pool, or a pool of your own hardware if you’re running it that way.

If it breaks: if a “residential” pool seems unusually cheap, check whether it’s a peer-to-peer network built on an SDK bundled into free apps or a free VPN running on someone else’s phone. Some of the cheapest residential proxy pools on the market are exactly that. Worth knowing what you’re buying before you build anything on top of it.

3. build your rotation logic

Don’t reuse one proxy for many queries in a row. Rotate per request, or at minimum every few requests, with a randomised delay between them.

import random
import time
import requests

PROXIES = [
    "http://user:[email protected]:8000",
    "http://user:[email protected]:8000",
    # load this from your provider's pool endpoint instead of a static list
]

def get_serp(query, location):
    proxy = random.choice(PROXIES)
    headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64)"}
    resp = requests.get(
        "https://www.google.com/search",
        params={"q": query, "gl": location},
        proxies={"http": proxy, "https": proxy},
        headers=headers,
        timeout=10,
    )
    time.sleep(random.uniform(4, 12))
    return resp

Expected output: 200 responses with SERP HTML, spread across different exit IPs on each call.

If it breaks: a wall of 429 responses or captcha pages means your delay is too short or your pool is too small for the volume you’re pushing through it. Also check your headers. RFC 9110 defines what a User-Agent string actually is, and the default one requests sends is not it. Set a real browser string.

4. respect the target’s own rules

Before you point anything at a search engine or a client’s site, check its robots.txt and any stated rate limits. Google’s robots.txt documentation covers what’s disallowed and why. Most DIY rank tracking sits in a grey zone relative to a search engine’s terms of service regardless of how carefully you rotate IPs, which is true of most commercial rank trackers too. Not legal advice, just what I’ve observed running this kind of infrastructure. Weigh the risk yourself.

Expected output: a clear read on what’s technically disallowed versus what everyone quietly does anyway.

If it breaks: this doesn’t really “break” so much as get you rate-limited, or at enough volume, get your provider’s IP ranges blacklisted. That’s a business risk, not a bug in your script.

5. parse and store the result

Pull the position of your target domain from the SERP HTML, or the JSON payload if you’re using a SERP API, and write it to your database with the date, keyword, location, and device attached.

Expected output: one row per keyword per day, with a rank position or null if the domain didn’t show up in the pages you checked.

If it breaks: if positions swing wildly day to day for no obvious reason, check whether you’re catching a personalised or A/B-tested SERP. Clean, logged-out, no-cookie requests help, but you can’t fully control for experiments running on the other end.

6. schedule it

Cron it to run daily, staggered across a window rather than firing your entire keyword list in one five-minute burst.

# stagger a 500-keyword run across a few hours instead of one burst
0 3 * * * /usr/bin/python3 /opt/rank-tracker/run.py --shard 1
0 4 * * * /usr/bin/python3 /opt/rank-tracker/run.py --shard 2
0 5 * * * /usr/bin/python3 /opt/rank-tracker/run.py --shard 3

Expected output: fresh rows landing in your database every morning without you touching anything.

If it breaks: a silent failure, the script runs and exits clean but no new rows appear, usually means a proxy pool went stale overnight. Alert on row count, not just exit code.

7. cross-check against a source of truth

Once a week, spot-check ten keywords against Search Console’s own position data or a manual incognito search. I went into the tradeoffs between the two in Google Search Console vs third party rank trackers. Short version: Search Console gives you Google’s own average position for free, but only for your own verified property, never a competitor’s.

Expected output: your scraped numbers landing within a point or two of Search Console’s, for your own domain.

If it breaks: consistent, large gaps almost always trace back to a location or device mismatch from step 1, not a scraping bug.

8. watch for drift, not just downtime

The failure mode that actually costs money is the one from the intro. A proxy pool degrading quietly, exit locations drifting, or a provider swapping in different IP ranges without telling you. Log the exit IP and its geolocation alongside every result, not just the SERP data itself.

Expected output: an audit trail that lets you catch “why did this Singapore keyword suddenly move” and find a Jakarta IP behind it, rather than assuming an algorithm update.

If it breaks: if you can’t tell the difference between a real ranking change and infrastructure drift, you don’t have a rank tracker. You have a random number generator with a nice dashboard attached.

common pitfalls

  • treating datacenter proxies as good enough because they’re cheap, when they get identified and deprioritised faster than residential or mobile ranges
  • not tagging location and device per keyword, then wondering why the numbers don’t match what a client sees on their own phone
  • hammering one proxy with the full keyword list because rotation “felt slow”, which is the fastest way to get a whole subnet flagged at once
  • ignoring your own site’s crawl budget while worrying about the other side of the equation. if you’re running bots against a big site of your own too, the concerns rhyme, see how to fix crawl budget problems on a big site
  • assuming this is free forever because you built it yourself. proxy traffic, compute, and the hours spent babysitting blocks all cost something, it’s just not a monthly invoice you can point to

scaling this

At 10x, call it fifty to a few hundred keywords a day, a handful of residential proxies and a single cron script is genuinely fine. You can watch it manually and catch problems by eye.

At 100x, a few thousand keywords a day, you need an actual pool (rotating dozens to low hundreds of IPs), retry and backoff logic, a queue instead of a flat script, and monitoring that pages you when the daily row count drops. This is roughly where the economics start to flip. The engineering time to keep a DIY scraper alive starts competing directly with just paying a SERP API by the call.

At 1000x, tens of thousands of keywords a day across many domains, agency or tool-vendor territory, DIY proxy scraping stops making sense for most teams. I’ll say the unpopular part directly: if you’re past a few thousand keywords a day and still building this yourself, you’re probably not saving money, you’re paying an engineer to reinvent something three vendors already sell well. A licensed SERP API or an enterprise tracker’s own API is usually cheaper than the headcount it takes to keep a homegrown scraper from getting blocked. I should say, I haven’t personally run my own scripts past about four thousand keywords a day, past that point I’ve always switched to an API myself, so treat the 1000x numbers here as informed rather than tested. If you’re building the same kind of rotation logic for account logins instead of search queries, that’s a closely related problem, worth a look at how multiaccountops.com handles proxy hygiene for that use case.

where to go next

Written by Xavier Fok

disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-09-15.

for SEOs
Tracking rankings or scraping SERPs at scale?

Rank checkers and SERP crawlers get blocked and geo-skewed fast on datacenter IPs. Singapore Mobile Proxy runs real 4G/5G mobile IPs that search engines still trust, so your position data stays clean.

see plans →
read on
More from The SEO Desk

Technical SEO, link building, content and SERP strategy, and tool reviews for people who ship growth.

browse all articles →