Web Scraping with Free Proxies in Python (With Code Examples)
Every scraper has the same origin story. You write a beautiful little script, it rips through the first dozen pages flawlessly, and then — 403 Forbidden. Or worse, a captcha asking you to identify seventeen traffic lights. Your IP address has been noticed, and it is no longer welcome. This is the part where proxies enter the chat.
Why sites block scrapers in the first place
Websites aren't blocking you out of spite (mostly). They block patterns: fifty requests a second from one IP, identical headers on every hit, no mouse movements, no mercy. Rate limiters count how fast you knock; IP bans decide you've knocked too much. Sometimes you get a soft punishment first — captchas, throttled responses, mysteriously empty pages — and sometimes it's a straight ban with no appeal. A proxy doesn't make you invisible — it spreads the knocking across many doors, so no single door gets suspicious. If you're new to the whole concept, our beginner's guide to proxy servers covers the basics first.
The 30-second version: one proxy in requests
If you've never used a proxy in Python, here's the entire trick. The requests library takes a proxies dictionary — one key for http:// URLs, one for https://. Point both at your proxy and you're done:
import requests
proxies = {
"http": "http://203.0.113.45:8080",
"https": "http://203.0.113.45:8080",
}
r = requests.get("https://example.com", proxies=proxies, timeout=10)
print(r.status_code) # 200, wearing somebody else's IP
That's genuinely it. Every request through that proxy now arrives wearing a different IP address. One caveat: this example uses an HTTP proxy. If you want to route through a SOCKS5 proxy instead, install the extra dependency with pip install requests[socks] and use a socks5:// scheme in the address — the rest of the code stays exactly the same. The requests documentation on proxies has the full details if you want the advanced options.
One proxy dies. Rotate.
A single free proxy is a single point of failure, and free proxies fail the way weather changes — suddenly, and without consulting you. The fix is a list and a loop: pick one at random, try the request, and if it blows up, pick another:
import random
import requests
from requests.exceptions import RequestException
PROXIES = [
"http://203.0.113.45:8080",
"http://198.51.100.23:3128",
"http://192.0.2.90:1080",
]
def fetch(url, tries=5):
last_err = None
for _ in range(tries):
proxy = random.choice(PROXIES)
try:
r = requests.get(
url,
proxies={"http": proxy, "https": proxy},
timeout=10,
headers={"User-Agent": "Mozilla/5.0 (compatible; MyScraper/1.0)"},
)
r.raise_for_status()
return r.text
except RequestException as e:
last_err = e # dead proxy — shrug, try another one
raise last_err
html = fetch("https://example.com")
A few notes from the school of hard knocks: random.choice beats round-robin, because dead proxies cluster at the top of every list; raise_for_status() turns HTTP errors into exceptions your retry loop can actually catch; and five tries is plenty — if five proxies in a row fail, the problem is your target, not your proxies.
Feeding your scraper a live proxy list
Hard-coding proxies is fine for a weekend project. For anything longer-lived, pull a fresh list programmatically instead of letting it rot in your source code. Our API hands you live proxies as JSON — no key, no signup, no ceremony:
import requests
def get_proxies(limit=50):
r = requests.get(
"https://proxynest.live/api/proxies",
params={"limit": limit},
timeout=15,
)
r.raise_for_status()
proxies = []
for p in r.json()["proxies"]:
if p["protocol"] == "http": # keep this example simple
proxies.append(f"http://{p['ip']}:{p['port']}")
return proxies
PROXIES = get_proxies(50)
print(f"loaded {len(PROXIES)} proxies")
Drop that into the rotation script above — replace the hard-coded list with get_proxies() — and your scraper starts every run with fresh IPs instead of yesterday's corpses. The full reference, including country and protocol filters, lives in the API docs. Prefer picking manually? The proxy list has the same data with copy buttons.
Timeouts and headers: the two lines everyone skips
Two lines separate a robust scraper from one that hangs until the heat death of the universe. Always set a timeout — ten seconds is sane — because a dead proxy will otherwise hold your connection open forever, politely, indefinitely. And set a real User-Agent header. The default python-requests/2.x header is essentially a name tag that reads "HELLO I AM A BOT". You don't need to impersonate a specific browser; you just need to not announce yourself as a script.
While you're at it, add a small delay between requests — time.sleep(1), maybe with a little random jitter. It costs you almost nothing and it's the single most effective anti-ban measure there is. Scrapers get blocked for behaving like machines; a one-second pause is the cheapest humanity you can buy.
The honest limits of free proxies
Let's not oversell this. Free proxies are the public bicycles of the internet: useful, free, and occasionally missing a wheel.
- They die fast. Expect a meaningful chunk of any free list to be dead on arrival. That's exactly what the retry loop is for — design for failure and it stops hurting.
- They're slow. You're sharing bandwidth with everyone else who found the same list. For large jobs, grab the entries with the lowest latency first.
- Don't hammer anyone. Space your requests — a second or two between hits — and respect robots.txt. Scraping politely isn't just ethics, it's self-preservation: aggressive scrapers get whole IP ranges banned, which ruins the proxies for everyone. More on how flaky free proxies can be in the FAQ.
- Never send anything sensitive through one. No logins, no sessions, no personal data. Free proxy operators can see unencrypted traffic — we covered this properly in our safety guide.
A quick word on ethics
Just because you can scrape a site doesn't mean you should scrape all of it, constantly. Check the terms of service, respect robots.txt, keep your request rate human, and don't republish other people's content as your own. Proxies are a tool for polite data collection, not an invisibility cloak for bad behavior.
Go build it
You now know the whole playbook: route through a proxy, rotate on failure, refresh your list from an API, set timeouts, and don't be a jerk about it. Grab a fresh batch below and put that retry loop to work.