A proxy scraper is a program that collects proxy entries, each an IP address and a port, from public pages and files. It drops the duplicates and tests what is left. The collecting part is ten lines of code, as the script below shows. The testing decides whether the result is worth anything, because most of what a scraper collects does not work.
The proxies it finds are open proxies, servers that, in the words of Wikipedia, are "accessible by any Internet user." We run a scraper of our own, and it feeds our free proxy list. Since 13 May 2026 it has found 853,181 entries on 116 public sources. By 27 September, 44.3% of them had passed at least one of our checks, and 19,779 were working that day. This page shows how a proxy scraper works, what ours finds, and how to run a small one safely.
Proxy scraper or scraper proxy?
The two phrases sound alike and mean different things. A proxy scraper collects proxies. A scraper proxy is a proxy that a web scraper sends its requests through, and our guide to proxies for web scraping covers that one. If you were looking for the company ProxyScrape, our comparison is HProxy vs ProxyScrape.
How a proxy scraper works
Collect
pages, files and API answers that publish proxies
Parse
every IP address and port in the text
Deduplicate
the same entry sits on many lists
Check
a real request through each proxy
Repeat
the lists change every day
- Collect. The scraper downloads pages, text files and API answers that publish proxies. A list can come as a table, a plain file with one
ip:portper line, or JSON. - Parse. It picks out each IP address and port, usually with one regular expression, and throws away anything that is not a valid address.
- Deduplicate. The same entry sits on many lists. Across everything our scraper has found, the average entry was listed by 8.3 of our 116 sources.
- Check. It sends a real request through each proxy to a test server. A port that accepts a connection is not yet a proxy that relays traffic. When we tested one public list on 5 September 2026, 1,770 of its 2,211 entries accepted a connection, and 306 relayed a request.
- Repeat. Lists churn, so a scraper has to run again and again. The researchers behind the 2018 study below fetched their lists again every day for that reason.
What our proxy scraper finds
We counted everything our scraper has found since the pool opened on 13 May 2026, and what became of it:
Everything our scraper found, 13 May to 27 September 2026
| Measured | Entries | Share |
|---|---|---|
| Found on 116 public sources | 853,181 | 100% |
| Passed at least one of our checks | 378,011 | 44.3% |
| Checked, never worked once | 475,142 | 55.7% |
| Still in rotation on 27 September | 93,296 | 10.9% |
| Working on 27 September | 19,779 | 2.3% |
HProxy free proxy pool, one read-only query on 27 September 2026, 13:43 UTC. Counts only, with no addresses and no source names.
More than half of everything our scraper collected never passed a single check. Counted by address, the picture is starker. On 28 September 2026, 42,321 of 284,790 addresses had ever passed one, 14.9 percent, and 87.5 percent of the entries that had sat on 1,181 hosts with 20 or more ports each. The finds keep coming anyway. In the 28 full days to 26 September, it found a median of about 2,650 new entries a day. The quietest day brought 651 and the busiest 26,584. On 26 September alone it found 15,765.

One week of finds, followed
A fresh find is worth more than an old one, but not for long. We took the 18,893 entries our scraper first saw from 13 to 19 September 2026, and read what had become of them on 27 September:
26% lost from First found, 13 to 19 September
55% lost from Passed at least one check
93% lost from Still in rotation on 27 September
Of that week, 73.8% passed at least one check, but most of those passes came from a few hosts. Read again by address on 28 September at 01:19 UTC, 12,917 of the 18,893 entries sat on 22 hosts that each added 20 or more new entries that week, most of them in one network. On ordinary hosts, with fewer than 20 new entries each, 2,635 of 5,976 entries passed a check even once, 44.1 percent. The list also decays fast. Eight to fourteen days later, 464 entries were working, 2.5% of the week.
Do more lists mean a better proxy?
A scraper that reads more sources mostly finds the same entries again. We grouped everything our scraper has found by the number of sources that listed it, and counted how many worked at least once:
Share that worked at least once, by how many sources listed it
| Listed by | Entries | Worked once |
|---|---|---|
| 1 source | 225,114 | 37.0% |
| 2 to 4 sources | 269,436 | 49.4% |
| 5 to 9 sources | 100,406 | 35.7% |
| 10 to 19 sources | 152,024 | 30.1% |
| 20 or more | 106,201 | 75.4% |
HProxy free proxy pool, 27 September 2026, 13:43 UTC. How many sources listed an entry, as a count; no source is named.
Only the most widely copied entries stood out: 75.4% of those on 20 or more lists worked at least once. Between 1 and 19 lists, the share moved between 30.1% and 49.4%, with no steady rise. A long listing is a weak sign that a proxy works. A check is the strong one.
What independent research found
Our numbers describe one pool. Researchers at Georgetown University and Northeastern University measured the same thing on their own. Akshaya Mani, Tavish Vaidya, David Dworken and Micah Sherr followed "more than 107,000 listed open proxies" for 50 days. They published the results in 2018:
- Availability: "more than 92% of open proxies that are advertised on proxy aggregator sites are offline or otherwise unavailable."
- The lists held a median of 41,520 proxies a day, and a median of 3,283 a day responded.
- Of 21,385 proxies that returned files to them, 413 (1.93%) "maliciously modified the .exe file at least once", adding malware to a Windows program.
- Some proxies rewrote pages "to mine cryptocurrency", and they also detected "TLS man-in-the-middle (MitM) attacks".
Eight years apart, their lists and our pool tell the same story about availability. The malware findings are the reason to keep proxies from a scraper away from anything you install or log in to.
Run a small proxy scraper of your own
Four rules come before any code:
- Collect only from pages whose terms allow it. A robots.txt file lists what its owner asks crawlers to leave alone. RFC 9309, the standard for it, describes rules "that crawlers are requested to honor when accessing URIs." It adds: "These rules are not a form of access authorization." A page that robots.txt does not forbid is therefore not automatically yours to take.
- Say who you are. RFC 9309 asks crawlers to identify themselves in the User-Agent header, and says the identification "SHOULD describe the purpose of the crawler."
- Go slowly. One page a second is plenty for a list.
- Do not test proxies from your own machine. A hostile proxy sees what you send through it. Hand the list to a checker instead.
This script follows all four. It uses only the Python standard library:
"""A minimal proxy scraper: collect ip:port pairs from pages you may read,
drop duplicates, and let a proxy checker test each one, so your own machine
never sends traffic through an unknown proxy."""
import collections
import json
import re
import sys
import time
import urllib.request
PAIR = re.compile(r"\b((?:\d{1,3}\.){3}\d{1,3}):(\d{2,5})\b")
CHECKER = "https://hproxy.com/api/proxy-check"
HEADERS = {"User-Agent": "my-proxy-scraper/1.0 (+https://example.com/about-this-bot)"}
def collect(urls):
found = set()
for url in urls:
request = urllib.request.Request(url, headers=HEADERS)
with urllib.request.urlopen(request, timeout=30) as response:
text = response.read().decode("utf-8", "replace")
for ip, port in PAIR.findall(text):
if all(int(part) <= 255 for part in ip.split(".")) and 0 < int(port) < 65536:
found.add(f"{ip}:{port}")
time.sleep(1) # one page a second is plenty
return sorted(found)
def check(proxies, retry=True):
results = []
for start in range(0, len(proxies), 1000): # the checker takes 1,000 per request
body = json.dumps({"proxies": proxies[start:start + 1000]}).encode()
headers = {**HEADERS, "Content-Type": "application/json"}
request = urllib.request.Request(CHECKER, data=body, headers=headers)
with urllib.request.urlopen(request, timeout=120) as response:
results += json.load(response)["results"]
time.sleep(1)
slow = [r["input"] for r in results if r["status"] == "unchecked"]
if retry and slow: # no answer inside the time budget: ask once more
again = {r["input"]: r for r in check(slow, retry=False)}
results = [again.get(r["input"], r) for r in results]
return results
if __name__ == "__main__":
found = collect(sys.argv[1:])
results = check(found)
alive = [r["input"] for r in results if r["status"] == "alive"]
print(f"collected {len(found)} unique addresses")
for status, count in collections.Counter(r["status"] for r in results).most_common():
print(f" {status}: {count}")
with open("alive.txt", "w") as out:
out.write("\n".join(alive) + "\n")
print(f"wrote {len(alive)} working proxies to alive.txt")
Save it as proxy_scraper.py, then pass it the pages to read, one or more:
python proxy_scraper.py "https://example.com/a-page-you-may-read"
The checking happens on our server. The script sends the list to our free proxy checker API, 1,000 at a time, and asks once more for any that did not answer in time. The API needs no key, and it checks up to 300,000 proxies a day for each IP address that calls it.
We ran the script on 27 September 2026 against two slices of our own free list, the one set of pages we may certainly read. The first slice was 300 entries that had passed their latest check. The second was 300 entries that had answered within 48 hours, but not at their latest check:
$ python proxy_scraper.py "https://hproxy.com/api/proxy-list?format=txt&limit=300"
collected 300 unique addresses
dead: 173
alive: 127
wrote 127 working proxies to alive.txt
$ python proxy_scraper.py "https://hproxy.com/api/proxy-list?format=txt&recent=true&limit=300&offset=18696"
collected 300 unique addresses
dead: 296
alive: 4
wrote 4 working proxies to alive.txt
The offset in the second command skips the 18,696 entries that were working at that minute. Minutes after their latest check, 127 of the first 300 answered. Of the 300 that had gone quiet, 4 answered. That is the gap between a list and a checked list, in two commands.
Or use the output of ours
If you want the result rather than the program, our scraper publishes it:
- the free proxy list, with every entry checked around the clock and the time of its latest check one click away;
- the free proxy list API, the same list as text, JSON or CSV, with no key;
- a copy on GitHub, at hproxy-com/free-proxy-list;
- the proxy checker, for lists you collect elsewhere.
Keeping such a list honest is the expensive half of the work. Our pool has logged 291,638,315 checks since May, and 129,645 entries were checked in the hour before our measurement.
When a free proxy is the wrong tool
Free proxies suit tests and small jobs. For a scraper that has to finish its work, a proxy that stays up matters more than a long list. HProxy residential proxies start at $0.44 per GB on Residential Lite, which targets by country.
How we measured
The pool figures come from one read-only query on our free proxy pool on 27 September 2026 at 13:43 UTC. Every date on this page is in UTC. Worked means the proxy passed at least one of our checks. The split by address comes from a second read-only query on 28 September 2026 at 01:19 UTC. Working on 27 September means it passed its latest check. Source counts are counts only: we do not publish which sources our scraper reads. We ran the script from our own computer, which contacted only hproxy.com; the proxies were tested from our server in the United States.
Sources
- Akshaya Mani, Tavish Vaidya, David Dworken and Micah Sherr, An Extensive Evaluation of the Internet's Open Proxies, arXiv 1806.10258, 27 June 2018.
- IETF, RFC 9309: Robots Exclusion Protocol, September 2022.
- Wikipedia, Open proxy, as edited on 15 June 2026.
- HProxy documentation: the free proxy list API and the proxy checker API, read 27 September 2026.


