Explainer

Proxy Scraper: What One Collects, and How Little of It Works

A proxy scraper collects proxy addresses from public pages and tests them. What ours found since May 2026, why most of it fails, and a safe script to run one.

HProxy Team··Updated September 27, 2026·11 min read
HProxy.Explainer

Skip the dead lists.

Our free proxy list re-checks every exit every few minutes across 100+ countries, with a live last-checked time, so you copy IPs that worked moments ago, not a stale text dump.

Open the free proxy list→

A proxy scraper is a program that collects proxy entries, each an IP address and a port, from public pages and files. It drops the duplicates and tests what is left. The collecting part is ten lines of code, as the script below shows. The testing decides whether the result is worth anything, because most of what a scraper collects does not work.

The proxies it finds are open proxies, servers that, in the words of Wikipedia, are "accessible by any Internet user." We run a scraper of our own, and it feeds our free proxy list. Since 13 May 2026 it has found 853,181 entries on 116 public sources. By 27 September, 44.3% of them had passed at least one of our checks, and 19,779 were working that day. This page shows how a proxy scraper works, what ours finds, and how to run a small one safely.

Proxy scraper or scraper proxy?

The two phrases sound alike and mean different things. A proxy scraper collects proxies. A scraper proxy is a proxy that a web scraper sends its requests through, and our guide to proxies for web scraping covers that one. If you were looking for the company ProxyScrape, our comparison is HProxy vs ProxyScrape.

How a proxy scraper works

What a proxy scraper does, in order
  1. Collect

    pages, files and API answers that publish proxies

  2. Parse

    every IP address and port in the text

  3. Deduplicate

    the same entry sits on many lists

  4. Check

    a real request through each proxy

  5. Repeat

    the lists change every day

Source: How the scraper behind the HProxy free proxy list works, at the level of our public pages.
  1. Collect. The scraper downloads pages, text files and API answers that publish proxies. A list can come as a table, a plain file with one ip:port per line, or JSON.
  2. Parse. It picks out each IP address and port, usually with one regular expression, and throws away anything that is not a valid address.
  3. Deduplicate. The same entry sits on many lists. Across everything our scraper has found, the average entry was listed by 8.3 of our 116 sources.
  4. Check. It sends a real request through each proxy to a test server. A port that accepts a connection is not yet a proxy that relays traffic. When we tested one public list on 5 September 2026, 1,770 of its 2,211 entries accepted a connection, and 306 relayed a request.
  5. Repeat. Lists churn, so a scraper has to run again and again. The researchers behind the 2018 study below fetched their lists again every day for that reason.

What our proxy scraper finds

We counted everything our scraper has found since the pool opened on 13 May 2026, and what became of it:

Everything our scraper found, 13 May to 27 September 2026

MeasuredEntriesShare
Found on 116 public sources853,181100%
Passed at least one of our checks378,01144.3%
Checked, never worked once475,14255.7%
Still in rotation on 27 September93,29610.9%
Working on 27 September19,7792.3%

HProxy free proxy pool, one read-only query on 27 September 2026, 13:43 UTC. Counts only, with no addresses and no source names.

More than half of everything our scraper collected never passed a single check. Counted by address, the picture is starker. On 28 September 2026, 42,321 of 284,790 addresses had ever passed one, 14.9 percent, and 87.5 percent of the entries that had sat on 1,181 hosts with 20 or more ports each. The finds keep coming anyway. In the 28 full days to 26 September, it found a median of about 2,650 new entries a day. The quietest day brought 651 and the busiest 26,584. On 26 September alone it found 15,765.

Our own console window: the proxy scraper behind our free proxy list, measured on 27 September 2026, 13:43 UTC. It found 853,181 entries on 116 public sources since 13 May 2026. 378,011 passed at least one check, 475,142 never worked once, and 19,779 were working on 27 September. Of 18,893 entries first found from 13 to 19 September, 13,942 passed at least one check and 464 were working on 27 September. Below that, the share that worked at least once by how many sources listed an entry, and two runs of proxy_scraper.py: 127 of 300 alive, and 4 of 300 alive.
Captured on our own machine on 27 September 2026: Node printing the saved result of one read-only query on our free proxy pool and of two runs of the script on this page. Counts only, no addresses.

One week of finds, followed

A fresh find is worth more than an old one, but not for long. We took the 18,893 entries our scraper first saw from 13 to 19 September 2026, and read what had become of them on 27 September:

One week of new finds, followed to 27 September 2026
First found, 13 to 19 September18,893100%
Passed at least one check13,94274%

26% lost from First found, 13 to 19 September

Still in rotation on 27 September6,24933%

55% lost from Passed at least one check

Working on 27 September4642%

93% lost from Still in rotation on 27 September

Most finds that passed a check were ports on 22 busy hosts. Eight to fourteen days later, about one in forty still worked.Source: HProxy free proxy pool, one read-only query on 27 September 2026, 13:43 UTC. Counts only.

Of that week, 73.8% passed at least one check, but most of those passes came from a few hosts. Read again by address on 28 September at 01:19 UTC, 12,917 of the 18,893 entries sat on 22 hosts that each added 20 or more new entries that week, most of them in one network. On ordinary hosts, with fewer than 20 new entries each, 2,635 of 5,976 entries passed a check even once, 44.1 percent. The list also decays fast. Eight to fourteen days later, 464 entries were working, 2.5% of the week.

Do more lists mean a better proxy?

A scraper that reads more sources mostly finds the same entries again. We grouped everything our scraper has found by the number of sources that listed it, and counted how many worked at least once:

Share that worked at least once, by how many sources listed it

Listed byEntriesWorked once
1 source225,11437.0%
2 to 4 sources269,43649.4%
5 to 9 sources100,40635.7%
10 to 19 sources152,02430.1%
20 or more106,20175.4%

HProxy free proxy pool, 27 September 2026, 13:43 UTC. How many sources listed an entry, as a count; no source is named.

Only the most widely copied entries stood out: 75.4% of those on 20 or more lists worked at least once. Between 1 and 19 lists, the share moved between 30.1% and 49.4%, with no steady rise. A long listing is a weak sign that a proxy works. A check is the strong one.

What independent research found

Our numbers describe one pool. Researchers at Georgetown University and Northeastern University measured the same thing on their own. Akshaya Mani, Tavish Vaidya, David Dworken and Micah Sherr followed "more than 107,000 listed open proxies" for 50 days. They published the results in 2018:

  • Availability: "more than 92% of open proxies that are advertised on proxy aggregator sites are offline or otherwise unavailable."
  • The lists held a median of 41,520 proxies a day, and a median of 3,283 a day responded.
  • Of 21,385 proxies that returned files to them, 413 (1.93%) "maliciously modified the .exe file at least once", adding malware to a Windows program.
  • Some proxies rewrote pages "to mine cryptocurrency", and they also detected "TLS man-in-the-middle (MitM) attacks".

Eight years apart, their lists and our pool tell the same story about availability. The malware findings are the reason to keep proxies from a scraper away from anything you install or log in to.

Run a small proxy scraper of your own

Four rules come before any code:

  • Collect only from pages whose terms allow it. A robots.txt file lists what its owner asks crawlers to leave alone. RFC 9309, the standard for it, describes rules "that crawlers are requested to honor when accessing URIs." It adds: "These rules are not a form of access authorization." A page that robots.txt does not forbid is therefore not automatically yours to take.
  • Say who you are. RFC 9309 asks crawlers to identify themselves in the User-Agent header, and says the identification "SHOULD describe the purpose of the crawler."
  • Go slowly. One page a second is plenty for a list.
  • Do not test proxies from your own machine. A hostile proxy sees what you send through it. Hand the list to a checker instead.

This script follows all four. It uses only the Python standard library:

"""A minimal proxy scraper: collect ip:port pairs from pages you may read,
drop duplicates, and let a proxy checker test each one, so your own machine
never sends traffic through an unknown proxy."""
import collections
import json
import re
import sys
import time
import urllib.request

PAIR = re.compile(r"\b((?:\d{1,3}\.){3}\d{1,3}):(\d{2,5})\b")
CHECKER = "https://hproxy.com/api/proxy-check"
HEADERS = {"User-Agent": "my-proxy-scraper/1.0 (+https://example.com/about-this-bot)"}


def collect(urls):
    found = set()
    for url in urls:
        request = urllib.request.Request(url, headers=HEADERS)
        with urllib.request.urlopen(request, timeout=30) as response:
            text = response.read().decode("utf-8", "replace")
        for ip, port in PAIR.findall(text):
            if all(int(part) <= 255 for part in ip.split(".")) and 0 < int(port) < 65536:
                found.add(f"{ip}:{port}")
        time.sleep(1)  # one page a second is plenty
    return sorted(found)


def check(proxies, retry=True):
    results = []
    for start in range(0, len(proxies), 1000):  # the checker takes 1,000 per request
        body = json.dumps({"proxies": proxies[start:start + 1000]}).encode()
        headers = {**HEADERS, "Content-Type": "application/json"}
        request = urllib.request.Request(CHECKER, data=body, headers=headers)
        with urllib.request.urlopen(request, timeout=120) as response:
            results += json.load(response)["results"]
        time.sleep(1)
    slow = [r["input"] for r in results if r["status"] == "unchecked"]
    if retry and slow:  # no answer inside the time budget: ask once more
        again = {r["input"]: r for r in check(slow, retry=False)}
        results = [again.get(r["input"], r) for r in results]
    return results


if __name__ == "__main__":
    found = collect(sys.argv[1:])
    results = check(found)
    alive = [r["input"] for r in results if r["status"] == "alive"]
    print(f"collected {len(found)} unique addresses")
    for status, count in collections.Counter(r["status"] for r in results).most_common():
        print(f"  {status}: {count}")
    with open("alive.txt", "w") as out:
        out.write("\n".join(alive) + "\n")
    print(f"wrote {len(alive)} working proxies to alive.txt")

Save it as proxy_scraper.py, then pass it the pages to read, one or more:

python proxy_scraper.py "https://example.com/a-page-you-may-read"

The checking happens on our server. The script sends the list to our free proxy checker API, 1,000 at a time, and asks once more for any that did not answer in time. The API needs no key, and it checks up to 300,000 proxies a day for each IP address that calls it.

We ran the script on 27 September 2026 against two slices of our own free list, the one set of pages we may certainly read. The first slice was 300 entries that had passed their latest check. The second was 300 entries that had answered within 48 hours, but not at their latest check:

$ python proxy_scraper.py "https://hproxy.com/api/proxy-list?format=txt&limit=300"
collected 300 unique addresses
  dead: 173
  alive: 127
wrote 127 working proxies to alive.txt

$ python proxy_scraper.py "https://hproxy.com/api/proxy-list?format=txt&recent=true&limit=300&offset=18696"
collected 300 unique addresses
  dead: 296
  alive: 4
wrote 4 working proxies to alive.txt

The offset in the second command skips the 18,696 entries that were working at that minute. Minutes after their latest check, 127 of the first 300 answered. Of the 300 that had gone quiet, 4 answered. That is the gap between a list and a checked list, in two commands.

Or use the output of ours

If you want the result rather than the program, our scraper publishes it:

Keeping such a list honest is the expensive half of the work. Our pool has logged 291,638,315 checks since May, and 129,645 entries were checked in the hour before our measurement.

When a free proxy is the wrong tool

Free proxies suit tests and small jobs. For a scraper that has to finish its work, a proxy that stays up matters more than a long list. HProxy residential proxies start at $0.44 per GB on Residential Lite, which targets by country.

How we measured

The pool figures come from one read-only query on our free proxy pool on 27 September 2026 at 13:43 UTC. Every date on this page is in UTC. Worked means the proxy passed at least one of our checks. The split by address comes from a second read-only query on 28 September 2026 at 01:19 UTC. Working on 27 September means it passed its latest check. Source counts are counts only: we do not publish which sources our scraper reads. We ran the script from our own computer, which contacted only hproxy.com; the proxies were tested from our server in the United States.

Sources

Frequently asked questions

What is a proxy scraper?
A program that collects proxy entries, each an IP address and a port, from public pages and files. It removes the duplicates and tests what is left. The testing decides the value: of 853,181 entries our own scraper found from 13 May to 27 September 2026, 44.3% passed even one of our checks.
Is a proxy scraper the same as a scraper proxy?
No. A proxy scraper collects proxies. A scraper proxy is a proxy that a web scraper sends its requests through, so the sites it reads see the address of the proxy. Search tools often report the two phrases as one, but they are different tools.
Are free proxy scrapers safe to use?
The scraper is as safe as its code. The proxies it finds are another matter. A 2018 study of more than 107,000 listed open proxies found proxies that added malware to downloads and mining code to pages. Let a checker test them, and never log in or download files through one.
How often should a proxy scraper run?
Many times a day. Of 18,893 entries our scraper first found between 13 and 19 September 2026, 13,942 passed at least one check, most of them ports on 22 busy hosts, and 464 still worked on 27 September. A list that was right last week is mostly wrong today.
What is the difference between a proxy scraper and a proxy checker?
A scraper finds entries, and a checker tests them. A good scraper runs a checker on everything it finds, because a published entry is only a claim. Our free proxy checker tests up to 1,000 proxies in one request, from our server rather than from your computer.
Is there a free proxy scraper API?
Our free proxy list API returns what our scraper found and checked: text, JSON or CSV, with filters for protocol, country and anonymity, and no key. A copy of the list is also published on GitHub, at hproxy-com/free-proxy-list.

Get proxies that are alive right now

Our free proxy list re-checks every exit every few minutes across 100+ countries, with a live last-checked time, so you copy IPs that worked moments ago, not a stale text dump. When the location has to survive a real check, the paid network holds up.

129M+ proxy checks run · 100+ countries · HTTP / HTTPS / SOCKS · re-checked every few minutes · no signup

HProxy.

Honest guides and comparisons on proxies, scraping and staying unblocked, from the team that runs the network.

RSS feed