Use case

Proxies for Scrapy: how a spider takes a proxy, and what Scrapy does with it

How Scrapy takes a proxy, from its documentation: the meta key, credentials, middleware order, retries, SOCKS5 and the settings that set the pace.

HProxy Team··Updated September 28, 2026·5 min read
HProxy.Use case

Free proxies won't hold up here.

Shared datacenter IPs get flagged and dropped fast. When it has to hold, gaming, streaming, accounts, you need mobile and residential IPs that read as a real device, from $0.44/GB, pay as you go.

Proxies for Web Scraping→

Search Console shows 3 impressions in 28 days for this page and 32 in 90 days, with no query it will name. The Keyword Planner counts 260 searches a month each for "scrapy proxy" and "scrapy proxies", 90 for "scrapy rotating proxies", 20 for "scrapy proxy middleware" and 10 for "scrapy socks5". Google also suggests searches about a proxy pool, the settings and SOCKS5. So this page answers from the Scrapy documentation for version 2.19.0 of 10 September 2026, which we read on 28 September 2026.

Free proxy lists have their own page, free proxies for Scrapy. It tested 300 of them and shows a middleware that drops dead entries. This page covers how Scrapy takes any proxy, and what it does with it.

The short answer: Scrapy reads the proxy from the request or from the environment, and one built-in middleware applies it. Credentials go in the proxy URL, and Scrapy does not rotate proxies by itself.

How Scrapy takes a proxy

The built-in HttpProxyMiddleware "sets the HTTP proxy to use for requests, by setting the proxy meta value for Request objects". It also obeys the environment variables http_proxy, https_proxy and no_proxy. So a proxy for every request can come from the shell:

P="http://USER:PASS@proxy.example:8080"
export http_proxy="$P" https_proxy="$P"

For one request, set the meta key to a URL with the login before the host, such as http://user:pass@host:port. The documentation says the meta value "will take precedence over http_proxy/https_proxy environment variables".

proxy = "http://USER:PASS@proxy.example:8080"
yield scrapy.Request(
    url, meta={"proxy": proxy}
)

The documentation also notes that "Usually a proxy URL uses the http:// scheme". The default download handler "supports HTTPS proxies only for HTTP destinations", so write the proxy URL with http:// even for https:// pages.

Middleware order, and what a retry does

To add a middleware of your own, "add it to the DOWNLOADER_MIDDLEWARES setting", not to a setting of another name. Scrapy calls process_request() in increasing order. HttpProxyMiddleware sits at 750 and RetryMiddleware at 550, so a middleware that sets the proxy needs a lower number than 750.

A retry does not pick a new proxy by itself. The retry middleware "Returns a new Request object to retry the specified request", and the meta of a request is "shallow copied when the request is cloned". The copy keeps the same proxy URL, so with a list of single proxies, a dead entry gets the retry too, unless a middleware changes it. Our page on free proxies for Scrapy shows one that drops the failed entry in process_exception().

Rotation

None of the middlewares Scrapy turns on by default rotates proxies. Each request goes through the proxy set for it. For a list, a middleware that sets meta["proxy"] per request does the rotation.

Many guides send readers to the scrapy-rotating-proxies package, and it tops the search results for "scrapy rotating proxies". Its last release on PyPI is 0.6.2, uploaded on 25 May 2019.

SOCKS5

"HttpxDownloadHandler supports SOCKS proxies while other built-in handlers don’t", the documentation says. Scrapy added that handler in 2.15.0 and gave it SOCKS support in 2.17.0, on 7 July 2026. It needs the httpx extra and supports SOCKS5. The documentation warns: "This handler is experimental, and not yet recommended for production environments". The settings to switch it on are on our page about free proxies for Scrapy.

Certificates through a proxy

"By default Scrapy does not verify the TLS certificate of HTTPS servers", the security page says. With DOWNLOAD_VERIFY_CERTIFICATES = True, requests to servers with bad certificates "fail instead of silently succeeding". Any proxy sits in the middle of the connection, so the setting is worth turning on.

The settings that set the pace

A proxy changes the address a request leaves from. These settings decide how often requests go out:

  • CONCURRENT_REQUESTS: 16, in a new project or not.
  • CONCURRENT_REQUESTS_PER_DOMAIN: 1 in a project made with scrapy startproject, 8 without it.
  • DOWNLOAD_DELAY: 1 second in a new project, 0 without it.
  • ROBOTSTXT_OBEY: True in a new project, False without it.
  • DOWNLOAD_TIMEOUT: 180 seconds either way.

The limits count per domain, whatever proxy a request goes through. DOWNLOAD_DELAY exists "to avoid hitting servers too hard". With ROBOTSTXT_OBEY on, "Scrapy will respect robots.txt policies". The AutoThrottle extension, off by default, can "automatically adjust Scrapy to the optimum crawling speed".

The default user agent is Scrapy/VERSION (+https://scrapy.org). The documentation asks you to "Set it to a value that identifies you, including a URL or an email address where website owners can reach you". Then they can "ask you to adjust your crawler rather than block it".

What this page leaves out

This page gives no steps for getting past bot protection, such as dressing Scrapy up as a browser. Several ranking guides sell that. The first result for "scrapy proxy" calls rotation "handy for bypassing detection methods". A proxy does not change what a site allows: its terms and its robots.txt still apply.

If you use proxies for other work

Our free proxy list and proxy checker cost nothing, and how to check if a proxy is working explains the results. For the basics, see what a residential proxy is and whether free proxies are safe. Our residential, ISP and mobile proxies are on the pricing page, and residential traffic starts at $0.44/GB.

The limits of this page

This page follows the Scrapy documentation for version 2.19.0, and later versions may change the handlers and settings. We did not run a crawl for it; our test of 300 free proxies is on the free page. This page is due for a check by 28 December 2026.

Sources

We read every source below on 28 September 2026.

Frequently asked questions

How do I set a proxy in Scrapy?
Set request.meta['proxy'] to a proxy URL with the login in it, such as http://user:pass@host:port, or export the http_proxy and https_proxy environment variables. The built-in HttpProxyMiddleware applies it, and the meta key wins over the variables.
Does Scrapy rotate proxies by itself?
No. None of the middlewares Scrapy turns on by default rotates proxies, so each request goes through the proxy set for it. The scrapy-rotating-proxies package, which many guides use, has had no release since 0.6.2 in May 2019.
Does Scrapy support SOCKS5 proxies?
Only through HttpxDownloadHandler, which Scrapy added in 2.15.0 and gave SOCKS support in 2.17.0. It needs the httpx extra, supports SOCKS5, and its documentation calls it experimental and not yet recommended for production. The other built-in handlers do not support SOCKS.
What happens to the proxy when Scrapy retries a request?
The retry is a new request copied from the old one, and Request.meta is shallow copied, so the retry keeps the same proxy. With a list of single proxies, change it in process_exception, as our page on free proxies for Scrapy shows.
What user agent should a Scrapy spider send?
The Scrapy documentation says to set USER_AGENT to a value that identifies you, with a URL or an email address where website owners can reach you, so that they can ask you to adjust your crawler rather than block it.

Proxies that don't die mid-job

Residential, ISP, datacenter and mobile, verified by the same engine that runs tens of millions of checks. They read as a real device and hold up under load. Pay as you go, and your balance never expires. $0.44/GB is the 2,000 GB+ rate; a single gigabyte is $0.50/GB, with no minimum order.

129M+ proxy checks run · 100+ countries · HTTP / HTTPS / SOCKS · re-checked every few minutes · no signup

HProxy.

Honest guides and comparisons on proxies, scraping and staying unblocked, from the team that runs the network.

RSS feed