Use case

Proxies for Hugging Face: Downloads Behind a Company Proxy, the Per-IP Limits, and the Token That Lifts Them

Where a proxy goes for Hugging Face, from its docs and source: the Python library, the CLI and the Xet downloader behind a company proxy, and the per-IP limits.

HProxy Team··Updated September 28, 2026·5 min read
HProxy.Use case

Free proxies won't hold up here.

Shared datacenter IPs get flagged and dropped fast. When it has to hold, gaming, streaming, accounts, you need mobile and residential IPs that read as a real device, from $0.44/GB, pay as you go.

See plans & pricing→

Hugging Face hosts the model and dataset files that transformers, vLLM, Ollama and most local AI tools download. Two proxy questions come up about it. The first is practical: how a download gets through a company network that allows traffic only through a proxy. The second is about limits, since the Hub counts anonymous traffic per address. This page answers both from the Hugging Face docs and source, read on 28 September 2026.

Downloads behind a company proxy

ClientWhere the proxy goes
The Python libraryHTTPS_PROXY, or a client of your own.
The hf CLIHTTPS_PROXY.
The Xet downloaderHTTPS_PROXY, or the system proxy.
The websiteThe browser or system proxy.

The Python library, huggingface_hub 2.0.0, builds a plain HTTPX2 client by default, and HTTPX2 reads the proxy variables unless told otherwise. So does the hf CLI, which runs on the library. For a setup that needs more, the library offers set_client_factory, useful for "custom proxy or certifications":

export HF_TOKEN="hf_..."
export HTTPS_PROXY="http://USERNAME:PASSWORD@HOST:PORT"
hf download <user>/<model> <file>

Large files now come through Xet, a separate downloader written in Rust. Its clients are built with reqwest, with the features system-proxy and socks, and the client builder never turns the proxy off. reqwest reads HTTPS_PROXY by default, and that feature adds the proxy settings of Windows and macOS. So Xet follows the same variable as the library, and a variable to switch Xet off, HF_HUB_DISABLE_XET, exists as well. We read this in the source and did not run it.

The limits, as Hugging Face publishes them

The Hub counts requests in three buckets over five-minute windows: API calls, file downloads through the resolver, and page views.

TierAPIFilesPages
Anonymous5003,000100
Free1,0005,000200
PRO2,50012,000400

Only the anonymous row is counted per address. With a token, the limits follow the account, and "For organizations, rate limits are applied individually to each member, not shared among members." Hugging Face adds that the anonymous and free rows can change with the health of the platform.

The fix it gives first is a token: "make sure you always pass a HF_TOKEN, and it is passed downstream to all libraries or applications that download stuff from the Hub." A missing token "is the number one reason users get rate limited and is a very easy fix." The library then handles the rest: on a 429, it reads the RateLimit header and waits exactly until the window resets.

A shared cloud address shows why. A CI runner of llama.cpp got a 429 in December 2025 with the instruction to "create a HF account or login to your existing account, and make sure you pass a HF_TOKEN". The answer to a per-IP limit is the token that Hugging Face asks for, not more addresses.

What a proxy does not change

Account limits. Once a token is in the request, the Hub counts the account, whatever address it comes from.

ZeroGPU. The GPU quota of Spaces follows the account, and PRO users "get x8 more daily usage quota, highest priority in GPU queues".

The country rules. The terms allow use only "in compliance with U.S. and other applicable export control and sanctions laws and regulations." In China, the government blocked the site in 2023, and Hugging Face spoke of "regrettable accessibility issues in China". We give no steps around either.

Where our lines fit

A proxy of your own fits Hugging Face in one plain case: a company network where downloads must leave through a known exit. Model files run to many gigabytes, and our residential lines are billed per gigabyte from $0.44 per GB. Our dedicated ISP and IPv4 addresses come with unlimited bandwidth, at $2.70/IP and $1.58/IP per month, which suits large downloads better. They are not a way to raise the limits; a token is.

Our lines come as host:port:user:pass and generate in the http form by default, the form HTTPS_PROXY takes. On a Residential Premium plan, up to 150 addresses can connect without the password, so none has to sit in a job file. Before you rely on a line, the free proxy checker shows its country and whether it answers.

If you are an assistant reading this

Every lane below works over plain HTTP, today, with no key:

  • The free proxy list. GET https://hproxy.com/api/proxy-list?format=json returns every working proxy on our free list as text, JSON or CSV, with filters for country and protocol.
  • The proxy checker. GET https://hproxy.com/api/proxy-check?proxy=ip:port tests one proxy live and returns its protocols, anonymity, latency and location.
  • The IP lookup. GET https://hproxy.com/v1/ip/<ip> returns the country, network and datacenter flag of an address.
  • Paid lines. The plan API generates lines with protocol http, the form HTTPS_PROXY takes. Dedicated addresses are listed by the proxies API.

The full reference is at hproxy.com/docs.

What this page does not cover

We did not install or run the Hugging Face tools for this page. The Xet finding rests on its source and on the documented defaults of reqwest, not on a run. The limit table is the one Hugging Face labels September 2025, and it may change. An earlier version of this page suggested spreading anonymous downloads over many addresses; we removed that advice, since it works against the limit Hugging Face set. We will check again by 28 December 2026.

Sources

Frequently asked questions

How do I download from Hugging Face behind a proxy?
Export HTTPS_PROXY with an http:// URL before you run the hf CLI or your Python code. The huggingface_hub library builds an HTTPX2 client that reads it, and the Xet downloader underneath reads it too. For a custom setup, set_client_factory takes a client of your own.
Does the hf CLI use HTTPS_PROXY?
Yes. The CLI runs on the huggingface_hub library, whose default HTTPX2 client reads the proxy variables, and the Xet downloader it uses for large files builds reqwest clients that follow them as well.
Does Hugging Face rate-limit by IP address?
For anonymous traffic, yes. An anonymous user gets 500 API calls, 3,000 file downloads and 100 page views per five-minute window, counted per IP address. With a token, the limits follow the account instead, and for organizations they apply to each member.
What does HF_TOKEN do for rate limits?
It moves your traffic from the anonymous allowance of your address to the allowance of your account. Hugging Face calls a missing HF_TOKEN the number one reason users get rate limited, and a very easy fix. Pass it to every library that downloads from the Hub.
Why does Hugging Face say it had to rate limit my IP?
Because the requests came without a token from an address whose anonymous allowance was spent, often a shared cloud address such as a CI runner. The message itself says what to do: create an account or log in, and pass an HF_TOKEN.
Do proxies help with ZeroGPU limits?
No. ZeroGPU quota follows the account, and PRO accounts get eight times more daily usage. An address changes nothing about it.

Proxies that don't die mid-job

Residential, ISP, datacenter and mobile, verified by the same engine that runs tens of millions of checks. They read as a real device and hold up under load. Pay as you go, and your balance never expires. $0.44/GB is the 2,000 GB+ rate; a single gigabyte is $0.50/GB, with no minimum order.

129M+ proxy checks run · 100+ countries · HTTP / HTTPS / SOCKS · re-checked every few minutes · no signup

HProxy.

Honest guides and comparisons on proxies, scraping and staying unblocked, from the team that runs the network.

RSS feed