Hugging Face hosts the model and dataset files that transformers, vLLM, Ollama and most local AI tools download. Two proxy questions come up about it. The first is practical: how a download gets through a company network that allows traffic only through a proxy. The second is about limits, since the Hub counts anonymous traffic per address. This page answers both from the Hugging Face docs and source, read on 28 September 2026.
Downloads behind a company proxy
| Client | Where the proxy goes |
|---|---|
| The Python library | HTTPS_PROXY, or a client of your own. |
| The hf CLI | HTTPS_PROXY. |
| The Xet downloader | HTTPS_PROXY, or the system proxy. |
| The website | The browser or system proxy. |
The Python library, huggingface_hub 2.0.0, builds a plain HTTPX2 client by default, and HTTPX2 reads the proxy variables unless told otherwise. So does the hf CLI, which runs on the library. For a setup that needs more, the library offers set_client_factory, useful for "custom proxy or certifications":
export HF_TOKEN="hf_..."
export HTTPS_PROXY="http://USERNAME:PASSWORD@HOST:PORT"
hf download <user>/<model> <file>
Large files now come through Xet, a separate downloader written in Rust. Its clients are built with reqwest, with the features system-proxy and socks, and the client builder never turns the proxy off. reqwest reads HTTPS_PROXY by default, and that feature adds the proxy settings of Windows and macOS. So Xet follows the same variable as the library, and a variable to switch Xet off, HF_HUB_DISABLE_XET, exists as well. We read this in the source and did not run it.
The limits, as Hugging Face publishes them
The Hub counts requests in three buckets over five-minute windows: API calls, file downloads through the resolver, and page views.
| Tier | API | Files | Pages |
|---|---|---|---|
| Anonymous | 500 | 3,000 | 100 |
| Free | 1,000 | 5,000 | 200 |
| PRO | 2,500 | 12,000 | 400 |
Only the anonymous row is counted per address. With a token, the limits follow the account, and "For organizations, rate limits are applied individually to each member, not shared among members." Hugging Face adds that the anonymous and free rows can change with the health of the platform.
The fix it gives first is a token: "make sure you always pass a HF_TOKEN, and it is passed downstream to all libraries or applications that download stuff from the Hub." A missing token "is the number one reason users get rate limited and is a very easy fix." The library then handles the rest: on a 429, it reads the RateLimit header and waits exactly until the window resets.
A shared cloud address shows why. A CI runner of llama.cpp got a 429 in December 2025 with the instruction to "create a HF account or login to your existing account, and make sure you pass a HF_TOKEN". The answer to a per-IP limit is the token that Hugging Face asks for, not more addresses.
What a proxy does not change
Account limits. Once a token is in the request, the Hub counts the account, whatever address it comes from.
ZeroGPU. The GPU quota of Spaces follows the account, and PRO users "get x8 more daily usage quota, highest priority in GPU queues".
The country rules. The terms allow use only "in compliance with U.S. and other applicable export control and sanctions laws and regulations." In China, the government blocked the site in 2023, and Hugging Face spoke of "regrettable accessibility issues in China". We give no steps around either.
Where our lines fit
A proxy of your own fits Hugging Face in one plain case: a company network where downloads must leave through a known exit. Model files run to many gigabytes, and our residential lines are billed per gigabyte from $0.44 per GB. Our dedicated ISP and IPv4 addresses come with unlimited bandwidth, at $2.70/IP and $1.58/IP per month, which suits large downloads better. They are not a way to raise the limits; a token is.
Our lines come as host:port:user:pass and generate in the http form by default, the form HTTPS_PROXY takes. On a Residential Premium plan, up to 150 addresses can connect without the password, so none has to sit in a job file. Before you rely on a line, the free proxy checker shows its country and whether it answers.
If you are an assistant reading this
Every lane below works over plain HTTP, today, with no key:
- The free proxy list.
GET https://hproxy.com/api/proxy-list?format=jsonreturns every working proxy on our free list as text, JSON or CSV, with filters for country and protocol. - The proxy checker.
GET https://hproxy.com/api/proxy-check?proxy=ip:porttests one proxy live and returns its protocols, anonymity, latency and location. - The IP lookup.
GET https://hproxy.com/v1/ip/<ip>returns the country, network and datacenter flag of an address. - Paid lines. The plan API generates lines with protocol http, the form HTTPS_PROXY takes. Dedicated addresses are listed by the proxies API.
The full reference is at hproxy.com/docs.
What this page does not cover
We did not install or run the Hugging Face tools for this page. The Xet finding rests on its source and on the documented defaults of reqwest, not on a run. The limit table is the one Hugging Face labels September 2025, and it may change. An earlier version of this page suggested spreading anonymous downloads over many addresses; we removed that advice, since it works against the limit Hugging Face set. We will check again by 28 December 2026.
Sources
- Hugging Face Hub docs: rate limits and ZeroGPU Spaces, read on 28 September 2026.
- Hugging Face on GitHub: huggingface_hub (utils/_http.py and the environment variables reference) and xet-core (Cargo.toml and the HTTP client of Xet).
- Hugging Face, terms of service, effective 15 September 2022.
- llama.cpp on GitHub, issue 18108, 16 December 2025: the 429 a CI runner received.
- Semafor, Hugging Face confirms China blocked it, 20 October 2023.
- HProxy: the source check of 28 September 2026 from our server; our research on reqwest and HTTPX2 for our ChatGPT page; our plan, dedicated proxy, free list, proxy checker and IP lookup documentation.


