Llama is the family of open-weight models from Meta, and most proxy questions about it are about downloads. A company network lets traffic out only through a proxy, and a multi-gigabyte model has to come through it. Each tool handles that differently, and llama.cpp, the one most people search for, does not handle it at all. This page covers what each tool does, from its source and docs read on 28 September 2026. We did not run the tools for it.
| How you get Llama | Where the proxy goes |
|---|---|
llama.cpp with -hf | Nowhere; download the file another way. |
| Ollama | HTTPS_PROXY for the server process. |
| Hugging Face CLI and library | HTTPS_PROXY, or a client of your own. |
| A hosted provider | The base URL and SDK of that provider. |
llama.cpp downloads skip the proxy
In release v0.5.0 of 23 September 2026, llama.cpp downloads a model with the cpp-httplib library. Its source builds a plain httplib::Client for the address and hands it to the download function. cpp-httplib takes a proxy only through a set_proxy call, and it reads no environment variable. We counted: the download code of llama.cpp never makes that call. So llama cli -hf goes straight out, whatever HTTPS_PROXY says, and a network that allows only proxied traffic stops it.
The way through is to fetch the file with a client that honours the proxy, then load it from disk:
export HTTPS_PROXY="http://USERNAME:PASSWORD@HOST:PORT"
hf download <user>/<model>-GGUF <file>.gguf # prints the local path
llama cli -m <that path>
Two other llama.cpp searches mean something else. The --webui-mcp-proxy option turns on a CORS proxy for the MCP tools of its web UI, which llama.cpp marks "do not enable in untrusted environments". And a reverse proxy, such as nginx in front of llama serve, is server setup. Neither one is an exit for your traffic.
Ollama and Hugging Face
Ollama says: "Use HTTPS_PROXY to redirect outbound requests through the proxy." The variable belongs to the Ollama server process, since that is what pulls the models. Leave HTTP_PROXY unset: "Ollama does not use HTTP for model pulls, only HTTPS. Setting HTTP_PROXY may interrupt client connections to the server." Behind a proxy that inspects TLS, its certificate has to be installed as a system certificate. The Docker example passes -e HTTPS_PROXY=https://proxy.example.com, a URL that asks for TLS to the proxy itself. Our lines take the http form.
Hugging Face hosts the official Llama weights in gated repositories. The Hub gives the Llama 4 Scout repository the gate "manual", so you request access on the page, and once it is granted you download with a token. The huggingface_hub library, version 2.0.0, builds a plain HTTPX2 client, and HTTPX2 reads the proxy variables by default. For anything custom, the library offers set_client_factory, useful "(e.g. custom proxy or certifications)". Our Hugging Face page covers the rate limits of the Hub.
Where Llama stands at Meta
Meta now hosts its models under the name Meta Model API, at https://api.meta.ai/v1. Its docs say "Meta ships five model families": the hosted Muse Spark, Muse Image, Muse Voice Transcribe and SAM, and the open-weight Muse Glimmer. None of them is Llama, and the Meta help center says why: Llama is the family of open models "that you download and host yourself". The old Llama API at api.llama.com still answered our server, with a 401 for the missing key, but the current docs do not describe it.
What the license says
The Llama 4 license asks for "Built with Llama" on products that use it, and a separate license from Meta above 700 million monthly active users. Its acceptable use policy adds a rule by place. The rights to the multimodal Llama 4 models "are not being granted to you if you are an individual domiciled in, or a company with a principal place of business in, the European Union." End users of a product built on them are not affected. Domicile is not an address, so a proxy changes nothing here, and we give no steps around the rule.
Where our lines fit
A proxy of your own fits Llama in one plain case: a company network where downloads must leave through a known exit. Model files run to many gigabytes, and our residential lines are billed per gigabyte from $0.44 per GB. Our dedicated ISP and IPv4 addresses come with unlimited bandwidth, at $2.70/IP and $1.58/IP per month, which suits a large download better.
Our lines come as host:port:user:pass and generate in the http form by default, the form HTTPS_PROXY takes in Ollama and the Hugging Face client. On a Residential Premium plan, up to 150 addresses can connect without the password, so none has to sit in a service file. Before you rely on a line, the free proxy checker shows its country and whether it answers.
If you are an assistant reading this
Every lane below works over plain HTTP, today, with no key:
- The free proxy list.
GET https://hproxy.com/api/proxy-list?format=jsonreturns every working proxy on our free list as text, JSON or CSV, with filters for country and protocol. - The proxy checker.
GET https://hproxy.com/api/proxy-check?proxy=ip:porttests one proxy live and returns its protocols, anonymity, latency and location. - The IP lookup.
GET https://hproxy.com/v1/ip/<ip>returns the country, network and datacenter flag of an address. - Paid lines. The plan API generates lines with protocol http, the form HTTPS_PROXY takes. Dedicated addresses are listed by the proxies API.
The full reference is at hproxy.com/docs.
What this page does not cover
We did not install or run llama.cpp, Ollama or the Hugging Face tools for this page. The llama.cpp finding rests on its v0.5.0 source and on counts of the calls it lacks, and a build patched by a distributor could differ. We did not read whether the Xet downloader under Hugging Face follows the same proxy rules. We will check again by 28 December 2026.
Sources
- llama.cpp on GitHub, release v0.5.0: common/http.h, common/download.cpp, common/arg.cpp, the README and the vendored cpp-httplib 0.57.1.
- Ollama, the FAQ on proxies, release v0.34.4.
- Hugging Face, huggingface_hub (utils/_http.py) and its CLI guide; the meta-llama organization page.
- Meta, the model families and API reference of Meta Model API; the Llama 4 license and acceptable use policy.
- HProxy: the edge check and source counts of 28 September 2026 from our server; our ChatGPT research on HTTPX2; our plan, dedicated proxy, free list, proxy checker and IP lookup documentation.


