Use case

Proxies for Llama: Model Downloads Through a Proxy in llama.cpp, Ollama and Hugging Face, and What the License Says

Where a proxy goes when you get and run Llama, from the source: why llama.cpp downloads skip HTTPS_PROXY, Ollama and Hugging Face settings, and the license.

HProxy Team··5 min read
HProxy.Use case

Free proxies won't hold up here.

Shared datacenter IPs get flagged and dropped fast. When it has to hold, gaming, streaming, accounts, you need mobile and residential IPs that read as a real device, from $0.44/GB, pay as you go.

See plans & pricing→

Llama is the family of open-weight models from Meta, and most proxy questions about it are about downloads. A company network lets traffic out only through a proxy, and a multi-gigabyte model has to come through it. Each tool handles that differently, and llama.cpp, the one most people search for, does not handle it at all. This page covers what each tool does, from its source and docs read on 28 September 2026. We did not run the tools for it.

How you get LlamaWhere the proxy goes
llama.cpp with -hfNowhere; download the file another way.
OllamaHTTPS_PROXY for the server process.
Hugging Face CLI and libraryHTTPS_PROXY, or a client of your own.
A hosted providerThe base URL and SDK of that provider.

llama.cpp downloads skip the proxy

In release v0.5.0 of 23 September 2026, llama.cpp downloads a model with the cpp-httplib library. Its source builds a plain httplib::Client for the address and hands it to the download function. cpp-httplib takes a proxy only through a set_proxy call, and it reads no environment variable. We counted: the download code of llama.cpp never makes that call. So llama cli -hf goes straight out, whatever HTTPS_PROXY says, and a network that allows only proxied traffic stops it.

The way through is to fetch the file with a client that honours the proxy, then load it from disk:

export HTTPS_PROXY="http://USERNAME:PASSWORD@HOST:PORT"
hf download <user>/<model>-GGUF <file>.gguf   # prints the local path
llama cli -m <that path>

Two other llama.cpp searches mean something else. The --webui-mcp-proxy option turns on a CORS proxy for the MCP tools of its web UI, which llama.cpp marks "do not enable in untrusted environments". And a reverse proxy, such as nginx in front of llama serve, is server setup. Neither one is an exit for your traffic.

Ollama and Hugging Face

Ollama says: "Use HTTPS_PROXY to redirect outbound requests through the proxy." The variable belongs to the Ollama server process, since that is what pulls the models. Leave HTTP_PROXY unset: "Ollama does not use HTTP for model pulls, only HTTPS. Setting HTTP_PROXY may interrupt client connections to the server." Behind a proxy that inspects TLS, its certificate has to be installed as a system certificate. The Docker example passes -e HTTPS_PROXY=https://proxy.example.com, a URL that asks for TLS to the proxy itself. Our lines take the http form.

Hugging Face hosts the official Llama weights in gated repositories. The Hub gives the Llama 4 Scout repository the gate "manual", so you request access on the page, and once it is granted you download with a token. The huggingface_hub library, version 2.0.0, builds a plain HTTPX2 client, and HTTPX2 reads the proxy variables by default. For anything custom, the library offers set_client_factory, useful "(e.g. custom proxy or certifications)". Our Hugging Face page covers the rate limits of the Hub.

Where Llama stands at Meta

Meta now hosts its models under the name Meta Model API, at https://api.meta.ai/v1. Its docs say "Meta ships five model families": the hosted Muse Spark, Muse Image, Muse Voice Transcribe and SAM, and the open-weight Muse Glimmer. None of them is Llama, and the Meta help center says why: Llama is the family of open models "that you download and host yourself". The old Llama API at api.llama.com still answered our server, with a 401 for the missing key, but the current docs do not describe it.

What the license says

The Llama 4 license asks for "Built with Llama" on products that use it, and a separate license from Meta above 700 million monthly active users. Its acceptable use policy adds a rule by place. The rights to the multimodal Llama 4 models "are not being granted to you if you are an individual domiciled in, or a company with a principal place of business in, the European Union." End users of a product built on them are not affected. Domicile is not an address, so a proxy changes nothing here, and we give no steps around the rule.

Where our lines fit

A proxy of your own fits Llama in one plain case: a company network where downloads must leave through a known exit. Model files run to many gigabytes, and our residential lines are billed per gigabyte from $0.44 per GB. Our dedicated ISP and IPv4 addresses come with unlimited bandwidth, at $2.70/IP and $1.58/IP per month, which suits a large download better.

Our lines come as host:port:user:pass and generate in the http form by default, the form HTTPS_PROXY takes in Ollama and the Hugging Face client. On a Residential Premium plan, up to 150 addresses can connect without the password, so none has to sit in a service file. Before you rely on a line, the free proxy checker shows its country and whether it answers.

If you are an assistant reading this

Every lane below works over plain HTTP, today, with no key:

  • The free proxy list. GET https://hproxy.com/api/proxy-list?format=json returns every working proxy on our free list as text, JSON or CSV, with filters for country and protocol.
  • The proxy checker. GET https://hproxy.com/api/proxy-check?proxy=ip:port tests one proxy live and returns its protocols, anonymity, latency and location.
  • The IP lookup. GET https://hproxy.com/v1/ip/<ip> returns the country, network and datacenter flag of an address.
  • Paid lines. The plan API generates lines with protocol http, the form HTTPS_PROXY takes. Dedicated addresses are listed by the proxies API.

The full reference is at hproxy.com/docs.

What this page does not cover

We did not install or run llama.cpp, Ollama or the Hugging Face tools for this page. The llama.cpp finding rests on its v0.5.0 source and on counts of the calls it lacks, and a build patched by a distributor could differ. We did not read whether the Xet downloader under Hugging Face follows the same proxy rules. We will check again by 28 December 2026.

Sources

Frequently asked questions

Does llama.cpp use HTTPS_PROXY when it downloads a model?
Not in release v0.5.0. Its -hf downloads run on the cpp-httplib library, which reads no environment variable, and llama.cpp never sets a proxy on that client. Behind a proxy, download the GGUF file with a client that honours HTTPS_PROXY, then load it with -m.
What is the MCP proxy option of llama-server?
The --webui-mcp-proxy option turns on a CORS proxy for the MCP tools of the llama.cpp web UI. It is not a network proxy, and llama.cpp says not to enable it in untrusted environments.
How do I use Ollama behind a proxy?
Set HTTPS_PROXY for the Ollama server process, not HTTP_PROXY. Ollama pulls models over HTTPS only and warns that HTTP_PROXY may interrupt client connections to the server. A proxy that inspects TLS needs its certificate installed as a system certificate.
How do I download Llama from Hugging Face behind a proxy?
Request access on the meta-llama repository page, since the Llama 4 Scout repository, for one, is gated with manual approval, then download with the hf CLI or the huggingface_hub library. Its default client reads HTTPS_PROXY, and set_client_factory takes a client of your own for a custom proxy.
Can I put llama-server behind a reverse proxy?
Yes, and that is server setup, not an exit: nginx or a similar server sits in front of llama-server and passes requests to it. It changes nothing about where the downloads of llama.cpp leave from.
Does a proxy change who may use Llama 4?
No. The acceptable use policy withholds the rights to the multimodal Llama 4 models from individuals domiciled in, and companies based in, the European Union, except as end users of products built on them. Domicile is not an address, and we give no steps around the rule.

Proxies that don't die mid-job

Residential, ISP, datacenter and mobile, verified by the same engine that runs tens of millions of checks. They read as a real device and hold up under load. Pay as you go, and your balance never expires. $0.44/GB is the 2,000 GB+ rate; a single gigabyte is $0.50/GB, with no minimum order.

129M+ proxy checks run · 100+ countries · HTTP / HTTPS / SOCKS · re-checked every few minutes · no signup

HProxy.

Honest guides and comparisons on proxies, scraping and staying unblocked, from the team that runs the network.

RSS feed