Guide

KoboldAI as a Janitor AI Proxy: A Local Model, No Key, No Token Cost, and Why the Tunnel Is the Hard Part

KoboldAI as a Janitor AI proxy: a free local model with no key and no token cost, exposed on an OpenAI endpoint.

HProxy Team··5 min read
HProxy.Guide

Skip the dead lists.

Our free proxy list re-checks every exit every few minutes across 100+ countries, with a live last-checked time, so you copy IPs that worked moments ago, not a stale text dump.

Open the free proxy list→

Every other backend in this series is a company's API: you paste a URL, a key, and a model name, and you pay or hit a cap. KoboldAI is the exception. KoboldCpp runs an open-weights model on your own computer, serves it on an OpenAI-compatible endpoint, and charges nothing per token, because there is no company in the loop. There is no key to leak, no daily cap, no privacy trade, and no bill. It is the most self-contained way to run Janitor AI.

It is also the one with a catch that none of the others have, and the catch is not the model. It is the address: your model answers on localhost, and Janitor does not document where its requests to the Proxy URL come from. This guide deals with that from the start.

Why a tunnel, and what it does

KoboldCpp, once running, serves its OpenAI-compatible API at an address like http://localhost:5001/v1. localhost always means the machine that sends the request. Janitor does not document whether its requests to the Proxy URL leave from your browser or from its own servers, and a request that leaves from a server finds nothing at localhost.

The dependable route is to give your local endpoint a public address that can be reached over the internet from anywhere, which is what a tunnel does. A tunnel opens a connection from your machine out to a public HTTPS URL and forwards anything that arrives there back to your local server. Cloudflare Tunnel and ngrok are the common ones; a reverse proxy on a server you own works too. The public URL the tunnel gives you, ending in /v1/chat/completions, is what goes in Janitor's Proxy URL field, not localhost.

Bridging Janitor to a model on your desk
  1. KoboldCpp on your PC

    OpenAI endpoint at localhost:5001/v1

  2. A tunnel

    Cloudflare Tunnel, ngrok, or your own reverse proxy

  3. A public HTTPS URL

    reachable from anywhere

  4. Janitor's Proxy URL

    the public URL, ending in /v1/chat/completions

  5. Your model answers

    no key, no token cost, chats stay on your machine

Source: KoboldCpp OpenAI-compatible API and standard tunnelling; Janitor AI proxy field

The values

FieldValue
Config NameLocal Kobold
Proxy URLyour tunnel's public URL, ending in /v1/chat/completions
API Keyblank, or a placeholder, or your tunnel password if you set one
Modelthe model KoboldCpp loaded (its name as the endpoint reports it)

KoboldCpp has no API key by default, which is why the field can be empty. If you protected the tunnel with a password, put that password in the field. Save, refresh, and test as with any backend; a network error here almost always means the tunnel is down or the URL is missing /chat/completions, per the pattern in our network error guide.

What you get, and what you give up

The appeal is real and worth stating plainly.

  • No cost per token. Nothing to fund, no cap to hit, no 429 from a provider. The only meter is your electricity bill.
  • Privacy. Your messages go from Janitor to your own machine and back. No provider stores them, and there is no key that a leak could expose. This is the strongest privacy position of any backend, stronger than a paid API, because there is no third party at the model at all.
  • No provider policy. An open-weights model runs what it runs; there is no external content policy declining scenes. What the model was trained to do is what it does.

The costs are equally real.

  • Setup is work. You install KoboldCpp, download a model file, and stand up a tunnel. SillyTavern's own documentation, describing the same class of local backend, notes that the installation can be complex and that its team does not provide support for it. Budget an afternoon the first time.
  • Your hardware is the ceiling. The model you can run is the one your GPU or RAM can hold, which is smaller than the frontier models on OpenRouter. For many roleplayers a good mid-size local model is plenty; for others the quality gap is the reason they pay. That is a taste call, and your machine decides the menu.
  • The machine has to stay on. Close the laptop or lose the connection and the tunnel drops and the endpoint dies mid-chat. A desktop left running is the natural home for this.

Where a network proxy fits, which is almost nowhere

This is the backend where our own product has the least to do. The model is on your machine; the bridge is a tunnel; neither is something an IP-changing proxy touches. Janitor AI has one rule that follows the address. Users connecting from Australia, Brazil or the UK must pass a one-time age check, and Janitor says "Detection is IP-based only." Before that check, it blocked UK addresses from July 2025 until June 2026. If janitorai.com is blocked on your network or in your country, we give no steps around that.

Where our lines do fit is a network that needs a known exit, such as a company network, with a static residential or ISP address in the country where you are, from $0.44 per GB on a usage meter. For the model, the tunnel and the cost, a network proxy is irrelevant, and we would rather say so than imply a local setup needs us.

If you want the frontier-model quality that a local model cannot match, the free hosted routes are the alternative, at the cost of a cap and a provider policy: OpenRouter's :free models and Google's Gemini free tier, both in best free Janitor AI proxy. The local route trades that quality for cost, privacy and no cap, and for a lot of people that is the trade worth making.

Sources

Frequently asked questions

Can Janitor AI use a local KoboldAI model?
Yes, with a bridge. KoboldCpp runs an open-weights model on your machine and serves an OpenAI-compatible endpoint, usually at http://localhost:5001/v1, with no API key and no per-token cost. The catch is the address: localhost means the machine that sends the request, and Janitor does not say where its requests come from. You expose the local endpoint through a tunnel with a public HTTPS address, then paste that address into Janitor's Proxy URL. The model is free; the tunnel is the work.
Is a local KoboldAI model really free?
Free of per-token cost, yes. There is no API key to buy, no rate limit set by a provider, and no daily cap: the only limits are your hardware and your electricity. It is also the most private option, because your chats never leave your machine to reach the model. What you trade is convenience and quality: you install and maintain it, and the model your GPU can hold is smaller than the frontier models on OpenRouter.
Why use a tunnel instead of localhost?
Because localhost always means the machine that sends the request, and Janitor does not document whether that is your browser or its own servers. A tunnel such as Cloudflare Tunnel or ngrok, or a reverse proxy on a server you own, gives your endpoint a public HTTPS address that answers from anywhere. That public address is what goes in Janitor's Proxy URL.
Does KoboldAI need an API key on Janitor AI?
No. A local KoboldCpp server has no key by default, so the API Key field can be left blank or set to any placeholder. That is the whole appeal alongside the cost: nothing to sign up for, nothing to fund, nothing to leak. If you add your own password to the tunnel, that password goes in the field instead.
Do I need a proxy or VPN for KoboldAI with Janitor AI?
Not the network kind, and this is the one backend where our product is almost entirely beside the point. The model is on your machine and the bridge is a tunnel, neither of which an IP-changing proxy touches. If janitorai.com is blocked on your network or in your country, we give no steps around that.
What are the downsides of the local route?
Three. Setup is real work: install KoboldCpp, download a model, and stand up a tunnel, which SillyTavern's own documentation notes is complex and unsupported. Quality is capped by your GPU, so the model is smaller than a frontier one. And your machine has to stay on and connected for the whole chat, because when it sleeps, the endpoint dies. In exchange you get free, private, uncapped roleplay.

Get proxies that are alive right now

Our free proxy list re-checks every exit every few minutes across 100+ countries, with a live last-checked time, so you copy IPs that worked moments ago, not a stale text dump. When the location has to survive a real check, the paid network holds up.

129M+ proxy checks run · 100+ countries · HTTP / HTTPS / SOCKS · re-checked every few minutes · no signup

HProxy.

Honest guides and comparisons on proxies, scraping and staying unblocked, from the team that runs the network.

RSS feed