A local-first API gateway for open-source LLMs: auth, RAG memory extension, DB connectors, token accounting, and a built-in coding agent.
Project description
localgate
Turn any local LLM into a managed API — real API keys, token accounting, and RAG memory that makes a small model remember far more than its context window holds.
The built-in dashboard at /dashboard/ — token spend, latency, and per-model
usage across every key. Shown with sample data.
Why
Ollama, LM Studio and LocalAI solve model serving. They deliberately don't solve anything around it:
- No API key management. No per-user keys, no revocation, no usage tracking.
- No memory past the context window. Your 8K model forgets everything beyond 8K tokens.
- No token accounting. You guess at what you've spent.
- No database story. You wire up Postgres yourself.
localgate is the management layer. It sits between your app and your inference server and adds all four — without touching how you serve models.
Installation
localgate is on PyPI. Pick whichever of these you already have — you don't need all three:
uv tool install localgate # no Python/pip setup needed if you have uv
# or
pipx install localgate # isolated install, doesn't need a venv
# or
pip install localgate # into your current environment/venv
All three put a localgate command on your PATH. If pip or pipx say "command not
found": macOS doesn't ship them by default — pip3 (from python3 -m ensurepip or
Homebrew's python) and pipx (brew install pipx) both need to be installed first. If
you don't already have Python tooling set up, uv tool install is the path of least
resistance: install uv (one
command, no Python required first), then uv tool install localgate.
A container image is also published, at ghcr.io/anjalllll/localgate.
Quick start
ollama serve # your inference backend
ollama pull llama3
ollama pull nomic-embed-text # enables RAG memory
localgate db upgrade
localgate keys create --name my-app # prints your key, once
localgate serve
Developing localgate itself, instead of just using it:
git clone https://github.com/AnjalLLL/localgate.git && cd localgate
uv sync --all-extras
uv run localgate db upgrade
uv run localgate serve
Now use it like OpenAI, because it is the OpenAI API:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="lg_9f3a...")
response = client.chat.completions.create(
model="llama3",
messages=[{"role": "user", "content": "Hello!"}],
)
Full walkthrough: Getting Started.
The memory bit
This is the part that isn't a proxy. Send an X-Session-ID and the gateway stores each
turn, embeds it, and retrieves what's relevant on later turns:
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="lg_9f3a...",
default_headers={"X-Session-ID": "conversation-1"},
)
client.chat.completions.create(
model="llama3",
messages=[{"role": "user", "content": "My name is Ana and I prefer Postgres."}],
)
# A separate request. No history sent. The model still knows.
client.chat.completions.create(
model="llama3",
messages=[{"role": "user", "content": "What database do I prefer?"}],
)
# → "You prefer Postgres."
The model answers correctly not because you sent the history, but because the gateway retrieved it. Past a threshold, older turns are folded into a rolling summary, so the context window holds the useful part of a long conversation rather than the most recent part of it. See RAG Memory.
Features
- OpenAI-compatible — works with any OpenAI SDK, LangChain, or curl. Unknown fields are forwarded to the backend, so your sampling knobs keep working.
- API key management — create, revoke, and rate-limit per key. Hashed, never stored raw.
- RAG memory — automatic chunking, embedding, retrieval, and rolling summarization.
- Token accounting — prompt/completion tokens per key, per model, over time.
- Any database — SQLite with zero config; Postgres or Neon with one env var.
- Any backend — Ollama, vLLM, llama.cpp, or any OpenAI-compatible server. Third parties can add more via an entry point, no fork required.
- Model aliasing — map
fast→phi4-miniand swap models without touching clients. - Prompt caching — opt-in; identical prompts skip inference entirely.
- Production-ready — structured JSON logs with correlation IDs, Prometheus metrics, liveness/readiness split, graceful shutdown, fail-fast config validation.
- Dashboard — keys, usage and conversations in the browser, at
/dashboard/.
Dashboard
Served at /dashboard/ — no build step, no separate deploy. It talks to the same /admin
API the CLI does, so anything it can do is equally scriptable.
Create and revoke keys, watch token spend per model, browse stored conversations with their rolling summaries, and point the gateway at a new database — with the connection tested before it is saved.
CLI
localgate serve # start the gateway
localgate health # is the backend up? is the DB migrated?
localgate backends # what adapters are installed
localgate keys create --name my-app # create a key (printed once)
localgate keys list # every key, active and revoked
localgate keys revoke <id> # revoke (history is kept)
localgate keys usage <id> # token usage for one key
localgate db upgrade # apply migrations
localgate db current # current schema revision
localgate code # interactive coding session in the current directory
localgate code "add a health check" # one-off task, then exit
The CLI talks to the database (and, for code, the inference backend) directly, not to a
running server — because keys create has to work before you have a key, and db upgrade
has to work when the server won't start.
Shell completion: localgate --install-completion (bash/zsh/fish/PowerShell, via Typer).
localgate code
A minimal coding agent that reads and edits files in the current project, backed by whatever
model localgate is already pointed at — no separate API key needed, since it talks to the
backend directly rather than through the gateway.
localgate code # REPL: /exit, /clear, /model <name>, /undo
localgate code "fix the off-by-one in parser.py" # one-shot
localgate code "..." --auto-approve --auto-commit # unattended, with every write committed
- Every write is shown as a colored diff and asks for confirmation, unless
--auto-approve. - On a dirty working tree, it warns once before writing anything (
--forceto skip). --auto-commitcommits each turn's writes, taggedlocalgate-agent:;/undoin the REPL reverts the last write (or the last agent commit, with--auto-commit) via git.- Tools:
read_file,write_file,list_directory,search_files(grep-like),git_status,git_diff. All confined to the project directory;.gitignoreand.localgateignorekeep secrets and generated directories out of the model's reach. No shell/run_commandtool. - Conversation history and recalled context persist per project (
.localgate/session_id), reusing the same RAG memory layer as the HTTP API — re-running it in a project you worked on before picks up where you left off. Disable with--no-memoryorLOCALGATE_MEMORY_ENABLED=false.
Tool-calling quality depends entirely on the model — verify yours actually returns structured tool calls (not JSON printed as text) before relying on this day to day.
Documentation
| Getting Started | Zero to working gateway |
| Configuration | Every setting |
| API Reference | Every endpoint |
| Database Setup | SQLite → Postgres → Neon |
| RAG Memory | How memory works, and how to tune it |
| Architecture | How it's built, and why |
| Deployment | Running it somewhere real |
| Decisions | ADRs for the choices that shaped the codebase |
Contributing
See CONTRIBUTING.md. Adding a backend means writing one class. Issues
tagged good-first-issue are a good place to start.
License
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file localgate-0.7.1.tar.gz.
File metadata
- Download URL: localgate-0.7.1.tar.gz
- Upload date:
- Size: 431.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
58774592b6968976981ecfc2c3a1b16aae56dfae51f3979dd5e07be248b489fb
|
|
| MD5 |
ad473f9ac38573be5b071d35484a5cb8
|
|
| BLAKE2b-256 |
98d453dde49380de51fd66cc596600d2a8aaacab0e17c37143cf242e0c3776a4
|
Provenance
The following attestation bundles were made for localgate-0.7.1.tar.gz:
Publisher:
release.yml on AnjalLLL/localgate
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
localgate-0.7.1.tar.gz -
Subject digest:
58774592b6968976981ecfc2c3a1b16aae56dfae51f3979dd5e07be248b489fb - Sigstore transparency entry: 2195017816
- Sigstore integration time:
-
Permalink:
AnjalLLL/localgate@53f921d2c8cefe80263a3d107444f9c12ac1f4cb -
Branch / Tag:
refs/tags/v0.7.1 - Owner: https://github.com/AnjalLLL
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@53f921d2c8cefe80263a3d107444f9c12ac1f4cb -
Trigger Event:
push
-
Statement type:
File details
Details for the file localgate-0.7.1-py3-none-any.whl.
File metadata
- Download URL: localgate-0.7.1-py3-none-any.whl
- Upload date:
- Size: 113.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4832b8465151955b43750d678a6bac17d74612848b1dce00caab87e7c0e6b398
|
|
| MD5 |
cd3546d25eabbb958e4b931417a09204
|
|
| BLAKE2b-256 |
f11e4573322d4f6f479eb7b51ce59425fcf22d19493287c86c7f58fb702e237f
|
Provenance
The following attestation bundles were made for localgate-0.7.1-py3-none-any.whl:
Publisher:
release.yml on AnjalLLL/localgate
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
localgate-0.7.1-py3-none-any.whl -
Subject digest:
4832b8465151955b43750d678a6bac17d74612848b1dce00caab87e7c0e6b398 - Sigstore transparency entry: 2195017841
- Sigstore integration time:
-
Permalink:
AnjalLLL/localgate@53f921d2c8cefe80263a3d107444f9c12ac1f4cb -
Branch / Tag:
refs/tags/v0.7.1 - Owner: https://github.com/AnjalLLL
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@53f921d2c8cefe80263a3d107444f9c12ac1f4cb -
Trigger Event:
push
-
Statement type: