Skip to main content

local-llm-router

A Python routing library for local LLM agent loops.

Give local-llm-router a prompt and your model list. It returns a complexity tier and which model to call. You keep your agent runner, gateway, or Ollama client — local-llm-router only decides which local model each step should use.

Zero runtime dependencies. Works offline. No inference, no agent framework, no chat UI.

Install

pip install local-llm-router
pip install "local-llm-router[ollama]"   # optional: Ollama discovery, llm-router ask

Quick start

import local_llm_router

local_llm_router.configure(vram_gb=16, quant="qat")  # once — or local_llm_router_VRAM_GB=16

for step in agent_steps:
    tier, model = local_llm_router.route(step.prompt, hint=step.hint)
    response = your_llm.complete(model=model, prompt=step.prompt)

Power-user path (explicit tier map):

from local_llm_router import assign_tiers, route_prompt

tiers = assign_tiers(["gemma4:e4b", "qwen3:8b", "qwen3:14b"])
tier, model = route_prompt(step.prompt, tiers, hint=step.hint)

Check your stack before routing:

llm-router doctor --check-stack --vram-gb 16 --quant qat

Integration guide: docs/for-app-authors.md · docs/integration.md

What it does

Piece Role
configure(vram_gb=..., quant=...) Set GPU budget + Gemma quant assumption once
route(prompt, hint=...) Pick (tier, model_name) for one agent step
assign_tiers / route_prompt Same routing with your own tier map
llm-router compare / stack benchmark CLI evidence — dry by default

Step hints: lookup, explain, design, code, reason — override keyword scoring when you know the agent phase.

VRAM and quant

local_llm_router.configure(vram_gb=16, quant="qat")
VRAM Profile
≤8 / 12 / 16 / 24 / 32 GB workstation_8gb … workstation_32gb

quant= is not per-prompt routing — it tells local-llm-router which pull format you use so VRAM filters and stack suggestions stay honest (QAT vs default Ollama Q4). Details: docs/local-models.md.

What it is not

Routing primitives only — not a chat app, agent framework, or multi-cloud proxy. Optional browser demo and VS Code panel are examples, not the product.

Examples (optional)

Clone the repo only if you want demos or to contribute:

Example Command
Agent runner python examples/agent_runner/run.py
Compare POC llm-router compare
Browser demo python examples/demo_ui/server.py → http://127.0.0.1:8765
Quickstart tour examples/quickstart/try_it.ps1

See examples/*/README.md for each.

Contributors

git clone https://github.com/edwardjbaumel/local-llm-router.git
cd local-llm-router
pip install -e ".[dev,ollama]"
pytest
llm-router setup --profile 12gb --dry-run

Docs: docs/publishing.md · docs/security.md · docs/backlog.md

Local Recruiting Ops — separate project by the same author; local-llm-router does not depend on it.

Metadata

Release files for local-llm-router 0.4.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for local-llm-router 0.4.1
File Size Uploaded
local_llm_router-0.4.1.tar.gz 52.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for local-llm-router 0.4.1
File Interpreter ABI Platform
local_llm_router-0.4.1-py3-none-any.whl Python 3 none any Details

Total release size: 106.1 kB

Release files / local_llm_router-0.4.1.tar.gz

Download URL local_llm_router-0.4.1.tar.gz
Size 52.2 kB
Tags Source
SHA-256 checksum
How to use checksums
2f3cd05d56235b7a02f81a8847243355488773bc14b7541e1959ab0f36b784d7
BLAKE2b-256 checksum
How to use checksums
49477951fead641628cefd1f79e35a4dbea435f0bdf4e51de56a4f1b54bc5ac9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.10

Release files / local_llm_router-0.4.1-py3-none-any.whl

Download URL local_llm_router-0.4.1-py3-none-any.whl
Size 53.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
65fa3837ffa045283f6b0240818c09dd8a47fe2d729f48e5b12fda1cb51ae7c1
BLAKE2b-256 checksum
How to use checksums
8792045f3f67060ff9edc0a9bd4950cfb9231184f6534a0dff6aa704bffcfe35
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.10

Release history Release notifications | RSS feed

This release

0.4.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page