Hanzo LLM
Call many LLM providers in the OpenAI format, from Python or through a proxy.
Forked from BerriAI/litellm (MIT).
Install
Python 3.10 or newer.
pip install hanzo-llm # the Python library
pip install 'hanzo-llm[proxy]' # the library and the proxy server
It imports as llm and puts an llm command on the path. Python 3.9 gets
1.81.15, the last release that allowed it, and that release's [proxy] extra
cannot install.
First call
completion takes a model name prefixed with its provider and answers in the
OpenAI response shape. This call goes to a local Ollama with qwen3:0.6b
pulled:
from llm import completion
r = completion(
model="ollama_chat/qwen3:0.6b",
api_base="http://127.0.0.1:11434",
messages=[{"role": "user", "content": "Say hello in five words."}],
)
print(r.choices[0].message.content)
The hosted Hanzo API takes the same call with model="openai/<model id>",
api_base="https://api.hanzo.ai/v1" and a Hanzo API key as api_key. Model
ids are listed at https://api.hanzo.ai/v1/models.
Proxy
The proxy serves the same format over HTTP. Name each model and where it runs in a config file:
model_list:
- model_name: qwen3
litellm_params:
model: ollama_chat/qwen3:0.6b
api_base: http://127.0.0.1:11434
Start it, then call it:
llm --config config.yaml --port 4000
curl -s http://127.0.0.1:4000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model": "qwen3", "messages": [{"role": "user", "content": "Say hello in five words."}]}'
This config sets no key, so the proxy answers any caller.
Use Hanzo LLM for
Agents - Invoke A2A Agents (Python SDK + AI Gateway)
Supported Providers - LangGraph, Vertex AI Agent Engine, Azure AI Foundry, Bedrock AgentCore, Pydantic AI
Python SDK - A2A Protocol
from llm.a2a_protocol import A2AClient
from a2a.types import SendMessageRequest, MessageSendParams
from uuid import uuid4
client = A2AClient(base_url="http://localhost:10001")
request = SendMessageRequest(
id=str(uuid4()),
params=MessageSendParams(
message={
"role": "user",
"parts": [{"kind": "text", "text": "Hello!"}],
"messageId": uuid4().hex,
}
)
)
response = await client.send_message(request)
AI Gateway (Proxy Server)
Step 1. Add your Agent to the AI Gateway
Step 2. Call Agent via A2A SDK
from a2a.client import A2ACardResolver, A2AClient
from a2a.types import MessageSendParams, SendMessageRequest
from uuid import uuid4
import httpx
base_url = "http://localhost:4000/a2a/my-agent" # Hanzo LLM gateway + agent name
headers = {"Authorization": "Bearer sk-1234"} # Hanzo LLM Virtual Key
async with httpx.AsyncClient(headers=headers) as httpx_client:
resolver = A2ACardResolver(httpx_client=httpx_client, base_url=base_url)
agent_card = await resolver.get_agent_card()
client = A2AClient(httpx_client=httpx_client, agent_card=agent_card)
request = SendMessageRequest(
id=str(uuid4()),
params=MessageSendParams(
message={
"role": "user",
"parts": [{"kind": "text", "text": "Hello!"}],
"messageId": uuid4().hex,
}
)
)
response = await client.send_message(request)
MCP Tools - Connect MCP servers to any LLM (Python SDK + AI Gateway)
Python SDK - MCP Bridge
from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client
from llm import experimental_mcp_client
import llm
server_params = StdioServerParameters(command="python", args=["mcp_server.py"])
async with stdio_client(server_params) as (read, write):
async with ClientSession(read, write) as session:
await session.initialize()
# Load MCP tools in OpenAI format
tools = await experimental_mcp_client.load_mcp_tools(session=session, format="openai")
# Use with any Hanzo LLM model
response = await llm.acompletion(
model="gpt-4o",
messages=[{"role": "user", "content": "What's 3 + 5?"}],
tools=tools
)
AI Gateway - MCP Gateway
Step 1. Add your MCP Server to the AI Gateway
Step 2. Call MCP tools via /chat/completions
curl -X POST 'http://0.0.0.0:4000/v1/chat/completions' \
-H 'Authorization: Bearer sk-1234' \
-H 'Content-Type: application/json' \
-d '{
"model": "gpt-4o",
"messages": [{"role": "user", "content": "Summarize the latest open PR"}],
"tools": [{
"type": "mcp",
"server_url": "llm_proxy/mcp/github",
"server_label": "github_mcp",
"require_approval": "never"
}]
}'
Use with Cursor IDE
{
"mcpServers": {
"HanzoLLM": {
"url": "http://localhost:4000/mcp/",
"headers": {
"x-llm-api-key": "Bearer sk-1234"
}
}
}
}
Supported Providers (Docs)
Run in developer mode
Start the database and Prometheus from the repo root; .env.example lists the
settings:
docker compose up db prometheus
Then run the proxy from a virtual environment:
python -m venv .venv && source .venv/bin/activate
pip install -e '.[proxy]'
pip install prisma
prisma generate
python llm/proxy/proxy_cli.py
Contributing
We welcome contributions to Hanzo LLM! Whether you're fixing bugs, adding features, or improving documentation, we appreciate your help.
Quick Start for Contributors
This requires poetry to be installed.
git clone https://github.com/hanzoai/llm.git
cd llm
make install-dev # Install development dependencies
make format # Format your code
make lint # Run all linting checks
make test-unit # Run unit tests
make format-check # Check formatting only
For detailed contributing guidelines, see CONTRIBUTING.md.
Code Quality / Linting
Hanzo LLM follows the Google Python Style Guide.
Our automated checks include:
- Black for code formatting
- Ruff for linting and code quality
- MyPy for type checking
- Circular import detection
- Import safety checks
All these checks must pass before your PR can be merged.
Support
Contributors
Metadata
Release files for hanzo-llm 1.82.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| hanzo_llm-1.82.2.tar.gz | 13.4 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| hanzo_llm-1.82.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 28.3 MB
Release files / hanzo_llm-1.82.2.tar.gz
| Download URL | hanzo_llm-1.82.2.tar.gz |
|---|---|
| Size | 13.4 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
3fd8335f73361a3c4775e18c20bfb477de734aece0db77131b2cd044d333dabf
|
|
BLAKE2b-256 checksum How to use checksums |
b5ebc7c92e8953eb0275a71046e928dd727d7528d24e105271c55192680efda3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.13
|
Release files / hanzo_llm-1.82.2-py3-none-any.whl
| Download URL | hanzo_llm-1.82.2-py3-none-any.whl |
|---|---|
| Size | 14.9 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a63e2a7e001c005bb8df80940ccacc335edb77e9d5bb6810394402cf1a9b8a10
|
|
BLAKE2b-256 checksum How to use checksums |
17948fba65e7e261a7b6bcc6b1eb5fa1cd4064100671489a18940e99487ac629
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.13
|