sandhi-gateway
Python binding for Sandhi — the metering layer for AI agents. The Rust core, in-process via PyO3: virtual keys, budgets, and neutral usage-event metering with zero network hop. Keep making your own provider calls; hand the response to Sandhi to meter it.
pip install sandhi-gateway # import as: import sandhi_gateway
The bare name
sandhion PyPI is an unrelated Sanskrit-linguistics library; this binding is published assandhi-gateway. The crate and GitHub repo aresandhi.
Usage
import json
import sandhi_gateway as sg
gw = sg.Gateway(sink_path="usage.jsonl") # events append as JSONL (+ in-memory)
gw.add_virtual_key("vk_alice", subject="alice", group="platform", upstream="anthropic")
gw.set_budget("group:platform", 1_000_000)
# ... you make your own provider call and get the raw response JSON ...
event = gw.meter(
"vk_alice", "anthropic", "claude-x", response_json,
session_id="conv_7",
)
# event["tokens_in"], event["cache_read_tokens"], event["subject_id"], ...
print(gw.spent("group:platform")) # budget recorded
print(gw.check_budget("group:platform", 5000)) # True/False
# Just parse usage (same Rust parsers as the proxy), no attribution:
sg.parse_usage("openai", response_json) # {tokens_in, tokens_out, cache_*}
Typed persistent provider runtime (0.1.2+)
New integrations should reuse a typed provider handle. Its inputs and outputs are Sandhi's versioned neutral chat documents; provider-native JSON is encoded and decoded in Rust.
runtime = sg.ProviderRuntime()
provider = runtime.provider("openrouter", "openai/gpt-4o", api_key)
request = {"model": "openai/gpt-4o", "messages": [{"role": "user", "content": "hello"}]}
response = json.loads(await provider.complete_json(json.dumps(request)))
async for event_json in provider.stream_json(json.dumps(request)):
event = json.loads(event_json) # response_start, text_delta, tool_call_*, usage, finish
The JSON bridge is ABI-stable typed v1 data, not provider-native JSON. The handle retains its HTTP
pool, circuit breaker, retry policy, and timeouts. Invalid documents fail before network I/O;
runtime failures use the serialized ProviderErrorV1 shape in the exception message.
runtime.provider() resolves a known endpoint from Sandhi's catalog;
runtime.openai_compat() is the explicit custom-endpoint escape hatch.
Legacy provider-native transport (0.1.2+)
Sandhi also owns the provider wire layer: endpoint routing, headers, HTTP/SSE, resilience, wire errors, and neutral usage extraction. Callers keep model policy, prompt/tool assembly, and their framework-facing response types.
import asyncio
import json
import sandhi_gateway as sg
async def main():
api_key = "..."
spec = sg.provider_spec("kimi", model="kimi-k3")
body = {"model": "kimi-k3", "messages": [{"role": "user", "content": "hello"}]}
result = await sg.complete(
spec["slug"], "kimi-k3", spec["base_url"], api_key, json.dumps(body),
max_retries=3,
)
# result = {"status": ..., "body": raw_json, "usage": neutral_cache_split}
openrouter_model = "meta-llama/llama-3.3-70b-instruct"
openrouter_body = {**body, "model": openrouter_model}
async for item in sg.stream(
"openrouter", openrouter_model, sg.provider_spec("openrouter")["base_url"], api_key,
json.dumps(openrouter_body), max_retries=3,
headers_json=json.dumps({"HTTP-Referer": "https://example.app", "X-Title": "My App"}),
):
print(item["data"])
# The terminal item carries finalized neutral usage.
asyncio.run(main())
provider_spec() exposes stable Rust-owned wire facts (canonical slug, aliases,
base URL, and model endpoint routing), not a model/capability catalog. Custom
Authorization, Content-Type, and Host values are ignored so callers cannot
override transport-owned headers.
The OpenAI-compatible transport accepts the Chat Completions roles developer,
system, user, assistant, tool, and legacy function. A tool result must
carry its tool_call_id; a legacy function result must carry name. Sandhi
validates these wire invariants before HTTP but deliberately does not rewrite roles:
whether a specific compatible model accepts developer, for example, is caller-owned
model policy.
Custom / unknown providers (host escape hatch)
# (a) register a host parser callback for a provider Sandhi doesn't know:
gw.register_parser("myprovider", lambda body: {"tokens_in": 30, "tokens_out": 12,
"cache_creation_tokens": 0, "cache_read_tokens": 0})
gw.meter("vk_alice", "myprovider", "model", response_json) # uses your callback
# (b) or skip parsing and pass counts directly:
gw.meter_tokens("vk_alice", "myprovider", "model", tokens_in=30, tokens_out=12)
meter() parses the usage at the source (the same cache-split logic as the reverse
proxy), attributes it to the virtual key's subject/group, records the budget, emits the
neutral usage event (matching usage-event.v1.schema.json),
and returns it for local display. Unknown key → KeyError; bad JSON → ValueError.
Usage snapshots (in-process aggregation)
import json
rows = json.loads(gw.usage_snapshot_json("subject")) # busiest subject first
rows[0]["billable_tokens"] # the quantity budgets enforce on
json.loads(gw.usage_snapshot_json("total"))[0] # one grand-total row
json.loads(gw.usage_snapshot_json("session", 256)) # bound distinct keys to 256
Folds the events recorded so far into
usage-aggregate.v1
rows for one dimension — subject (user), group, provider, model, key
(virtual_key), session, or total — using the same fold the reverse proxy, the
sandhi CLI, and the dashboard read. Neutral units only, never dollars. The optional
second argument caps distinct keys (default 1024); everything past it folds into a single
"(overflow)" row, so a long-lived process loses per-key detail but never the sum.
Unknown dimension → ValueError.
Apache-2.0. See the main README and ADR-0001.
Release files for sandhi-gateway 0.2.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distributions (wheels)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| sandhi_gateway-0.2.1-cp311-abi3-win_amd64.whl | CPython 3.11 | abi3 | Windows x86-64 | Details |
| sandhi_gateway-0.2.1-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl | CPython 3.11 | abi3 | Linux glibc 2.17+ x86-64 | Details |
| sandhi_gateway-0.2.1-cp311-abi3-macosx_11_0_arm64.whl | CPython 3.11 | abi3 | macOS 11.0+ ARM64 | Details |
Total release size: 7.6 MB
Release files / sandhi_gateway-0.2.1-cp311-abi3-win_amd64.whl
| Download URL | sandhi_gateway-0.2.1-cp311-abi3-win_amd64.whl |
|---|---|
| Size | 2.4 MB |
| Tags | CPython 3.11 Windows x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
193826aedb610f9deadd0b19467b3f2c7961b164f8b773579037afef26850da6
|
|
BLAKE2b-256 checksum How to use checksums |
6dfdc9f1849410cbf218b4f8544da565bd788f3501c6099f0e85a5fc58b491ef
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 1, 2026.
Transparency logRelease files / sandhi_gateway-0.2.1-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
| Download URL | sandhi_gateway-0.2.1-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl |
|---|---|
| Size | 2.7 MB |
| Tags | CPython 3.11 Linux glibc 2.17+ x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
7009a3a3ec81c4d53e83ca70d7d0568e733b2de50a8e61bc1986b59e297f9a12
|
|
BLAKE2b-256 checksum How to use checksums |
18b6e4dfba666847e27b0b3afa38a7a8979eb7d74b748dafbf04cebf4768b27d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 1, 2026.
Transparency logRelease files / sandhi_gateway-0.2.1-cp311-abi3-macosx_11_0_arm64.whl
| Download URL | sandhi_gateway-0.2.1-cp311-abi3-macosx_11_0_arm64.whl |
|---|---|
| Size | 2.5 MB |
| Tags | CPython 3.11 abi3 macOS 11.0+ ARM64 |
|
SHA-256 checksum How to use checksums |
b9195f6a3efe0daca479fbbfb9adbe99f32d5e300e309ed68b6da6028406a380
|
|
BLAKE2b-256 checksum How to use checksums |
4612850743d024c8c5700f5693b49d2fde92785bbbafda3845a31f2607212b70
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 1, 2026.
Transparency log