neutral-llm-gateway
A small, honest gateway for LLM calls: typed contracts, thin provider adapters, and accounting of retries, fallbacks, tokens and cost that refuses to lie to you.
It knows about providers, never about products. It holds no credentials, reads no environment variables, ships no prompts and stores no business schemas. Everything product-shaped — ledgers, tenants, alerting, history, prompts — stays in your application, wired in through optional ports.
your application → your facade → llm_gateway → provider SDK
Never the reverse.
Why this exists
It was extracted from several applications that had each grown their own version of the same "call an LLM" function — the largest close to two thousand lines, mixing provider calls, retry policy, cost maths, a usage ledger and business alerting in one place. Writing that from scratch a fourth time is how subtle accounting bugs get copied around.
The bugs it is built to prevent are all the same shape: a number that looks like a fact but isn't.
- A provider returns no usage, the code records
0tokens, and the call bills as free. - A model is missing from the price table, so its cost is
USD 0.00— indistinguishable from a genuinely free call. - A retry fails after the model produced tokens, and only the successful attempt gets counted.
- A fallback quietly answers with a different model, and the metrics attribute it to the one you asked for.
Here, unreported usage is None, unknown cost is UNAVAILABLE, every billable
attempt is counted, and a fallback is never silent.
Is this for you?
Probably yes if you want a thin, auditable layer you can read in an afternoon, you want to own your credentials, you care about per-call cost being reconcilable against an invoice, and you prefer typed results to dictionaries.
Probably not if you want routing across dozens of providers, a proxy server, streaming, or an agent framework. LiteLLM and friends cover far more surface than this does. This covers deliberately less, and is explicit about what it does not know.
Install
Provider SDKs are optional extras. Install only what you call.
# uv
uv add "neutral-llm-gateway[gemini]"
# pip
pip install "neutral-llm-gateway[gemini]"
Available extras: openai, gemini, groq, assemblyai, openrouter,
replicate, wavespeed, all. Combine them as [openai,assemblyai].
openrouter installs the openai SDK, and assemblyai and wavespeed
install the small HTTP transport used by their REST adapters.
Importing the package with no extra installed works by design; asking for a provider you have not installed raises a typed error naming the exact extra.
To try an unreleased commit, the PEP 508 git form works everywhere and pins a tag or revision:
pip install "neutral-llm-gateway[gemini] @ git+https://github.com/jmgb/llm-gateway-python.git@v0.5.0"
Use
from pydantic import BaseModel
from llm_gateway import (
FallbackPolicy,
LLMGateway,
LLMRequest,
Message,
ResponseFormat,
RetryPolicy,
)
from llm_gateway.factories import build_registry, create_gemini_client
class Answer(BaseModel):
verdict: str
# You build the client, so you keep the key. Prices come from the built-in
# versioned catalogue unless you pass your own.
gateway = LLMGateway(
registry=build_registry(gemini_client=create_gemini_client(api_key=my_key)),
)
result = await gateway.generate(
LLMRequest(
model="gemini-3.5-flash-lite",
system_prompt="Answer strictly from the supplied evidence.",
messages=(Message("user", question),),
response_format=ResponseFormat.JSON_SCHEMA,
response_schema=Answer,
temperature=0,
retry_policy=RetryPolicy.transient(max_attempts=2),
fallback_policy=FallbackPolicy.disabled(),
request_id=request_id,
source="my-feature",
)
)
result.output # Answer instance — no metadata mixed in
result.usage.input_tokens # None means "not reported", not zero
result.cost.amount_usd # None when unavailable, never a fake 0
result.cost.measurement # ACTUAL | ESTIMATED | UNAVAILABLE
result.execution.model_used # provider-reported model that actually answered
result.execution.fallback_used # whether the gateway used its fallback plan
result.execution.attempts # every attempt, including the failed ones
If the model answers with something that is not valid JSON, or with JSON that
violates Answer, that attempt is recorded as failed and billed — the
tokens were spent — and the next model in fallback_policy is tried. When no
model produces a usable answer the call raises AllAttemptsFailed, carrying
every attempt, with the parsing or schema error as its __cause__.
Schema validation errors name each Pydantic loc and type, while dynamic
response keys and values stay out of the message.
Transcription
Speech-to-text is a separate operation with duration usage and audio pricing:
from llm_gateway import AudioInput, LLMGateway, TranscriptionRequest
from llm_gateway.factories import build_registry, create_openai_client
gateway = LLMGateway(
registry=build_registry(openai_client=create_openai_client(api_key=my_key)),
)
transcript = await gateway.transcribe(
TranscriptionRequest(
model="gpt-transcribe",
audio=AudioInput(
data=audio_bytes,
filename="voice.webm",
mime_type="audio/webm",
duration_seconds=12.5,
),
language="es",
source="voice-note",
)
)
transcript.text
transcript.usage.duration_seconds
transcript.cost.amount_usd # audio minutes, never token pricing
language is optional and defaults to provider detection. Use
assemblyai-universal-3-pro or assemblyai-universal-2 with a public
AudioInput.url, and whisper-large-v3-turbo/whisper-large-v3 with Groq.
AssemblyAI Universal-3 Pro supports prompt and speaker labels; OpenAI and
Groq reject speaker labels rather than ignoring them. Fallbacks are explicit
through FallbackPolicy.models_in_order(...).
Image generation
Images are a separate operation too, for the same reason: the output is bytes or a URL, and providers bill it either per image or with image-output tokens whose rate can differ from text output.
from llm_gateway import ImageInput, ImageRequest, LLMGateway
from llm_gateway.factories import build_registry, create_gemini_client
gateway = LLMGateway(
registry=build_registry(gemini_client=create_gemini_client(api_key=my_key)),
)
result = await gateway.generate_image(
ImageRequest(
model="gemini-3.1-flash-image",
prompt="a cat wearing a hat, studio lighting",
image=ImageInput(data=original_bytes, mime_type="image/jpeg"), # optional: an edit
source="whatsapp-image",
)
)
result.images[0].data # bytes from Gemini; `.url` from Replicate and WaveSpeed
result.usage.images
result.cost.amount_usd # per image, or per token where the model bills that way
Which form comes back is the provider's, not a choice: Gemini returns inline bytes, Replicate and WaveSpeed return a URL. Downloading it, re-hosting it, watermarking it and enforcing a per-user quota stay in the application, exactly as decoding audio does.
Editing needs a source image in the form the provider accepts — bytes for
Gemini, a URL for Replicate — and an adapter that cannot use the form it was
given raises rather than dropping it. WaveSpeed's text-to-image models refuse
an edit outright. Sending an image model through generate() raises too: its
reply carries no text, and returning it as an empty success is the failure this
separation exists to prevent.
Video generation
Video is a third seam, billed by the second:
from llm_gateway import ImageInput, LLMGateway, VideoRequest
from llm_gateway.factories import build_registry, create_wavespeed_client
gateway = LLMGateway(
registry=build_registry(wavespeed_client=create_wavespeed_client(api_key=my_key)),
)
result = await gateway.generate_video(
VideoRequest(
model="wavespeed-ai/minimax-h3/image-to-video",
prompt="the lioness sprints and leaps at the gazelle",
image=ImageInput(data=first_frame_bytes, mime_type="image/png"),
resolution="480p",
duration_seconds=5,
source="wildlife-clip",
)
)
result.videos[0].url
result.usage.seconds, result.usage.resolution
result.cost.amount_usd # per second, at the rate for that resolution
The first frame may be bytes or a URL — WaveSpeed accepts an inline data URI, which is what lets an image from one provider be animated by another without hosting it anywhere first.
A clip takes minutes, so VideoRequest defaults to a fifteen-minute total
budget and the adapter owns the provider's polling loop. Resolution is part of
usage rather than a detail of the request because the rate depends on it: a
resolution the price table does not know costs UNAVAILABLE, never the
cheaper rate.
The guarantees
These are enforced by tests, not by convention:
| Guarantee | Why it matters |
|---|---|
Unreported usage is None, not 0 |
A zero token count silently under-bills |
Unknown cost is UNAVAILABLE, not USD 0 |
"Free" and "unknown" are different facts |
| Cost aggregates every billable attempt | A retry that failed may still be invoiced |
| An unusable answer is a billed, failed attempt | Invalid JSON still cost money, and the fallback still gets a turn |
Each attempt carries a typed failure_phase |
configuration, provider, timeout, output_parsing or schema_validation, without parsing a message |
| Every attempt sends only options its model accepts | A fallback must not fail on a temperature the next model rejects |
| Fallback is off by default and always visible | A silent model switch corrupts A/B comparisons and cost attribution |
| Exhausted calls raise | They never return something that looks like a success |
| Errors carry the attempts already made | A failure still accounts for the money it spent |
| Output, usage, execution and cost are separate | A token count can never be mistaken for a business field |
| Sinks never receive prompts or responses | Observability without storing content |
| No module reads the environment | The application owns its credentials |
| Importing needs no provider extra | Each application installs only the SDKs it calls |
Extending
Ports are optional and default to no-op: UsageSink, EventSink, AlertSink,
AudioUsageSink, PriceCatalog and AudioPriceCatalog. Implement what you
need; the package will not reach into your application to find them.
Adding to the public API follows the two-consumer rule: nothing is promoted into the core until two distinct applications need it. Until then it belongs in that application's local adapter.
Providers
| Provider | Extra | Notes |
|---|---|---|
| OpenAI | [openai] |
Responses API |
| Google Gemini | [gemini] |
google-genai async surface, not the retired google-generativeai |
| Groq | [groq] |
Chat Completions. Declares no schema enforcement; the schema is described in the messages and the gateway validates after |
| AssemblyAI | [assemblyai] |
REST submit/poll transcription API |
| OpenRouter | [openrouter] |
Chat Completions. Aggregator: declares the floor every route honours, not the best case |
| Replicate | [replicate] |
Image generation and editing. Answers with a URL, and fetches the source image from one |
| WaveSpeed | [wavespeed] |
REST submit/poll. Text-to-image, and image-to-video billed per second |
Capabilities are declared per provider and never faked as identical — query
adapter.capabilities before relying on one.
A provider that declares structured_outputs=False has no API field that binds
the answer to a shape, so its adapter states the requested schema in the system
prompt instead of dropping it. Otherwise the model answers valid JSON under keys
of its own choosing, validation rejects it, the attempt is billed anyway and the
fallback serves every structured call — a result that looks correct and shows up
only on the invoice.
The same adapters add one sentence asking for JSON when JSON_OBJECT is
requested, because Groq rejects that mode with HTTP 400 unless the word appears
in the messages. Setting response_format is what creates the obligation, so
the adapter meets it rather than the caller's prompt — and a prompt that already
says "json" is left as it is.
function_calling and inline_files remain unsupported. OpenAI declares
remote_files=True: LLMRequest.attachments accepts already-uploaded file IDs
and the Responses adapter appends them to the last user message as
input_file parts. Providers without that capability reject the request rather
than silently dropping the files. OpenAI, Groq and AssemblyAI expose
audio_transcription=True through the separate TranscriptionRequest API,
and Gemini, Replicate and WaveSpeed expose image_generation=True through
ImageRequest. WaveSpeed also declares video_generation=True and
video_from_image=True, reached through VideoRequest.
Request options are adapted per model before each API attempt. A model that
rejects temperature — the OpenAI 5.6 family — never receives it, including
when it is reached through a fallback that inherited it from another model.
Reasoning effort is checked the same way. OpenAI 5.6
models support none, low, medium, high, xhigh, and max; Gemini 3
Flash supports minimal, low, medium, and high; Gemini 3 Pro and Groq
GPT-OSS support low, medium, and high. If a fallback cannot honour the
requested effort, the gateway uses medium when available and otherwise omits
the reasoning option.
[openrouter] installs the openai SDK, because OpenRouter speaks the OpenAI
wire format and ships none of its own. That is a fact about the transport: the
adapter, the declared capabilities and the prices are OpenRouter's.
Routing to OpenRouter
Models reach it by their namespace, so nothing needs configuring:
from llm_gateway.factories import (
build_registry,
create_openai_client,
create_openrouter_client,
)
registry = build_registry(
openai_client=create_openai_client(api_key=...),
openrouter_client=create_openrouter_client(api_key=...),
)
registry.resolve("gpt-5.6-luna") # openai
registry.resolve("deepseek/deepseek-v4-pro") # openrouter, from the catalogue
registry.resolve("somevendor/brand-new") # raises: not in the catalogue
registry.resolve("openai/gpt-oss-120b") # groq — the prefix is not OpenAI
gemini-3.1-pro-preview and google/gemini-3.1-pro-preview are the same model on
two routes, catalogued separately because they are billed separately.
For any other OpenAI-compatible endpoint — Azure, vLLM, your own gateway —
pass base_url to create_openai_client and widen the routing with
build_registry(extra_openai_prefixes=...).
Model catalogue and prices
The package ships a versioned table of models — provider, and price in USD per
million tokens — used by default, so a call is priced without you wiring
anything up. Audio models such as gpt-transcribe, Whisper and AssemblyAI are
catalogued with their duration unit and priced through a separate
AudioPriceCatalog; image models carry their per-image rate, or their token
rate where the provider bills images as tokens, and are priced through an
ImagePriceCatalog; video models carry a per-second rate that depends on the
resolution and are priced through a VideoPriceCatalog. Override prices for negotiated rates, or implement either
catalog protocol yourself. See
docs/pricing.md.
Not in this version
Tools/function calling, inline files, streaming and Gemini File Search remain
absent, and so does any video provider that answers only through a webhook —
VideoRequest polls, so Replicate's predictions and Sora's jobs need a
two-phase contract this version does not have. Remote file IDs for OpenAI,
audio transcription, image generation and polled video generation are
supported with their own capability and cost contracts.
Documentation
docs/architecture.md— how the split was decideddocs/pricing.md— cost model and updating pricesdocs/migration.md— adopting it behind an existing functionCONTRIBUTING.md— what belongs here, and the non-negotiables
Releasing without GitHub Actions
The local release runner keeps versioning and publication independent from GitHub Actions minutes. Preview a release first:
uv run --offline python scripts/release.py --version 0.6.0 --dry-run
Prepare the release locally, including tests, the version in pyproject.toml
and uv.lock, the changelog, a release commit, and an annotated tag:
uv run --offline python scripts/release.py --version 0.6.0
Add --push to push main and the tag. Add --publish as well to publish
the matching wheel and sdist with uv publish and create the GitHub Release;
the latter requires GitHub CLI authentication and UV_PUBLISH_TOKEN.
--publish implies a real external release and therefore requires --push.
Every built artifact is audited before it is uploaded, and the release is refused if the archive contains an unexpected dotfile or a credential-shaped name. What reaches a package index cannot be recalled — the file is mirrored within minutes — so the check runs between the build and the upload, which is the last moment it is still worth anything.
The local runner is the normal publisher when Actions minutes are unavailable. The GitHub workflow is manual only; use one publisher per version to avoid uploading the same PyPI files twice.
Development
uv sync
uv run pytest # no network, no cost, no extras required
uv run ruff check .
uv run mypy
uv build
Python 3.11+, Pydantic v2.
Status
0.x: in production use, but the API may still change between minor versions.
Pin an exact version. Every release documents its changes, and cost-affecting
changes are called out explicitly.
License
MIT.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file neutral_llm_gateway-0.11.0.tar.gz.
File metadata
- Download URL: neutral_llm_gateway-0.11.0.tar.gz
- Upload date:
- Size: 202.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9aa22f85d4e988bc81f9340faf163086dacfb386e429fd7c68d38220273ea501
|
|
| MD5 |
bb7c208336f050bcc660a830011b5bc6
|
|
| BLAKE2b-256 |
778ce1c37131c3b5d6905e69509a05ce295622ecdc8104e2cbf213f0a9aa9026
|
Provenance
The following attestation bundles were made for neutral_llm_gateway-0.11.0.tar.gz:
Publisher:
release.yml on jmgb/llm-gateway-python
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
neutral_llm_gateway-0.11.0.tar.gz -
Subject digest:
9aa22f85d4e988bc81f9340faf163086dacfb386e429fd7c68d38220273ea501 - Sigstore transparency entry: 2350660647
- Sigstore integration time:
-
Permalink:
jmgb/llm-gateway-python@2af98683abfae5256d312fcfe5306d47485d10ba -
Branch / Tag:
refs/heads/main - Owner: https://github.com/jmgb
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@2af98683abfae5256d312fcfe5306d47485d10ba -
Trigger Event:
workflow_dispatch
-
Statement type:
File details
Details for the file neutral_llm_gateway-0.11.0-py3-none-any.whl.
File metadata
- Download URL: neutral_llm_gateway-0.11.0-py3-none-any.whl
- Upload date:
- Size: 75.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0f2596c8eadede03d5b4dfe98ebd8adcf73194e75ae7adec8c8c1779c57138c3
|
|
| MD5 |
8eeaa548098ebc96af7026fe0bee64ff
|
|
| BLAKE2b-256 |
e9a715574a9ebc8ffa6f1775a3445cea7ffe222a1abaabcc53b24959421f224c
|
Provenance
The following attestation bundles were made for neutral_llm_gateway-0.11.0-py3-none-any.whl:
Publisher:
release.yml on jmgb/llm-gateway-python
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
neutral_llm_gateway-0.11.0-py3-none-any.whl -
Subject digest:
0f2596c8eadede03d5b4dfe98ebd8adcf73194e75ae7adec8c8c1779c57138c3 - Sigstore transparency entry: 2350660743
- Sigstore integration time:
-
Permalink:
jmgb/llm-gateway-python@2af98683abfae5256d312fcfe5306d47485d10ba -
Branch / Tag:
refs/heads/main - Owner: https://github.com/jmgb
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@2af98683abfae5256d312fcfe5306d47485d10ba -
Trigger Event:
workflow_dispatch
-
Statement type: