Skip to main content

airmux

One self-hosted LLM gateway. Any compatible client. Multiple providers.

Connect through a supported inference API while airmux centralizes provider translation, routing, policy, credentials, failover, and usage accounting behind one endpoint.

CI PyPI Python 3.13+ Elastic License 2.0

Quickstart · Documentation · Contributing · Report a bug

airmux gives applications, agents, CLIs, and services one self-hosted origin for calling multiple provider families. Clients can send Chat Completions, Responses, or Messages requests through any SDK or integration that can target the corresponding HTTP API. airmux authenticates the workspace, applies policy, selects a model and scoped provider credential, translates the request, and records the result.

Quickstart

Run a local gateway with no Docker, Postgres, or control plane. You need Python 3.13+, uv, and an OpenAI API key.

1. Install and start the gateway

uv tool install airmux
mkdir airmux-demo
cd airmux-demo
export OPENAI_API_KEY='your-provider-key'
airmux gateway init
airmux gateway serve

init creates a local configuration, a model taxonomy, and a private inference key under .airmux/. The gateway reads provider credentials from the environment and reloads taxonomy edits while it runs.

2. Make a real model request

In another terminal:

cd airmux-demo
export AIRMUX_INFERENCE_KEY="$(cat .airmux/inference.key)"
curl --fail-with-body http://127.0.0.1:8080/inf/v1/chat/completions \
  -H "Authorization: Bearer $AIRMUX_INFERENCE_KEY" \
  -H 'Content-Type: application/json' \
  -d '{"model":"openai/gpt-4o-mini","messages":[{"role":"user","content":"Say hello in one word."}],"max_completion_tokens":16}'

The same gateway accepts streaming requests, tool calls, structured output, reasoning, images, and PDF inputs when the selected model supports them. Continue with the gateway-only guide for configuration and operation.

Use your existing client

Any client that can target one of airmux's exposed HTTP APIs and send an inference key through Authorization: Bearer or x-api-key can connect. That includes SDKs, agent frameworks, CLIs, services, and raw HTTP integrations.

API Endpoint
Chat Completions POST /inf/v1/chat/completions
Responses POST /inf/v1/responses
Messages POST /inf/v1/messages
Model discovery GET /inf/v1/models and GET /inf/v1/models/{model_id}

The OpenAI SDK is one example. Point it at /inf/v1 and replace the upstream key with an airmux inference key:

import os

from openai import OpenAI

client = OpenAI(
    base_url="http://127.0.0.1:8080/inf/v1",
    api_key=os.environ["AIRMUX_INFERENCE_KEY"],
)

response = client.chat.completions.create(
    model="openai/gpt-4o-mini",
    messages=[{"role": "user", "content": "Why use an LLM gateway?"}],
)

print(response.choices[0].message.content)

The client protocol does not constrain the provider route. A Messages request can target an OpenAI-compatible model, and a Chat Completions request can target an Anthropic model. airmux translates the request and returns the response and errors in the caller's dialect. See the OpenAI SDK and Anthropic SDK guides for complete examples, including streaming.

Why airmux

  • Protocol-first clients: connect any SDK, agent framework, CLI, service, or raw HTTP integration that speaks an exposed API
  • Policy at the gateway: compose model and provider allowlists, price ceilings, request limits, credential rules, denials, strict parameters, and fallbacks
  • Scoped provider secrets: separate instance, organization, and workspace credentials without exposing secret values to configuration bundles
  • Predictable failover: retry eligible credentials and route to bounded backup models without escaping workspace policy
  • Complete request records: capture tokens, cost, latency, status, credential scope, configuration version, and every fallback attempt
  • A resilient request path: gateways evaluate immutable local bundles and can keep serving through a control-plane outage

The shipped catalog includes Anthropic, Cerebras, DeepSeek, Fireworks, Groq, Mistral, OpenAI, Together, and xAI. Model IDs, prices, context windows, modalities, capabilities, and parameter support are explicit, inspectable data in the taxonomy.

Full-platform quickstart

Use the complete stack when you want the web console, organizations and workspaces, managed credentials, live policy, usage history, and audit activity. You need Docker with Compose 2.24.4+, Python 3.13+, uv, and at least one provider API key.

git clone https://github.com/michel-tricot/airmux.git
cd airmux
cp .env.example .env

Add a provider key such as OPENAI_API_KEY or ANTHROPIC_API_KEY to .env, then run:

uv tool install airmux
docker compose up -d --build --wait
airmux quickstart --url http://localhost:8080

quickstart creates or resumes the owner account, organization, and workspace; imports missing provider credentials; mints an inference key; and proves the installation with a real model request. Open localhost:8080 for the console.

Goal Command
List catalog models airmux models list
Inspect the installation airmux doctor
Follow gateway activity airmux events tail --interval 2 --keep 30
Follow service logs docker compose logs -f airmux
Stop while preserving state docker compose down

Architecture

airmux separates mutable management work from the inference request path.

flowchart LR
  A[Application] -->|Inference key| G[Data plane]
  U[Operator] --> C[Console or CLI]
  C --> M[Control plane]
  M --> P[(Postgres)]
  M --> S[(Secret store)]
  M -->|Versioned bundles| G
  G -->|Cold secret resolution| S
  G -->|Provider request| L[LLM provider]
  G -->|Usage and health| M

Caller dialects and provider protocols cross through one canonical model. Adding a caller dialect requires one ingress adapter; adding a provider family requires one egress adapter. Policy, routing, streaming, and metering stay provider-neutral instead of multiplying into a translator for every caller and provider pair.

The control plane compiles complete, versioned organization bundles. Data-plane workers validate them, build immutable indexes, and atomically adopt them. Inference therefore avoids management database reads and an in-flight request never observes partially updated policy. Cold provider-secret resolution is the only database-capable exception, and secret values never enter a bundle.

Read the architecture guide for the full data flow, failure boundaries, and deployment shapes.

Documentation

I want to... Start here
Try the complete stack Quickstart
Connect an application OpenAI SDK or Anthropic SDK
Add routing and access rules Policy workflow
Understand supported inference shapes Inference reference
Deploy airmux Deployment overview
Upgrade or roll back a deployment Upgrade and rollback
Call the management API Management API
Work on the project Contributing and Development guide

The complete management API is generated from lib/api-spec/openapi.yaml.

Release files for airmux 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for airmux 0.1.2
File Size Uploaded
airmux-0.1.2.tar.gz 202.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for airmux 0.1.2
File Interpreter ABI Platform
airmux-0.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 492.0 kB

Release files / airmux-0.1.2.tar.gz

Download URL airmux-0.1.2.tar.gz
Size 202.2 kB
Tags Source
SHA-256 checksum
How to use checksums
4c32f1bd345aad574b5b10e32dc058e7597153938e024ab872b302b380c2ecfc
BLAKE2b-256 checksum
How to use checksums
ead0b99e12bcb9744d4e98cfc7a01918aca6859325a90b7cacf9015fc366c95b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 18, 2026.

Transparency log

Release files / airmux-0.1.2-py3-none-any.whl

Download URL airmux-0.1.2-py3-none-any.whl
Size 289.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3551db091230415d114ca53f369f0a0a96926263278577289b916bb65a0f1763
BLAKE2b-256 checksum
How to use checksums
df056993aea65fdca6c86a1b1544e2b0cbb5d710f6ca7a6b52ef609c967cbd8e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 18, 2026.

Transparency log

Release history Release notifications | RSS feed

0.2.3

2 release files

0.2.0

2 release files

This release

0.1.2 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page