Know what your AI work costs.
Without keeping what it said.
Inferrail is a gateway you run yourself, between your app and OpenAI or Anthropic. It shows what each AI job costs, blocks calls that would push a job past its spending limit, and doesn't copy prompts or responses into its records.
Quickstart · Where your data goes · How it works · Docs · tryinferrail.com
A real run of the Inferrail gateway and dashboard, with a local stand-in model: no API key, no provider charges. How it was recorded
Developer preview. Everything below is implemented and tested. CLI flags, config, and receipt fields may still change before 1.0.
Quickstart
Requires Python 3.11+.
Protect one run, inside your Python app. No config file and no second terminal:
pip install inferrail
export OPENAI_API_KEY=sk-...
import inferrail
from openai import OpenAI
base_url = inferrail.start() # the gateway, on a background thread in this process
client = OpenAI(base_url=base_url, api_key="unused")
client.chat.completions.create(
model="gpt-4o-mini", # example: any model your account can use (`inferrail models` lists them)
max_tokens=200,
messages=[{"role": "user", "content": "Summarize this contract."}],
extra_headers={
"X-Inferrail-Attribute-Work-Id": "contract-review-42", # the run
"X-Inferrail-Budget-Usd": "0.50", # its dollar ceiling
},
)
Every call that carries the same run id shares that budget, including parallel calls. A call that would push the run past it gets HTTP 402 before it reaches OpenAI. Then:
inferrail work contract-review-42
shows what the run cost.
Inferrail doesn't choose a model: whatever model id you send is passed to
the provider. A dollar budget needs a price for that model; inferrail models shows which models have one, and inferrail.start(pricing=...)
adds a price for a new model. Framework snippets (LangChain, LangGraph,
OpenAI Agents SDK, CrewAI, Haystack, LlamaIndex, Microsoft Agent Framework):
recipe.
Try it offline. No API key, no network calls, no provider charges:
pip install inferrail
inferrail demo
inferrail report --by customer --receipts ./inferrail-demo-receipts.jsonl
The demo sends scripted requests through the real engine to a fake
provider with made-up prices labeled DEMO, and prints what they cost by
customer (recording).
Run it as its own process, with the local dashboard. For a gateway shared by several apps, or to watch calls live, set your key in the gateway's terminal:
export OPENAI_API_KEY=sk-... # and/or ANTHROPIC_API_KEY
inferrail serve --quickstart --app-mode
Open the Dashboard: URL it prints, then point your app at the gateway
and tag the work:
from openai import OpenAI
# api_key is a placeholder: your provider key stays in the gateway
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")
client.chat.completions.create(
model="gpt-4o-mini", # example model id; Inferrail passes yours through
messages=[{"role": "user", "content": "Summarize this contract."}],
extra_headers={
"X-Inferrail-Attribute-Customer": "acme",
"X-Inferrail-Attribute-Work-Id": "contract-review-42",
"X-Inferrail-Budget-Usd": "0.50", # optional: a dollar ceiling for this work
},
)
The call appears in the Live Feed, and the Work screen totals what
contract-review-42 cost. The Anthropic SDK works the same way with
base_url="http://127.0.0.1:8000" (no /v1). More clients and
frameworks: docs/integrations.md.
Setting up Python, or inferrail: command not found
python3 -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
python -m pip install inferrail
If inferrail is still not found, the environment isn't active or pip
installed into a different Python. More in
docs/self-hosting.md.
What you get
- Cost per unit of work. Tag calls with any attribute (
customer,workflow,work_id, …). The dashboard andinferrail report --by <tag>total known cost by those tags, andinferrail work <id>shows one job. - A dollar budget for one agent run. Declare it in a header. Parallel calls share it safely, and calls that would exceed it get HTTP 402 before they reach the provider. Global, project, daily, and monthly budgets too. Recipe.
- Numbers you can trust. A cost is recorded only when the provider
reports usage and a price is on file. Otherwise it's
unknown, never a guessed$0. - Local records. Receipts are SQLite or JSONL files on your machine. No Inferrail account or hosted service is involved.
- Ask your agent. A read-only MCP server answers questions like "How much did work contract-review-42 cost?"
Privacy boundary
| Where | What it sees or keeps |
|---|---|
| Your provider | The full request, exactly as it would without Inferrail, under the provider's own policies. |
| The gateway, in memory | Your provider key (from its own environment) and the prompts and responses it forwards. |
| Receipts, on your disk | Provider, model, token usage, the price used and the cost, status, timing, and the attribution tags you send. |
| Not in receipts | Prompt and response bodies. |
Attribution tags are stored exactly as sent, so use identifiers and keep
secrets and message content out of them. inferrail verify-payload-free
prints the live receipt schema, and canary tests check that message
bodies never reach receipts. Neither is a security audit. The optional
usage beacon sends nothing unless an endpoint is configured
(details). How to check all of this
yourself: docs/PRODUCT.md.
How it works
The gateway routes each request by model to a configured provider,
checks any budget in scope, forwards the call, and writes one receipt per
request, including requests a budget refused. Reports, the dashboard, and MCP tools all read those
receipts. Details: docs/ARCHITECTURE.md.
MCP
inferrail mcp is a stdio MCP server with two read-only tools over your
local receipts: get_spend (known cost and tokens grouped by provider,
model, route, or any tag such as work_id) and get_health. They don't
run inference or write files.
claude mcp add inferrail \
-e INFERRAIL_RECEIPTS_PATH=/absolute/path/to/inferrail-receipts.jsonl \
-- uvx --with "mcp>=2.0" inferrail mcp
Other clients, and where --app-mode keeps its receipts:
docs/integrations.md.
Current support
- Endpoints:
POST /v1/chat/completions(OpenAI-compatible) andPOST /v1/messages(Anthropic-compatible), with streaming and tool calls. Any client that lets you set a base URL and sends these request shapes can use them. - Providers: OpenAI, Anthropic, and endpoints compatible with either. Built-in prices cover OpenAI and Anthropic models. Other endpoints need a price declared in your config, or their cost stays unknown.
- Images in chat messages (
image_urlparts, e.g. a browser agent's screenshots) are forwarded to OpenAI-compatible providers. Under a budget each image reserves a fixed 3,000-token estimate; the actual cost comes from the provider's reported usage. - Not supported: the OpenAI Responses API, embeddings, image generation, audio and the Realtime API, batch, and native Gemini or Bedrock APIs. Request fields the gateway can't account for are rejected with a clear error, never silently dropped.
- Scope: only calls that go through a running gateway are counted, not everything on your provider account.
Exact contract and non-goals: docs/PRODUCT.md.
Experimental capabilities
These are separate from the gateway above and not needed to use it.
- AP invoice-exception recovery (experimental): decide and record a
retry or human review for one invoice-extraction exception. Try
inferrail ap demo. Docs. - Hosted trial (preview): a short-lived hosted gateway at tryinferrail.com/try. If you add a real key there, the hosted process holds it (key handling).
- Work Economics and Economic Authority (experimental, hosted, Base Sepolia testnet only): Work Economics docs · Economic Authority docs · example.
Documentation
- Give one AI agent run a dollar budget
- Integrations: SDKs, frameworks, attribution, MCP, voice
- Self-hosting: install, configuration, storage, budgets, dashboard
- Product scope and architecture
- openapi.json, config.schema.json, llms.txt
Contributing, security, license
Questions and bugs: GitHub issues.
Security or privacy vulnerabilities: report privately, as described in
SECURITY.md. Contributing: CONTRIBUTING.md
(pytest needs no API key or network). License: Apache-2.0.
Metadata
Release files for inferrail 0.4.13
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| inferrail-0.4.13.tar.gz | 2.0 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| inferrail-0.4.13-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 2.3 MB
Release files / inferrail-0.4.13.tar.gz
| Download URL | inferrail-0.4.13.tar.gz |
|---|---|
| Size | 2.0 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
40a37acac1b05c292a0b24b86e06c3476d3590cdab6a69acb2cd59afc82906d3
|
|
BLAKE2b-256 checksum How to use checksums |
7d913ecc909dfd487e2dc31f39c533242981c7f67c83992746d251d3a28d6113
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.
Transparency logRelease files / inferrail-0.4.13-py3-none-any.whl
| Download URL | inferrail-0.4.13-py3-none-any.whl |
|---|---|
| Size | 278.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
3f493e901b2b17cb11ba4d2cd347feb3a5fceb4830c57f5abd1489212d8b7d2e
|
|
BLAKE2b-256 checksum How to use checksums |
b7c3c7c588db3cedb3f495f954f5f27e82e5e686e23ff48d7f4d4b7db05b0b36
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.
Transparency log