Skip to main content

Inferrail logo

Know what your AI work costs.
Without keeping what it said.

Inferrail is a gateway you run yourself, between your app and OpenAI or Anthropic. It shows what each AI job costs, blocks calls that would push a job past its spending limit, and doesn't copy prompts or responses into its records.

CI PyPI version License: Apache-2.0

Quickstart · Where your data goes · How it works · Docs · tryinferrail.com

Recording of a real Inferrail run. The gateway starts with its local dashboard. An agent script gives one AI job, contract-review-42, a $0.04 spending limit and makes six model calls: four are answered and two are blocked with HTTP 402 before reaching the model. The dashboard lists each call with its cost, then shows the job spent $0.0325 of its $0.0400 budget, with 2 requests blocked before reaching the provider.

A real run of the Inferrail gateway and dashboard, with a local stand-in model: no API key, no provider charges. How it was recorded

Developer preview. Everything below is implemented and tested. CLI flags, config, and receipt fields may still change before 1.0.

Quickstart

Requires Python 3.11+.

Protect one run, inside your Python app. No config file and no second terminal:

pip install inferrail
export OPENAI_API_KEY=sk-...
import inferrail
from openai import OpenAI

base_url = inferrail.start()   # the gateway, on a background thread in this process
client = OpenAI(base_url=base_url, api_key="unused")
client.chat.completions.create(
    model="gpt-4o-mini",   # example: any model your account can use (`inferrail models` lists them)
    max_tokens=200,
    messages=[{"role": "user", "content": "Summarize this contract."}],
    extra_headers={
        "X-Inferrail-Attribute-Work-Id": "contract-review-42",   # the run
        "X-Inferrail-Budget-Usd": "0.50",                        # its dollar ceiling
    },
)

Every call that carries the same run id shares that budget, including parallel calls. A call that would push the run past it gets HTTP 402 before it reaches OpenAI. Then:

inferrail work contract-review-42

shows what the run cost.

Inferrail doesn't choose a model: whatever model id you send is passed to the provider. A dollar budget needs a price for that model; inferrail models shows which models have one, and inferrail.start(pricing=...) adds a price for a new model. Framework snippets (LangChain, LangGraph, OpenAI Agents SDK, CrewAI, Haystack, LlamaIndex, Microsoft Agent Framework): recipe.

Try it offline. No API key, no network calls, no provider charges:

pip install inferrail
inferrail demo
inferrail report --by customer --receipts ./inferrail-demo-receipts.jsonl

The demo sends scripted requests through the real engine to a fake provider with made-up prices labeled DEMO, and prints what they cost by customer (recording).

Run it as its own process, with the local dashboard. For a gateway shared by several apps, or to watch calls live, set your key in the gateway's terminal:

export OPENAI_API_KEY=sk-...          # and/or ANTHROPIC_API_KEY
inferrail serve --quickstart --app-mode

Open the Dashboard: URL it prints, then point your app at the gateway and tag the work:

from openai import OpenAI

# api_key is a placeholder: your provider key stays in the gateway
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")
client.chat.completions.create(
    model="gpt-4o-mini",   # example model id; Inferrail passes yours through
    messages=[{"role": "user", "content": "Summarize this contract."}],
    extra_headers={
        "X-Inferrail-Attribute-Customer": "acme",
        "X-Inferrail-Attribute-Work-Id": "contract-review-42",
        "X-Inferrail-Budget-Usd": "0.50",   # optional: a dollar ceiling for this work
    },
)

The call appears in the Live Feed, and the Work screen totals what contract-review-42 cost. The Anthropic SDK works the same way with base_url="http://127.0.0.1:8000" (no /v1). More clients and frameworks: docs/integrations.md.

Setting up Python, or inferrail: command not found
python3 -m venv .venv && source .venv/bin/activate   # Windows: .venv\Scripts\activate
python -m pip install inferrail

If inferrail is still not found, the environment isn't active or pip installed into a different Python. More in docs/self-hosting.md.

What you get

  • Cost per unit of work. Tag calls with any attribute (customer, workflow, work_id, …). The dashboard and inferrail report --by <tag> total known cost by those tags, and inferrail work <id> shows one job.
  • A dollar budget for one agent run. Declare it in a header. Parallel calls share it safely, and calls that would exceed it get HTTP 402 before they reach the provider. Global, project, daily, and monthly budgets too. Recipe.
  • Numbers you can trust. A cost is recorded only when the provider reports usage and a price is on file. Otherwise it's unknown, never a guessed $0.
  • Local records. Receipts are SQLite or JSONL files on your machine. No Inferrail account or hosted service is involved.
  • Ask your agent. A read-only MCP server answers questions like "How much did work contract-review-42 cost?"

Privacy boundary

Where What it sees or keeps
Your provider The full request, exactly as it would without Inferrail, under the provider's own policies.
The gateway, in memory Your provider key (from its own environment) and the prompts and responses it forwards.
Receipts, on your disk Provider, model, token usage, the price used and the cost, status, timing, and the attribution tags you send.
Not in receipts Prompt and response bodies.

Attribution tags are stored exactly as sent, so use identifiers and keep secrets and message content out of them. inferrail verify-payload-free prints the live receipt schema, and canary tests check that message bodies never reach receipts. Neither is a security audit. The optional usage beacon sends nothing unless an endpoint is configured (details). How to check all of this yourself: docs/PRODUCT.md.

How it works

Data flow. On your machine, your application sends requests to the Inferrail gateway, which reads the provider API key from its environment, forwards request content and the key to OpenAI or Anthropic, and returns the response. Separately, the gateway writes a metadata receipt (tokens, cost, status, timing, attributes, no message bodies) to a local JSONL or SQLite store that reports, the local dashboard, and MCP tools read. A usage beacon collector receives lifecycle events only if an endpoint is configured.

The gateway routes each request by model to a configured provider, checks any budget in scope, forwards the call, and writes one receipt per request, including requests a budget refused. Reports, the dashboard, and MCP tools all read those receipts. Details: docs/ARCHITECTURE.md.

MCP

inferrail mcp is a stdio MCP server with two read-only tools over your local receipts: get_spend (known cost and tokens grouped by provider, model, route, or any tag such as work_id) and get_health. They don't run inference or write files.

claude mcp add inferrail \
  -e INFERRAIL_RECEIPTS_PATH=/absolute/path/to/inferrail-receipts.jsonl \
  -- uvx --with "mcp>=2.0" inferrail mcp

Other clients, and where --app-mode keeps its receipts: docs/integrations.md.

Current support

  • Endpoints: POST /v1/chat/completions (OpenAI-compatible) and POST /v1/messages (Anthropic-compatible), with streaming and tool calls. Any client that lets you set a base URL and sends these request shapes can use them.
  • Providers: OpenAI, Anthropic, and endpoints compatible with either. Built-in prices cover OpenAI and Anthropic models. Other endpoints need a price declared in your config, or their cost stays unknown.
  • Images in chat messages (image_url parts, e.g. a browser agent's screenshots) are forwarded to OpenAI-compatible providers. Under a budget each image reserves a fixed 3,000-token estimate; the actual cost comes from the provider's reported usage.
  • Not supported: the OpenAI Responses API, embeddings, image generation, audio and the Realtime API, batch, and native Gemini or Bedrock APIs. Request fields the gateway can't account for are rejected with a clear error, never silently dropped.
  • Scope: only calls that go through a running gateway are counted, not everything on your provider account.

Exact contract and non-goals: docs/PRODUCT.md.

Experimental capabilities

These are separate from the gateway above and not needed to use it.

  • AP invoice-exception recovery (experimental): decide and record a retry or human review for one invoice-extraction exception. Try inferrail ap demo. Docs.
  • Hosted trial (preview): a short-lived hosted gateway at tryinferrail.com/try. If you add a real key there, the hosted process holds it (key handling).
  • Work Economics and Economic Authority (experimental, hosted, Base Sepolia testnet only): Work Economics docs · Economic Authority docs · example.

Documentation

Contributing, security, license

Questions and bugs: GitHub issues. Security or privacy vulnerabilities: report privately, as described in SECURITY.md. Contributing: CONTRIBUTING.md (pytest needs no API key or network). License: Apache-2.0.

Metadata

Release files for inferrail 0.4.13

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for inferrail 0.4.13
File Size Uploaded
inferrail-0.4.13.tar.gz 2.0 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for inferrail 0.4.13
File Interpreter ABI Platform
inferrail-0.4.13-py3-none-any.whl Python 3 none any Details

Total release size: 2.3 MB

Release files / inferrail-0.4.13.tar.gz

Download URL inferrail-0.4.13.tar.gz
Size 2.0 MB
Tags Source
SHA-256 checksum
How to use checksums
40a37acac1b05c292a0b24b86e06c3476d3590cdab6a69acb2cd59afc82906d3
BLAKE2b-256 checksum
How to use checksums
7d913ecc909dfd487e2dc31f39c533242981c7f67c83992746d251d3a28d6113
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.

Transparency log

Release files / inferrail-0.4.13-py3-none-any.whl

Download URL inferrail-0.4.13-py3-none-any.whl
Size 278.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3f493e901b2b17cb11ba4d2cd347feb3a5fceb4830c57f5abd1489212d8b7d2e
BLAKE2b-256 checksum
How to use checksums
b7c3c7c588db3cedb3f495f954f5f27e82e5e686e23ff48d7f4d4b7db05b0b36
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.4.13 This release

2 release files

0.4.9

2 release files

0.4.8

2 release files

0.4.7

2 release files

0.4.6

2 release files

0.4.5

2 release files

0.4.4

2 release files

0.4.3

2 release files

0.4.1

2 release files

0.2.0

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page