Skip to main content

BiVelio Savings Layer for Python

A thin client for the local bvsala gateway. It routes the official OpenAI and Anthropic Python SDKs through the gateway and hands you, for each request, the report the gateway recorded for your BiVelio dashboard: the provider's own token counts, the cost at the model's list price, and where the figure comes from (cost.basis).

  • The gateway does the work. It ships in the npm package @bivelio/savings-layer, holds your BiVelio license, forwards each request to your provider and measures it with the provider's usage counter. This package prices nothing itself, so its numbers are exactly the dashboard's.
  • Fail-closed. Without a running gateway and a live license, wrap() raises. It never falls back to calling your provider directly.
  • No dependencies. Standard library only. It does not import the OpenAI or Anthropic SDKs; you bring them.
  • Your prompts never reach BiVelio. The client only talks to the gateway on 127.0.0.1, ignoring any HTTP proxy settings, and the gateway forwards your requests to your provider. What it sends to BiVelio is per-request usage (token counts, cost, the model id and opaque ids), never your content: the LICENSE, section 5, lists every field.

This package measures. The saving levers of the TypeScript SDK do not run in Python; the gateway's own levers apply to Anthropic /v1/messages traffic with a PRO license.

Requirements

  • Python 3.10 or newer.
  • Node.js 20 or newer, for the gateway, and a BiVelio license with a service key (BIVELIO_SERVICE_KEY): the gateway does not start without them.

Install

pip install bivelio-savings-layer

The package is bivelio-savings-layer on PyPI and bvsala in your code (import bvsala). From a checkout of the repository, to try unreleased changes:

pip install ./packages/python

Start a gateway

One gateway per provider:

export BIVELIO_SERVICE_KEY=…
npx -p @bivelio/savings-layer bvsala gateway                                 # Anthropic, on 8402
npx -p @bivelio/savings-layer bvsala gateway --upstream openai --port 8403   # OpenAI, on 8403
python -m bvsala --url http://127.0.0.1:8403 status                          # version, upstream, license

Use it

import bvsala
from openai import OpenAI

gw = bvsala.Gateway("http://127.0.0.1:8403")
client = gw.wrap(OpenAI())          # a copy that goes through the gateway; OpenAI() is untouched

resp = client.chat.completions.create(
    model="gpt-4o-mini", messages=[{"role": "user", "content": "Hi"}],
)
report = gw.report(resp)
print(report.summary())             # gpt-4o-mini · measured · $0.000024 (in 100, cached 40, out 20)
print(report.cost.actual, report.cost.basis, report.tokens.actual_output)

Streaming works the same way. For OpenAI Chat Completions, ask for usage in the request, or the provider sends no counter and the report says so:

stream = client.chat.completions.create(
    model="gpt-4o-mini", messages=[{"role": "user", "content": "Hi"}],
    stream=True, stream_options={"include_usage": True},
)
for chunk in stream:
    if chunk.choices:
        print(chunk.choices[0].delta.content or "", end="")
print(gw.report(stream).summary())

Anthropic:

from anthropic import Anthropic

gw = bvsala.Gateway()                # $BVSALA_GATEWAY_URL or http://127.0.0.1:8402
claude = gw.wrap(Anthropic())
msg = claude.messages.create(model="claude-haiku-4-5", max_tokens=256,
                             messages=[{"role": "user", "content": "Hi"}])
print(gw.report(msg).summary())

gw.report() accepts a response, a stream you have read, any chunk or event of it, a raw response from with_raw_response, or an id string. It also reads a LiteLLM response sent with extra_headers=bvsala.REPORT_HEADERS, from the provider headers LiteLLM keeps in _hidden_params (checked against LiteLLM's source, not yet end to end).

gw.report() waits up to wait seconds (5 by default) for a request still in flight; after that the report comes back with status == "pending". In async code use await gw.areport(resp); wrap() also takes AsyncOpenAI, AzureOpenAI and AsyncAnthropic.

gw.totals() adds up the reports this process fetched. Savings stay split by basis (measured, reference, estimated), never added into one figure.

Any other client

The gateway side works with any HTTP client: send x-bvsala-report: 1, read the x-bvsala-request-id response header, then GET /_bvsala/v1/reports/<id> on the gateway. bvsala.REPORT_HEADERS is that header as a dict, for extra_headers= or default_headers=.

Errors

Exception When
GatewayUnavailable nothing answers at the URL, or it is not a bvsala gateway with local reports (0.4.26 or later is needed)
LicenseInactive the gateway's license does not serve new requests; .state says why
GatewayMismatch the client points at a different provider than the one the gateway forwards to
ReportNotFound the gateway holds no report for that id (only requests sent with x-bvsala-report: 1 have one, kept for at most an hour)

License

Commercial software, © 2026 BiVelio Inc. Free to install, paid to run: see LICENSE and savings.bivelio.com.

Metadata

Release files for bivelio-savings-layer 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for bivelio-savings-layer 0.1.0
File Size Uploaded
bivelio_savings_layer-0.1.0.tar.gz 26.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for bivelio-savings-layer 0.1.0
File Interpreter ABI Platform
bivelio_savings_layer-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 47.5 kB

Release files / bivelio_savings_layer-0.1.0.tar.gz

Download URL bivelio_savings_layer-0.1.0.tar.gz
Size 26.4 kB
Tags Source
SHA-256 checksum
How to use checksums
79ff6e521384d77052d42d03b0d3d481c819ed15ac3c0fb7e6480ea9df7c4e77
BLAKE2b-256 checksum
How to use checksums
5a753a807b4f933e05a3418ee6050d9e143ae28ab3918c200cf84183277ec46e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / bivelio_savings_layer-0.1.0-py3-none-any.whl

Download URL bivelio_savings_layer-0.1.0-py3-none-any.whl
Size 21.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
59ca1f4a714f5411587e03d723af5b0560f467631d7847ee1cfea7e0a5edfa05
BLAKE2b-256 checksum
How to use checksums
dc33565854a515d2746fa5f8398c40f7b3c602d1d1b2e0a7fdb3a5ae467f6418
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page