Skip to main content

usageflow-vibe (beta)

usageflow-vibe is one Python client for Anthropic and OpenAI that checks every call against your UsageFlow limits and policies before it runs, then records the real token usage.

It is separate from the Flask/FastAPI route middleware: use it when you call LLMs from your own code and want to meter and govern those calls per customer.

Beta: this package is new and its APIs may still change between minor versions. If you hit an issue, please report it.

Usage is recorded asynchronously right after each call, so a balance you read back from credits() can lag by a moment; withdraw, credit, and close return as soon as the request is sent, not once it has been applied.

Install

pip install usageflow-vibe
export USAGEFLOW_API_KEY="your-api-key"
export OPENAI_API_KEY="..."      # only the provider(s) you call
export ANTHROPIC_API_KEY="..."

Python 3.9+. The only dependency is websocket-client; provider calls use the standard library.

Chat

from usageflow.vibe import VibeClient, Message, RejectionError

client = VibeClient()  # reads USAGEFLOW_API_KEY

try:
    result = client.chat(
        identity="cust_acme",          # the customer you meter
        workflow="support-agent",      # optional: the slug of a Vibe policy
        model="claude-sonnet-5",       # provider inferred: claude-* → Anthropic, else OpenAI
        messages=[Message("user", "Summarize this ticket.")],
        customer_metadata={"plan": "pro"},  # optional: string/number/bool values only
    )
    print(result.content, result.usage)
except RejectionError as e:
    print("blocked by UsageFlow:", e.message)
  • Denied calls never reach the provider. A quota or policy denial raises RejectionError.
  • Fails closed. If UsageFlow can't be reached, UsageFlowUnavailableError is raised and the provider is not called.
  • Policies can reroute the call. If a matched policy tier names a model, the call runs on that model instead. result.model / result.provider show what actually ran, result.requested_model what you asked for, and result.vibe_policy the tier that fired.
  • max_tokens defaults to 1024 and is enforced. For OpenAI reasoning models (o1, o3, o4, gpt-5) it is sent as max_completion_tokens, and temperature is dropped.
  • tools and tool_choice are passed to the provider as-is; tool calls come back in result.tool_calls.
  • Results have to_dict() for JSON responses.

A workflow that doesn't match an active policy is ignored. Usage is counted per identity, so all workflows for one identity share one total; each workflow only defines its own rules.

Stream

stream = client.stream(identity="cust_acme", model="gpt-4o-mini",
                       messages=[Message("user", "Tell me a story")])
for chunk in stream:
    print(chunk, end="", flush=True)
result = stream.result()  # usage is recorded when the stream ends

A denial raises when you call stream(), before any text is produced. Always drain the stream.

Embeddings (OpenAI)

result = client.embed(identity="cust_acme", model="text-embedding-3-small", input=["hello"])

Adjust a customer's balance

client.withdraw(identity="cust_acme", amount=50, idempotency_key="order-123", reason="export")
client.credit(identity="cust_acme", amount=50, idempotency_key="refund-123")  # reverses

idempotency_key is recorded with the event, but duplicates are not rejected yet — don't retry blindly.

Hold now, charge later

hold = client.withdraw_async(identity="cust_acme", amount=100, idempotency_key="job-7")
# ... do the work ...
client.close(hold.capture_id, 60)   # charge 60; omit the amount to charge the full 100

withdraw_async checks limits and policies up front; nothing is charged until close. A hold can be closed once. To close it from another process, pass identity= and the amount.

A hold stays open for 24 hours by default. Pass hold_for (a timedelta or seconds) to change it, and close before hold.expires_at (epoch ms): a hold that expires unclosed is released without charging. Closing early charges the amount you pass and releases the rest right away.

from datetime import timedelta
hold = client.withdraw_async(identity="cust_acme", amount=100, idempotency_key="job-8",
                             hold_for=timedelta(hours=2))

Configuration

Argument Environment variable
api_key USAGEFLOW_API_KEY required
openai_api_key OPENAI_API_KEY read on first OpenAI call
anthropic_api_key ANTHROPIC_API_KEY read on first Anthropic call
timeout seconds to wait for UsageFlow (default 10)

The client connects on first use; call client.connect() at startup to surface a bad key early, and client.destroy() (or use it as a context manager) on shutdown. The client is thread-safe.

Not included yet

Image generation, speech, transcription, moderation and batch are not in the Python SDK. Async (asyncio) usage is not supported yet; call the client from a thread (e.g. asyncio.to_thread).

Example app

examples/vibe in the UsageFlow Python repository is a runnable HTTP app with the same routes as the JS and Go examples.

Release files for usageflow-vibe 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for usageflow-vibe 0.1.0
File Size Uploaded
usageflow_vibe-0.1.0.tar.gz 26.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for usageflow-vibe 0.1.0
File Interpreter ABI Platform
usageflow_vibe-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 46.0 kB

Release files / usageflow_vibe-0.1.0.tar.gz

Download URL usageflow_vibe-0.1.0.tar.gz
Size 26.7 kB
Tags Source
SHA-256 checksum
How to use checksums
49789ab56f07b535acf1dc3546876602bfad53005e793795d9b16e613550c4ab
BLAKE2b-256 checksum
How to use checksums
894d2f5facf163f7e74403a278a5591b24f9b11ef8847487124ace0b87bd2281
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.2

Release files / usageflow_vibe-0.1.0-py3-none-any.whl

Download URL usageflow_vibe-0.1.0-py3-none-any.whl
Size 19.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
219e8dcc46ad94b607c778b25c4875d312ffa95727373de41d558e460c0cab12
BLAKE2b-256 checksum
How to use checksums
f69f9191be2e48cd9ed18719017164fe0073583827c37f6639f7194754c5db20
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.2

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page