Skip to main content

Vantage LiteLLM Callback

Required Vantage collector

The callback works with the Vantage LiteLLM collector. To set up the Vantage LiteLLM integration, install both components:

  1. Run the collector beside each LiteLLM proxy. The collector stores usage events and uploads compacted batches to Vantage.
  2. Install this callback in the LiteLLM Python environment. The callback sends usage events to the local collector over its Unix socket.

The callback cannot upload usage to Vantage by itself. The collector cannot collect LiteLLM usage without the callback. See the collector repository README for deployment and configuration instructions.

Install

Install a released version with:

pip install vantage-litellm-callback

To work from a source checkout, install this directory with pip install .. Configure LiteLLM with:

litellm_settings:
  callbacks:
    - vantage_callback.callback_instance

See the repository README for collector setup and lifecycle billing semantics.

The callback is fail-open: collector outages, acknowledgement failures, malformed responses, queue pressure, and projection errors never block or fail a provider request. Usage events are placed into a bounded in-process queue and delivered by a background worker. When delivery is unhealthy, the worker opens a circuit, aggregates bounded loss information, and periodically probes the collector with a versioned delivery_gap control frame. Normal delivery resumes only after that frame receives a durable acknowledgement.

Configuration:

  • VANTAGE_COLLECTOR_DELIVERY_QUEUE_SIZE (default 1024)
  • VANTAGE_COLLECTOR_RECOVERY_PROBE_SECONDS (default 5)
  • VANTAGE_COLLECTOR_GAP_EVENT_ID_SAMPLE_SIZE (default 32)
  • VANTAGE_COLLECTOR_GAP_SUMMARY_INTERVAL_SECONDS (default 60)
  • VANTAGE_COLLECTOR_SOCKET_PATH (default /var/run/vantage-collector/collector.sock)
  • VANTAGE_COLLECTOR_ACK_TIMEOUT_SECONDS (default 5)
  • VANTAGE_COLLECTOR_CONNECTION_POOL_SIZE (default 16)

The queue and gap aggregate are intentionally process-local and bounded. A process exit may therefore lose queued events; the gap frame reports losses observed while the process remains alive.

Running the provider matrix locally

tests/test_provider_matrix.py is a regression matrix over the provider and mode combinations whose usage shapes disagree with one another. Every cell is a recorded fixture, so the matrix needs no API keys and runs on every pull request:

pip install --editable . pytest pytest-asyncio
PYTHONPATH=. pytest tests/test_provider_matrix.py -v

Run a single cell while iterating on a provider:

PYTHONPATH=. pytest tests/test_provider_matrix.py -k anthropic

Fixtures live in tests/fixtures/, one JSON file per cell:

Directory Cells
nonstreaming/ OpenAI chat + cache, OpenAI multimodal, Anthropic post-calculate_usage(), Bedrock converse, Responses API + cache (with and without a text_tokens detail), OCR / non-token
streaming/ OpenAI SSE, Anthropic /v1/messages SSE, Bedrock converse, Responses API response.completed

Each file carries its usage payload, the expected_usage projection, the expected_billable_total, and a note citing the LiteLLM transform the shape was taken from. Adding a provider means adding a JSON file — the tests parametrize over the directory, so a new file becomes a new cell with no test changes.

The matrix asserts three things per cell: the four projected Usage fields, that those fields are mutually exclusive (they must sum to the billable total, so no token is counted twice), and that the callback emits exactly one started and one terminal event. Two further tests assert that nothing the callback attaches to a request reaches the outbound provider body.

Two token-semantics rules are what the matrix exists to protect, both of which have regressed before:

  • The input count is gross wherever the cache is reported nested. LiteLLM's transforms fold cache reads and writes into prompt_tokens, and the Responses API counts input_tokens_details.cached_tokens inside input_tokens. Those buckets have to be subtracted or the cached tokens bill twice. Raw Anthropic usage is the exception: it puts cache_*_input_tokens beside a net input_tokens, so nothing is subtracted there.
  • text_tokens is not "uncached input". It means the text modality for OpenAI but the net raw input for Anthropic. Reading it verbatim silently drops image, video, and audio input tokens, so uncached input is derived from the gross count instead.

For an end-to-end check against a live proxy on :4000, demo/send_requests.sh sends tagged requests through LiteLLM; see demo/README.md. That path needs a running proxy and collector and is a manual smoke test, not part of CI.

Metadata

Release files for vantage-litellm-callback 0.0.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for vantage-litellm-callback 0.0.4
File Size Uploaded
vantage_litellm_callback-0.0.4.tar.gz 31.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for vantage-litellm-callback 0.0.4
File Interpreter ABI Platform
vantage_litellm_callback-0.0.4-py3-none-any.whl Python 3 none any Details

Total release size: 50.9 kB

Release files / vantage_litellm_callback-0.0.4.tar.gz

Download URL vantage_litellm_callback-0.0.4.tar.gz
Size 31.1 kB
Tags Source
SHA-256 checksum
How to use checksums
cbbe7852744b7bf86613fedc6967debd2f9977cf69b47aa90aaf6b1bdd554c2e
BLAKE2b-256 checksum
How to use checksums
b7207ee3233a4549ed997b7a893e6680a4396217fe15413775645e859e96b80c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.

Transparency log

Release files / vantage_litellm_callback-0.0.4-py3-none-any.whl

Download URL vantage_litellm_callback-0.0.4-py3-none-any.whl
Size 19.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
aafb87ad254d91996645080d55e81c8c78d23196001ed7f0f9888e61cb7dc989
BLAKE2b-256 checksum
How to use checksums
0ab23466d0da4d1ad91436d12979208ae5c75fcc8222a83e135bc09e5299ca07
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.

Transparency log

Release history Release notifications | RSS feed

0.0.7

2 release files

0.0.6

2 release files

0.0.5

2 release files

This release

0.0.4 This release

2 release files

0.0.3

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page