Vantage LiteLLM Callback
Required Vantage collector
The callback works with the Vantage LiteLLM collector. To set up the Vantage LiteLLM integration, install both components:
- Run the collector beside each LiteLLM proxy. The collector stores usage events and uploads compacted batches to Vantage.
- Install this callback in the LiteLLM Python environment. The callback sends usage events to the local collector over its Unix socket.
The callback cannot upload usage to Vantage by itself. The collector cannot collect LiteLLM usage without the callback. See the collector repository README for deployment and configuration instructions.
Install
Install a released version with:
pip install vantage-litellm-callback
To work from a source checkout, install this directory with pip install ..
Configure LiteLLM with:
litellm_settings:
callbacks:
- vantage_callback.callback_instance
See the repository README for collector setup and lifecycle billing semantics.
The callback is fail-open: collector outages, acknowledgement failures, malformed
responses, queue pressure, and projection errors never block or fail a provider
request. Usage events are placed into a bounded in-process queue and delivered by
a background worker. When delivery is unhealthy, the worker opens a circuit,
aggregates bounded loss information, and periodically probes the collector with a
versioned delivery_gap control frame. Normal delivery resumes only after that
frame receives a durable acknowledgement.
Configuration:
VANTAGE_COLLECTOR_DELIVERY_QUEUE_SIZE(default1024)VANTAGE_COLLECTOR_RECOVERY_PROBE_SECONDS(default5)VANTAGE_COLLECTOR_GAP_EVENT_ID_SAMPLE_SIZE(default32)VANTAGE_COLLECTOR_GAP_SUMMARY_INTERVAL_SECONDS(default60)VANTAGE_COLLECTOR_SOCKET_PATH(default/var/run/vantage-collector/collector.sock)VANTAGE_COLLECTOR_ACK_TIMEOUT_SECONDS(default5)VANTAGE_COLLECTOR_CONNECTION_POOL_SIZE(default16)
The queue and gap aggregate are intentionally process-local and bounded. A process exit may therefore lose queued events; the gap frame reports losses observed while the process remains alive.
Running the provider matrix locally
tests/test_provider_matrix.py is a regression matrix over the provider and
mode combinations whose usage shapes disagree with one another. Every cell is a
recorded fixture, so the matrix needs no API keys and runs on every pull
request:
pip install --editable . pytest pytest-asyncio
PYTHONPATH=. pytest tests/test_provider_matrix.py -v
Run a single cell while iterating on a provider:
PYTHONPATH=. pytest tests/test_provider_matrix.py -k anthropic
Fixtures live in tests/fixtures/, one JSON file per cell:
| Directory | Cells |
|---|---|
nonstreaming/ |
OpenAI chat + cache, OpenAI multimodal, Anthropic post-calculate_usage(), Bedrock converse, Responses API + cache (with and without a text_tokens detail), OCR / non-token |
streaming/ |
OpenAI SSE, Anthropic /v1/messages SSE, Bedrock converse, Responses API response.completed |
Each file carries its usage payload, the expected_usage projection, the
expected_billable_total, and a note citing the LiteLLM transform the shape
was taken from. Adding a provider means adding a JSON file — the tests
parametrize over the directory, so a new file becomes a new cell with no test
changes.
The matrix asserts three things per cell: the four projected Usage fields,
that those fields are mutually exclusive (they must sum to the billable total,
so no token is counted twice), and that the callback emits exactly one started
and one terminal event. Two further tests assert that nothing the callback
attaches to a request reaches the outbound provider body.
Two token-semantics rules are what the matrix exists to protect, both of which have regressed before:
- The input count is gross wherever the cache is reported nested. LiteLLM's
transforms fold cache reads and writes into
prompt_tokens, and the Responses API countsinput_tokens_details.cached_tokensinsideinput_tokens. Those buckets have to be subtracted or the cached tokens bill twice. Raw Anthropic usage is the exception: it putscache_*_input_tokensbeside a netinput_tokens, so nothing is subtracted there. text_tokensis not "uncached input". It means the text modality for OpenAI but the net raw input for Anthropic. Reading it verbatim silently drops image, video, and audio input tokens, so uncached input is derived from the gross count instead.
For an end-to-end check against a live proxy on :4000, demo/send_requests.sh
sends tagged requests through LiteLLM; see demo/README.md. That path needs a
running proxy and collector and is a manual smoke test, not part of CI.
Metadata
Release files for vantage-litellm-callback 0.0.4
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| vantage_litellm_callback-0.0.4.tar.gz | 31.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| vantage_litellm_callback-0.0.4-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 50.9 kB
Release files / vantage_litellm_callback-0.0.4.tar.gz
| Download URL | vantage_litellm_callback-0.0.4.tar.gz |
|---|---|
| Size | 31.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
cbbe7852744b7bf86613fedc6967debd2f9977cf69b47aa90aaf6b1bdd554c2e
|
|
BLAKE2b-256 checksum How to use checksums |
b7207ee3233a4549ed997b7a893e6680a4396217fe15413775645e859e96b80c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.
Transparency logRelease files / vantage_litellm_callback-0.0.4-py3-none-any.whl
| Download URL | vantage_litellm_callback-0.0.4-py3-none-any.whl |
|---|---|
| Size | 19.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
aafb87ad254d91996645080d55e81c8c78d23196001ed7f0f9888e61cb7dc989
|
|
BLAKE2b-256 checksum How to use checksums |
0ab23466d0da4d1ad91436d12979208ae5c75fcc8222a83e135bc09e5299ca07
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.
Transparency log