Skip to main content

audr-adapter-litellm

PyPI Python versions

The AUDR callback for LiteLLM turns every provider-reported completion, Responses API, embedding and rerank call made through the LiteLLM Python SDK or Router into one AUDR record for an audr.Client your application owns. It reads usage, cost, identifiers and timings only: never prompts, messages, responses, tool arguments, API keys or exception messages.

Status: alpha. The record model tracks AUDR v1.0.0; until 1.0.0, a minor release may change the public API.

Setup

pip install "audr-adapter-litellm[runtime]"

Requires Python 3.11 or later and litellm 1.95 or later within 1.x, which the runtime extra installs. The LiteLLM Proxy is not supported: it exposes no documented callback shutdown hook after which the adapter's records are guaranteed to have reached the client.

Usage

Construct the callback on the running event loop that owns the Client and register it through LiteLLM's callback manager, which keeps any other registered callbacks:

import asyncio

import litellm
from audr import Attribution, Client, FileSink
from audr_adapter_litellm import LiteLLMAudrCallback, LiteLLMConfig


async def main() -> None:
    async with Client(FileSink("audr.jsonl")) as client:
        callback = LiteLLMAudrCallback(
            client=client,
            config=LiteLLMConfig(attribution_defaults=Attribution(environment="production")),
        )
        litellm.logging_callback_manager.add_litellm_callback(callback)
        try:
            await litellm.acompletion(
                model="openai/gpt-4o-mini",
                messages=[{"role": "user", "content": "Where is order 42?"}],
                metadata={"audr": {"attribution": {"account_id": "acct_42"}}},
            )
            # LiteLLM invokes logging callbacks after the call returns; wait up to 5 s.
            for _ in range(500):
                if client.stats.submitted:
                    break
                await asyncio.sleep(0.01)
        finally:
            litellm.logging_callback_manager.remove_callback_from_all_lists(callback)
            await callback.drain(timeout=5)
            callback.close()


asyncio.run(main())

Leaving the async with block drains the client's queue and closes its sink. Runnable versions on mock responses, without network access or provider keys, are examples/litellm_completion.py and examples/router_fallback.py.

Attribution

Attribution is resolved for each call from metadata["audr"]["attribution"] merged field by field over LiteLLMConfig.attribution_defaults. The request wins, and labels merge by key:

from audr import Attribution
from audr_adapter_litellm import LiteLLMConfig

config = LiteLLMConfig(
    attribution_defaults=Attribution(environment="production", labels={"region": "us"}),
)
metadata = {
    "audr": {
        "attribution": {
            "account_id": "acct_42",
            "user_id": "u_8f14e45f",  # pseudonymous, never an email or a name
            "labels": {"feature": "support-chat"},
        },
        "run": {"run_id": "agent-run-123", "run_type": "agent_run"},
    }
}
# A call made with this metadata carries labels {"region": "us", "feature": "support-chat"}.
  • metadata["audr"] accepts attribution, run and resource objects, and each rejects unknown fields. The reference lists every field.
  • A call with no environment, or a production call with no account_id, is skipped with a warning rather than billed to a guess.
  • Pass the agent's run_id so model calls from one agent execution join downstream. Without it, the LiteLLM trace or call identifier becomes run.run_id.

Records

LiteLLM call resource.operation usage
completion, text_completion, Responses API generation llm tokens, requests: 1
embedding embedding llm tokens, requests: 1
rerank reranking llm tokens, requests: 1

Cache and reasoning tokens are counted apart from input_tokens and output_tokens, and LiteLLM's response_cost becomes cost.total_cost in USD. The provider-echoed model name wins over the requested LiteLLM alias. A Router fallback chain produces one record, for the attempt that reported usage.

Not metered: cache hits, failures and abandoned streams without reported usage or a positive cost, and tool executions, which LiteLLM returns to the application rather than running.

Documentation

  • Reference: options, request metadata, record fields, shutdown order, diagnostics, operational bounds
  • Examples: runnable on mock responses, without network access
  • Changelog
  • AUDR specification, which defines every record field

License

Apache-2.0. Contributions follow CONTRIBUTING.md.

Metadata

Release files for audr-adapter-litellm 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for audr-adapter-litellm 0.2.0
File Size Uploaded
audr_adapter_litellm-0.2.0.tar.gz 15.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for audr-adapter-litellm 0.2.0
File Interpreter ABI Platform
audr_adapter_litellm-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 34.4 kB

Release files / audr_adapter_litellm-0.2.0.tar.gz

Download URL audr_adapter_litellm-0.2.0.tar.gz
Size 15.9 kB
Tags Source
SHA-256 checksum
How to use checksums
25709d85e1179e6e876184d45cded0535e8d2d722443cfe41fc4a1fd23e06022
BLAKE2b-256 checksum
How to use checksums
eed614408f14cd18e3146ea1472d63e132eb259b2317f5d3a11efc6a77e63092
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release files / audr_adapter_litellm-0.2.0-py3-none-any.whl

Download URL audr_adapter_litellm-0.2.0-py3-none-any.whl
Size 18.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
46157e812f80caa87b6b3d38003fea09efbd5f3739360299b1b4f77a5cf89eda
BLAKE2b-256 checksum
How to use checksums
95a7c0ded7ce7074f460bae61b9a02147662767828cde2d38f7d39b0b4aec0ad
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page