axonflow-litellm
AxonFlow governance integration for LiteLLM. Enforce policies, audit LLM calls, and gate high-risk requests behind human approval — all through a drop-in wrapper around litellm.completion().
Installation
pip install axonflow-litellm
Quick Start
from axonflow_litellm import AxonFlowLogger, AxonFlowLoggerConfig, PolicyDeniedError
logger = AxonFlowLogger(
AxonFlowLoggerConfig(
endpoint="http://localhost:8080",
client_id="my-app",
client_secret="...",
)
)
try:
response = logger.completion(
model="gpt-4o",
messages=[{"role": "user", "content": "Summarize quarterly earnings"}],
)
print(response.choices[0].message.content)
except PolicyDeniedError as e:
print(f"Blocked: {e.reason}")
How It Works
AxonFlowLogger provides two integration modes:
Governance Mode (recommended)
Use logger.completion() or logger.acompletion() as drop-in replacements for litellm.completion() / litellm.acompletion():
- Pre-check — sends the prompt to AxonFlow for policy evaluation
- HITL — if the policy returns
require_approval, creates a human-in-the-loop review request and polls until approved, rejected, or timed out - LLM call — delegates to LiteLLM (all providers supported)
- Audit — records the response to AxonFlow for observability
# Async (recommended for production)
response = await logger.acompletion(
model="claude-sonnet-4-6",
messages=[{"role": "user", "content": "..."}],
user_token="jwt-from-your-auth",
)
Audit-Only Mode
Register as a LiteLLM callback for observability without blocking:
import litellm
litellm.callbacks = [logger]
response = litellm.acompletion(model="gpt-4o", messages=[...])
In this mode, every LLM call is recorded to AxonFlow for audit trail. Policy denials are logged as warnings but cannot block the request (a LiteLLM SDK limitation — callback exceptions are silently swallowed).
Configuration
| Parameter | Default | Description |
|---|---|---|
endpoint |
(required) | AxonFlow agent URL |
client_id |
(required) | AxonFlow client identifier |
client_secret |
"" |
AxonFlow client secret |
default_user_token |
"anonymous" |
Token for policy evaluation when none provided — accepted only by community-mode deployments; see User tokens |
tenant_id |
None |
AxonFlow tenant identifier |
fail_open |
True |
Allow LLM calls when AxonFlow is unreachable |
call_timeout_seconds |
5.0 |
Per-hook timeout for AxonFlow API calls |
breaker_failure_threshold |
5 |
Consecutive failures before circuit opens |
breaker_recovery_seconds |
30.0 |
Wait before attempting recovery probe |
enable_hitl_polling |
True |
Enable HITL approval flow for require_approval |
approval_poll_interval_seconds |
2.0 |
Polling interval for HITL status |
approval_max_wait_seconds |
300.0 |
Maximum wait for HITL decision |
extra_context |
{} |
Additional context sent with every pre-check |
User tokens: community vs. enterprise
What user_token must be depends on the AxonFlow deployment mode:
-
Community / community-SaaS — the token is not validated. The
default_user_token="anonymous"placeholder works out of the box. -
Enterprise / evaluation — the platform validates
user_tokenon/api/policy/pre-checkand rejects both absent and invalid tokens (including the"anonymous"placeholder) with 401. Pass a real per-user token on every call (or setdefault_user_tokento one):response = await logger.acompletion( model="gpt-4o", messages=[...], user_token=minted_token, # per-user token minted by your admin )
Admins mint per-user tokens via the customer-portal admin API (
POST /api/v1/admin/organizations/{org_id}/user-tokens) — see the per-user token provisioning guide. Admin-minted (HS256) tokens only: the pre-check plane pins the accepted algorithm to HS256, so tenant-OIDC access tokens (RS256) are rejected there. The audit trail then attributes each LLM call to that user.
A platform rejection (401/402/403 — bad credentials, rejected
user_token, budget block, tenant mismatch) raises PolicyDeniedError
whenever the pre-check reaches the platform, regardless of fail_open.
fail_open covers availability only: a healthy platform refusing the
request is a governance verdict, not an outage, and proceeding would
silently skip governance on every call. (In audit-only callback mode the
error is logged but cannot block — a LiteLLM callback limitation.)
Fail-Open vs. Fail-Closed
By default, fail_open=True: if AxonFlow is unreachable or times out, the LLM call proceeds normally. This ensures an AxonFlow outage does not break your application.
For high-stakes workloads where unapproved LLM calls must never proceed:
config = AxonFlowLoggerConfig(
endpoint="http://localhost:8080",
client_id="payments-service",
client_secret="...",
fail_open=False,
)
Sync vs. Async
Both litellm.completion() (sync) and litellm.acompletion() (async) are fully supported.
When registered via litellm.callbacks, sync hooks delegate to their async counterparts via asyncio.run(). This adds minor overhead (~1ms) per hook call in the sync path. For performance-critical sync workloads, use logger.completion() directly (governance wrapper) which amortizes the event loop creation.
If sync hooks are invoked inside a running event loop (unusual — e.g., sync callbacks from an async framework), a one-time RuntimeWarning is emitted directing you to acompletion().
Sync callback mode caveats
In sync callback mode (litellm.callbacks = [logger] + litellm.completion()), each callback hook creates an ephemeral asyncio event loop via asyncio.run(). Pre-check (governance) and post-LLM audit both fire and write to AxonFlow. However:
- Audit write failures are logged at WARNING level and do not raise to the caller (fail-open by default). If AxonFlow is temporarily unreachable during the audit phase, the LLM response is still returned but the audit row may be missing.
- Each hook creates a new event loop, so connection pooling is not shared across hooks within the same LLM call. This is slightly less efficient than the governance wrapper path.
For strict audit guarantees (every LLM call audited, failure = exception), use logger.completion() or logger.acompletion() instead of the callback registration path.
Exceptions
| Exception | When |
|---|---|
PolicyDeniedError |
Policy denied the request |
ApprovalRejected |
HITL approval was rejected |
ApprovalTimeout |
HITL approval timed out |
All exceptions carry .reason (string) and .policies (list of policy IDs).
These exceptions do NOT extend litellm.exceptions.APIError — catch governance denials via PolicyDeniedError, not LiteLLM's exception hierarchy.
MCP Governance
LiteLLM is LLM-completion-focused. For MCP tool governance, use AxonFlow's MCP server directly.
Requirements
- Python >= 3.10
litellm>= 1.40axonflow>= 8.2.0
Telemetry
This integration declares itself on the AxonFlow SDK's existing anonymous heartbeat, so aggregate adoption figures can tell LiteLLM-governed usage apart from bare SDK usage. Without it the two are indistinguishable: both report the same sdk, the same sdk_version and the same endpoint.
It adds no network request. adapter:litellm rides the features array of the heartbeat the SDK already sends — there is no second ping, no second endpoint, and no new configuration surface. The declaration happens once, immediately before the AxonFlow client is constructed, so it reaches the very first heartbeat.
What is and is not collected: the string litellm, and nothing else. No prompts, no completions, no model names, no user identities, no configuration. Everything else on the heartbeat is the SDK's own — see the AxonFlow Python SDK's telemetry section for the full field list and the opt-out.
AXONFLOW_TELEMETRY=off suppresses the heartbeat, and this declaration with it. An SDK older than 9.3.0 simply has nothing to declare to; the integration works normally either way.
License
MIT
Metadata
Release files for axonflow-litellm 1.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| axonflow_litellm-1.1.0.tar.gz | 31.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| axonflow_litellm-1.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 50.8 kB
Release files / axonflow_litellm-1.1.0.tar.gz
| Download URL | axonflow_litellm-1.1.0.tar.gz |
|---|---|
| Size | 31.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1c0843ebc890e53f8418bca4e6b59eb46be8bf6c1236105a12aa041645a09ab8
|
|
BLAKE2b-256 checksum How to use checksums |
f96c9f8beb4052f6a3d92057fafcaaac3cf78ee555245f845458534e44f2ff3a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 7, 2026.
Transparency logRelease files / axonflow_litellm-1.1.0-py3-none-any.whl
| Download URL | axonflow_litellm-1.1.0-py3-none-any.whl |
|---|---|
| Size | 19.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
14d6dc741b897e58cee0c65bd3719ab5536b8f4aaa09ca881dd53449479b06d6
|
|
BLAKE2b-256 checksum How to use checksums |
af924a1621b85dae6d646be96a46b7dbfbb7c36302dd5c55e1ccc7b0bdf28763
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 7, 2026.
Transparency log