Skip to main content

Agent Observability for LiteLLM

The agento11y-litellm package sends what your LiteLLM calls do to Grafana Cloud Agent Observability. Register one callback and every completion is exported as a generation, with its messages, tokens, cost, tool calls, and errors. Add a second callback and Agent Observability guards can block a request before it reaches the provider.

It works in your own process and inside the LiteLLM proxy. Guards need the proxy, because LiteLLM only runs call hooks there.

Before you begin

You need Python 3.10 or later, and an endpoint, instance ID, and token. To enable the product and collect those three values, refer to Set up Agent Observability.

Install the packages:

pip install agento11y agento11y-litellm litellm

Export generations from your app

The client reads its connection details from the environment, so Client() takes no arguments:

export AGENTO11Y_ENDPOINT=https://your-agento11y.grafana.net
export AGENTO11Y_PROTOCOL=http
export AGENTO11Y_AUTH_MODE=basic
export AGENTO11Y_AUTH_TENANT_ID=your-instance-id
export AGENTO11Y_AUTH_TOKEN=glc_your_token

Grafana Cloud needs both AGENTO11Y_PROTOCOL=http and AGENTO11Y_AUTH_MODE=basic. The SDK otherwise defaults to gRPC with no authentication, and a Cloud endpoint then answers 401 with no other signal. For content capture, batching, and the rest of the settings, refer to Configure the Agent Observability SDK.

To export generations, follow these steps:

  1. Create a client and a handler.

    import litellm
    from agento11y import Client
    from agento11y_litellm import Agento11yLiteLLMLogger
    
    client = Client()
    litellm.callbacks = [Agento11yLiteLLMLogger(client=client, agent_name="my-agent")]
    
  2. Call LiteLLM as you normally do.

    response = litellm.completion(
        model="openai/gpt-4o-mini",
        messages=[{"role": "user", "content": "Hello!"}],
    )
    print(response.choices[0].message.content)
    
  3. Flush buffered telemetry before your process exits.

    client.shutdown()
    

Streaming works the same way. A streamed call is exported as one generation in STREAM mode, carrying the time of the first token:

response = litellm.completion(
    model="openai/gpt-4o-mini",
    messages=[{"role": "user", "content": "Give me three reliability tips."}],
    stream=True,
)
for chunk in response:
    content = chunk.choices[0].delta.content
    if content:
        print(content, end="", flush=True)

Your generations appear under the agent name you passed. To find and read them, refer to Browse and debug conversations. For every option the handler takes, refer to the LiteLLM adapter reference.

Export generations from the LiteLLM proxy

The proxy loads callbacks by dotted path from a Python file next to config.yaml. Put a Client and a handler in that file, then name the handler in config.yaml:

litellm_settings:
  callbacks:
    - agento11y_callback.agento11y_handler

For the callback file, the Dockerfile, and the docker run command, refer to Deploy in the LiteLLM proxy. For a proxy you can start in one command, refer to the LiteLLM proxy example.

Name the calling agent

One proxy usually serves several agents. The handler reads the agent name from each request, so each caller gets its own agent in Agent Observability. When a request names no agent, the handler falls back to the agent_name you configured.

A client that knows about Agent Observability can pass metadata:

response = litellm.completion(
    model="openai/gpt-4o-mini",
    messages=[{"role": "user", "content": "Continue our chat."}],
    metadata={
        "agent_name": "search-agent",
        "agent_version": "v2",
        "conversation_id": "conv-abc-123",
    },
)

A client that doesn't can send the header LiteLLM already understands:

curl http://localhost:4000/v1/chat/completions \
  -H 'x-litellm-agent-id: search-agent' \
  -H 'Content-Type: application/json' \
  -d '{"model": "gpt-4o-mini", "messages": [{"role": "user", "content": "Hello!"}]}'

Both work on every recorded route. To name callers after their virtual key instead, or to learn which metadata containers the adapter reads, refer to Agent identity.

Enforce guards on proxy requests

A guard is a rule that runs on the request path and can fail the request. Agento11yLiteLLMGuardrail evaluates your Agent Observability guards inside the proxy: preflight before the provider is called, postflight against the provider's response.

Guards live in Agent Observability, not in config.yaml. To create one, refer to Set up guards.

To enforce them, follow these steps:

  1. Enable hooks on the client and build the guardrail next to the handler.

    from agento11y import Client, ClientConfig, HooksConfig
    from agento11y_litellm import Agento11yLiteLLMGuardrail, Agento11yLiteLLMLogger
    
    client = Client(
        ClientConfig(
            hooks=HooksConfig(
                enabled=True,
                timeout_seconds=3.0,
                # Drop "postflight" to evaluate request rules only.
                phases=["preflight", "postflight"],
            )
        )
    )
    
    agento11y_handler = Agento11yLiteLLMLogger(client=client, agent_name="litellm-proxy")
    agento11y_guards = Agento11yLiteLLMGuardrail(
        client=client,
        agent_name="litellm-proxy",
        default_on=True,
        event_hook=["pre_call", "post_call"],
    )
    
  2. List both objects in config.yaml.

    litellm_settings:
      callbacks:
        - agento11y_callback.agento11y_handler
        - agento11y_callback.agento11y_guards
    
  3. Send a request that a rule denies, and confirm the proxy answers 400 with the rule's reason.

LiteLLM runs any CustomGuardrail in litellm.callbacks on both hook paths, so the guardrail needs no separate guardrails block.

Three limits are worth knowing before you rely on guards:

  • A postflight deny can't stop a streamed response, because the caller already has it. Use preflight to block a streaming request.
  • A redact rule rewrites the request all or nothing. When the guardrail can't apply the whole rewrite, it forwards the original and logs a warning starting with agento11y: skipping.
  • A deny answers 400 from LiteLLM 1.87.0 on, and 500 before that.

For the routes each phase covers, the deny outcome per delivery, and every guardrail option, refer to the LiteLLM guard reference.

Learn more

Metadata

Release files for agento11y-litellm 0.18.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for agento11y-litellm 0.18.0
File Size Uploaded
agento11y_litellm-0.18.0.tar.gz 65.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for agento11y-litellm 0.18.0
File Interpreter ABI Platform
agento11y_litellm-0.18.0-py3-none-any.whl Python 3 none any Details

Total release size: 99.2 kB

Release files / agento11y_litellm-0.18.0.tar.gz

Download URL agento11y_litellm-0.18.0.tar.gz
Size 65.3 kB
Tags Source
SHA-256 checksum
How to use checksums
d1d572c59a74001342b64769dfa939efc04a7a64baba7784b72c87948b17bfc9
BLAKE2b-256 checksum
How to use checksums
ea812c3050dcf8ac17202a17337948be9cbbe6b5bc65ee916ed3a124c8dfad53
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / agento11y_litellm-0.18.0-py3-none-any.whl

Download URL agento11y_litellm-0.18.0-py3-none-any.whl
Size 33.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0bb016a4dcd261031f9f455eddb4bc8c18df4c01b616fe6d2ab7651f3a3229fc
BLAKE2b-256 checksum
How to use checksums
b922c9f8dcc5f170a77fd8a4025668eb07daf4d99f093338ef297e0ce8a5c9a4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.18.0 This release

2 release files

0.17.0

2 release files

0.16.0

2 release files

0.15.0

2 release files

0.14.0

2 release files

0.12.0

2 release files

0.10.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page