Skip to main content

Grafana Agent Observability Python SDK

agento11y records normalized LLM generation and tool-execution telemetry. It exports normalized generations to Agent Observability ingest and uses your OpenTelemetry tracer/meter setup for traces and metrics.

Use this package when you want:

  • A provider-agnostic generation record (same schema for OpenAI, Anthropic, Gemini, or custom adapters).
  • OTel-aligned tracing attributes for generation and tool spans.
  • Async export with retry/backoff, queueing, batching, and explicit shutdown semantics.

Installation

pip install agento11y

For a Grafana Cloud setup walkthrough (where to find the endpoint URL, instance ID, and API token), refer to the Grafana Cloud setup guide.

Validation

Run the shared core conformance suite for the Python SDK from the repo root:

mise run test:py:sdk-conformance

Run the cross-language aggregate core conformance suite from the repo root:

mise run sdk:conformance

Optional provider helper packages:

pip install agento11y-openai
pip install agento11y-anthropic
pip install agento11y-gemini

Optional framework modules:

pip install agento11y-langchain
pip install agento11y-langgraph
pip install agento11y-openai-agents
pip install agento11y-llamaindex
pip install agento11y-google-adk
pip install agento11y-strands
pip install agento11y-claude-agent-sdk
pip install agento11y-litellm
pip install agento11y-pydantic-ai

Framework handler usage:

from agento11y import Client
from agento11y_langchain import with_agento11y_langchain_callbacks
from agento11y_langgraph import with_agento11y_langgraph_callbacks
from agento11y_openai_agents import with_agento11y_openai_agents_hooks
from agento11y_llamaindex import with_agento11y_llamaindex_callbacks
from agento11y_google_adk import with_agento11y_google_adk_callbacks
from agento11y_strands import with_agento11y_strands_hooks
from agento11y_claude_agent import with_agento11y_claude_agent_options
from agento11y_pydantic_ai import with_agento11y_pydantic_ai_capability

client = Client()
chain_config = with_agento11y_langchain_callbacks(None, client=client, provider_resolver="auto")
graph_config = with_agento11y_langgraph_callbacks(None, client=client, provider_resolver="auto")
openai_agents_run_options = with_agento11y_openai_agents_hooks(None, client=client, provider_resolver="auto")
llamaindex_config = with_agento11y_llamaindex_callbacks(None, client=client, provider_resolver="auto")
google_adk_agent_config = with_agento11y_google_adk_callbacks(None, client=client, provider_resolver="auto")
strands_agent_config = with_agento11y_strands_hooks(None, client=client, provider_resolver="auto")
claude_agent_options = with_agento11y_claude_agent_options(None, client=client)
pydantic_ai_capabilities = with_agento11y_pydantic_ai_capability(None, client=client, provider_resolver="auto")

LiteLLM uses a callback class instead of a with_agento11y_* helper:

import litellm
from agento11y import Client
from agento11y_litellm import Agento11yLiteLLMLogger

client = Client()
litellm.callbacks = [Agento11yLiteLLMLogger(client=client)]

Framework handlers use the Client instance you pass in. If that client is configured with generation_sanitizer, the same redaction policy applies automatically to generations recorded through LangChain, LangGraph, OpenAI Agents, LlamaIndex, Google ADK, Strands, Claude Agent SDK, LiteLLM, and Pydantic AI integrations.

Framework handlers inject framework tags/metadata on recorded generations:

  • agento11y.framework.name (langchain, langgraph, openai-agents, llamaindex, google-adk, strands, claude-agent-sdk, litellm, or pydantic-ai)
  • agento11y.framework.source=handler (or hooks for Strands Agents and Claude Agent SDK)
  • agento11y.framework.language=python
  • metadata["agento11y.framework.run_id"]
  • metadata["agento11y.framework.thread_id"] (when present)
  • metadata["agento11y.framework.parent_run_id"] (when available)
  • metadata["agento11y.framework.component_name"]
  • metadata["agento11y.framework.run_type"]
  • metadata["agento11y.framework.tags"]
  • metadata["agento11y.framework.retry_attempt"] (when available)
  • metadata["agento11y.framework.event_id"] (when available)
  • metadata["agento11y.framework.langgraph.node"] (LangGraph when available)

Conversation mapping is conversation-first:

  • conversation_id / session_id / group_id from framework context first
  • then thread_id
  • deterministic fallback agento11y:framework:<framework_name>:<run_id>

When present in generation metadata, low-cardinality framework keys are copied onto generation span attributes.

For LangGraph persistence, pass configurable.thread_id and reuse it across invocations:

thread_config = {
    **with_agento11y_langgraph_callbacks(None, client=client, provider_resolver="auto"),
    "configurable": {"thread_id": "customer-42"},
}
graph.invoke({"prompt": "Remember my timezone is UTC+1.", "answer": ""}, config=thread_config)
graph.invoke({"prompt": "What timezone did I give you?", "answer": ""}, config=thread_config)

Full framework examples:

  • LangChain: ../python-frameworks/langchain/README.md
  • LangGraph: ../python-frameworks/langgraph/README.md
  • OpenAI Agents: ../python-frameworks/openai-agents/README.md
  • LlamaIndex: ../python-frameworks/llamaindex/README.md
  • Google ADK: ../python-frameworks/google-adk/README.md
  • Strands Agents: ../python-frameworks/strands/README.md
  • Claude Agent SDK: ../python-frameworks/claude-agent-sdk/README.md
  • LiteLLM: ../python-frameworks/litellm/README.md
  • Pydantic AI: ../python-frameworks/pydantic-ai/README.md

Quick Start (Sync Generation)

Client() reads AGENTO11Y_* env vars by default. See the Grafana Cloud setup guide for the variable names. Pass an explicit ClientConfig only when you need to override.

from agento11y import (
    Client,
    GenerationStart,
    ModelRef,
    assistant_text_message,
    user_text_message,
)

client = Client()  # reads AGENTO11Y_* env vars

with client.start_generation(
    GenerationStart(
        conversation_id="conv-1",
        agent_name="my-service",
        agent_version="1.0.0",
        model=ModelRef(provider="openai", name="gpt-5"),
    )
) as rec:
    rec.set_result(
        input=[user_text_message("What is the weather in Paris?")],
        output=[assistant_text_message("It is 18C and sunny.")],
    )

    # Recorder errors are local SDK errors (validation/enqueue/shutdown),
    # not provider call failures.
    if rec.err() is not None:
        raise rec.err()

client.shutdown()

Explicit configuration form:

import os
from agento11y import AuthConfig, Client, ClientConfig, GenerationExportConfig

client = Client(
    ClientConfig(
        generation_export=GenerationExportConfig(
            protocol="http",
            endpoint="https://agento11y-prod-<region>.grafana.net",
            auth=AuthConfig(
                mode="basic",
                tenant_id=os.environ["AGENTO11Y_AUTH_TENANT_ID"],
                basic_password=os.environ["AGENTO11Y_AUTH_TOKEN"],
            ),
        ),
    )
)

Pre-Ingest Redaction

Use generation_sanitizer when you want to redact substrings from normalized generations before validation, span sync, and export.

from agento11y import (
    Client,
    ClientConfig,
    SecretRedactionOptions,
    create_secret_redaction_sanitizer,
)

client = Client(
    ClientConfig(
        generation_sanitizer=create_secret_redaction_sanitizer(
            SecretRedactionOptions(
                redact_input_messages=False,  # None falls back to AGENTO11Y_REDACT_INPUT_MESSAGES, then False
                redact_email_addresses=True,
            )
        )
    )
)

The built-in sanitizer:

  • redacts high-confidence secret formats in assistant text and thinking
  • redacts secret formats plus key/value secrets in system prompts, tool call inputs, and tool results
  • redacts email addresses by default
  • redacts the conversation title and call error
  • redacts historic assistant turns and tool messages in input
  • leaves user input unchanged unless input redaction is enabled

To preserve email addresses, opt out explicitly:

client = Client(
    ClientConfig(
        generation_sanitizer=create_secret_redaction_sanitizer(
            SecretRedactionOptions(redact_email_addresses=False)
        )
    )
)

Configuring redaction via environment variables

create_secret_redaction_sanitizer() reads AGENTO11Y_REDACT_INPUT_MESSAGES (accepts 1/0, true/false, yes/no, on/off) when redact_input_messages is left None. Precedence is explicit option > env var > False. An unrecognised env value logs a warning through the agento11y logger and falls back to the next layer, so a typo cannot silently flip redaction.

from agento11y import (
    Client,
    ClientConfig,
    create_secret_redaction_sanitizer,
)

# Leave redact_input_messages unset so AGENTO11Y_REDACT_INPUT_MESSAGES decides.
client = Client(
    ClientConfig(
        generation_sanitizer=create_secret_redaction_sanitizer(),
    )
)

Hooks and Guards

Use hooks when you want Agent Observability guard rules to run before an LLM call. The SDK evaluates the hook on your request path; guard rules configured in Grafana Cloud decide whether to allow, deny, or transform the input.

Hooks are disabled by default. Enable them on the client and call evaluate_hook(...) before the provider request:

from agento11y import (
    Client,
    ClientConfig,
    HookContext,
    HookEvaluateRequest,
    HookInput,
    HookModel,
    HookPhase,
    HooksConfig,
    Message,
    MessageRole,
    hook_denied_from_response,
    text_part,
)

client = Client(ClientConfig(hooks=HooksConfig(enabled=True)))

messages = [
    Message(role=MessageRole.USER, parts=[text_part("Summarize this customer note...")]),
]
response = client.evaluate_hook(
    HookEvaluateRequest(
        phase=HookPhase.PREFLIGHT.value,
        context=HookContext(
            agent_name="support-agent",
            agent_version="1.0.0",
            model=HookModel(provider="openai", name="gpt-5"),
            conversation_id="support-case-42",
        ),
        input=HookInput(
            messages=messages,
            system_prompt="You are a helpful support agent.",
            conversation_preview="Summarize this customer note...",
        ),
    )
)

denied = hook_denied_from_response(response)
if denied is not None:
    raise denied

if response.transformed_input is not None:
    messages = response.transformed_input.messages or messages

HooksConfig defaults to phases=["preflight"], timeout_seconds=15.0, and fail_open=True. With fail-open enabled, hook transport errors resolve to allow so an unavailable evaluator does not block production traffic. Set fail_open=False for strict paths that should fail closed.

evaluate_hook also takes a keyword-only hooks override that replaces the client's resolved hook configuration for that one call. It does not mutate the client, so later calls use the client configuration again:

from dataclasses import replace

response = client.evaluate_hook(request, hooks=replace(client.hooks_config, fail_open=False))

Use it when your caller has to tell a server allow from an allow the SDK synthesized after a failure. The response carries no marker separating the two, so a fail-open allow otherwise looks like a completed evaluation. agento11y-litellm calls fail-closed this way and applies the configured fail_open policy itself, after recording the failure as a guardrail verdict.

Set HookContext.conversation_id to the same ID used by start_generation(...). The SDK also reads with_conversation_id(...) and the active OpenTelemetry span when explicit correlation fields are omitted. This lets Agent Observability retain denied preflight attempts even though no LLM generation is created.

If you use transformed input, pass the transformed messages/system prompt to the provider and record those same values in start_generation(...). For a runnable example, see ../examples/getting-started/python-hooks/.

Configure OTEL exporters (traces/metrics) in your application OTEL SDK setup. You can optionally pass tracer and meter via ClientConfig.

Quick OTEL setup pattern before creating the agento11y client:

from opentelemetry import metrics, trace
from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.trace import TracerProvider

trace.set_tracer_provider(TracerProvider())
metrics.set_meter_provider(MeterProvider())

The providers above have no exporters attached, so nothing leaves the process. See OpenTelemetry Setup for the full wiring, including the OTLP exporters, and why analytics stays empty when this step is skipped.

Streaming Generation

Use start_streaming_generation(...) when the upstream provider call is streaming.

from agento11y import GenerationStart, ModelRef

with client.start_streaming_generation(
    GenerationStart(
        conversation_id="conv-stream",
        model=ModelRef(provider="anthropic", name="claude-sonnet-4-5"),
    )
) as rec:
    rec.set_result(output=[assistant_text_message("partial stream summary")])

Embedding Observability

Use start_embedding(...) for embedding API calls. Embedding recording emits OTel spans and SDK metrics only, and does not enqueue generation exports.

from agento11y import EmbeddingResult, EmbeddingStart, ModelRef

with client.start_embedding(
    EmbeddingStart(
        agent_name="retrieval-worker",
        agent_version="1.0.0",
        model=ModelRef(provider="openai", name="text-embedding-3-small"),
    )
) as rec:
    response = openai.embeddings.create(model="text-embedding-3-small", input=["hello", "world"])
    rec.set_result(
        EmbeddingResult(
            input_count=2,
            input_tokens=response.usage.prompt_tokens,
            input_texts=["hello", "world"],  # captured only when embedding_capture.capture_input=True
            response_model=response.model,
        )
    )

Input text capture is opt-in:

from agento11y import ClientConfig, EmbeddingCaptureConfig

cfg = ClientConfig(
    embedding_capture=EmbeddingCaptureConfig(
        capture_input=True,
        max_input_items=20,
        max_text_length=1024,
    )
)

capture_input may expose PII/document content in spans. Keep it disabled by default and enable only for scoped debugging.

TraceQL examples:

  • traces{gen_ai.operation.name="embeddings"}
  • traces{gen_ai.operation.name="embeddings" && gen_ai.request.model="text-embedding-3-small"}
  • traces{gen_ai.operation.name="embeddings" && error.type!=""}

Tool Execution Span Recording

Tool spans are recorded independently of generation export.

from agento11y import ToolExecutionStart

with client.start_tool_execution(
    ToolExecutionStart(
        tool_name="weather",
        tool_call_id="call_weather_1",
        tool_type="function",
        include_content=True,
    )
) as rec:
    rec.set_result(arguments={"city": "Paris"}, result={"temp_c": 18})

SDK identity attributes

  • Generation and tool spans always include:
    • agento11y.sdk.name=sdk-python
  • Normalized generation metadata always includes the same key.
  • If caller metadata provides a conflicting value for this key, the SDK overwrites it.

Context Defaults

Use context helpers to set defaults once per request/task boundary.

from agento11y import with_agent_name, with_agent_version, with_conversation_id

with with_conversation_id("conv-ctx"), with_agent_name("planner"), with_agent_version("2026.02"):
    with client.start_generation(
        GenerationStart(model=ModelRef(provider="gemini", name="gemini-2.5-pro"))
    ) as rec:
        rec.set_result(output=[assistant_text_message("ok")])

Content Capture Mode

ContentCaptureMode controls what content is included in exported generation payloads and OTel span attributes. Use it to prevent sensitive text (prompts, tool I/O, model responses) from leaving the process. See Content Capture Modes for the cross-SDK reference, including the per-surface behavior matrix.

Mode Generation export Generation span Tool spans Embedding span
FULL Full content Content attributes included Arguments and results included Input texts included when capture is on
NO_TOOL_CONTENT (SDK default) Full content Content attributes included Arguments and results excluded Input texts included when capture is on
METADATA_ONLY Structure only; text and tool I/O stripped Content attributes omitted Arguments and results excluded Input texts omitted
FULL_WITH_METADATA_SPANS Full content Content attributes omitted Arguments and results excluded Input texts omitted

DEFAULT is a placeholder for "inherit from the next layer"; at the client level it resolves to NO_TOOL_CONTENT. The SDK default is NO_TOOL_CONTENT, which matches the SDK's behavior before this feature was added.

FULL_WITH_METADATA_SPANS is the right mode when the gRPC ingest destination is private but the OTel trace/metric destination is shared and must not receive any content. Tool execution and embedding spans behave like METADATA_ONLY under this mode because they have no separate gRPC export.

User-provided metadata and tags are not stripped by any capture mode; callers must avoid putting sensitive content in those dicts when using METADATA_ONLY or FULL_WITH_METADATA_SPANS. SDK-internal metadata keys that carry content (e.g. call_error, agento11y.conversation.title) are stripped along with the matching content. See Tags and Metadata for where client tags, per-generation tags, metadata, and user_id each show up (export vs spans vs metrics).

Client-level default

from agento11y import Client, ClientConfig, ContentCaptureMode

client = Client(ClientConfig(
    content_capture=ContentCaptureMode.METADATA_ONLY,
))

Per-generation override

from agento11y import ContentCaptureMode, GenerationStart, ModelRef

with client.start_generation(
    GenerationStart(
        model=ModelRef(provider="openai", name="gpt-5"),
        content_capture=ContentCaptureMode.FULL,
    )
) as rec:
    rec.set_result(
        input=[user_text_message("What is the weather?")],
        output=[assistant_text_message("18C and sunny.")],
    )

Context propagation

Child tool executions inherit the active capture mode from the parent generation via ContextVar. You can also set it explicitly for a block:

from agento11y import ContentCaptureMode, with_content_capture_mode

with with_content_capture_mode(ContentCaptureMode.METADATA_ONLY):
    with client.start_tool_execution(
        ToolExecutionStart(tool_name="search")
    ) as rec:
        rec.set_result(arguments={"q": "weather"}, result={"temp_c": 18})

Dynamic resolution via resolver

A callback on ClientConfig that resolves the capture mode per-recording at runtime. Useful for feature flags, per-tenant policies, or context-dependent decisions:

from agento11y import Client, ClientConfig, ContentCaptureMode

def resolve_capture(metadata: dict) -> ContentCaptureMode:
    if metadata.get("tenant") == "healthcare":
        return ContentCaptureMode.METADATA_ONLY
    return ContentCaptureMode.DEFAULT  # fall through to client default

client = Client(ClientConfig(
    content_capture_resolver=resolve_capture,
))

Resolution precedence

For generations, highest to lowest:

  1. GenerationStart.content_capture
  2. with_content_capture_mode(...) when set
  3. content_capture_resolver return value
  4. ClientConfig.content_capture (defaults to NO_TOOL_CONTENT; DEFAULT at the client level resolves to NO_TOOL_CONTENT)

For tool executions, highest to lowest:

  1. ToolExecutionStart.content_capture
  2. Parent generation's resolved mode, or with_content_capture_mode(...) when set
  3. content_capture_resolver return value
  4. ClientConfig.content_capture (defaults to NO_TOOL_CONTENT; DEFAULT at the client level resolves to NO_TOOL_CONTENT)

Exceptions in the resolver are caught and treated as METADATA_ONLY (fail-closed).

Export Configuration

HTTP generation export

import os
from agento11y import ApiConfig, AuthConfig, ClientConfig, GenerationExportConfig

cfg = ClientConfig(
    generation_export=GenerationExportConfig(
        protocol="http",
        endpoint="https://agento11y-prod-<region>.grafana.net",
        auth=AuthConfig(
            mode="basic",
            tenant_id=os.environ["AGENTO11Y_AUTH_TENANT_ID"],
            basic_password=os.environ["AGENTO11Y_AUTH_TOKEN"],
        ),
    ),
    api=ApiConfig(endpoint="https://agento11y-prod-<region>.grafana.net"),
)

generation_export.export_timeout bounds each HTTP or gRPC generation and workflow-step request. It defaults to 30 seconds.

Set AGENTO11Y_EXPORT_TIMEOUT_MS to a base-10 integer from 1 through 2147483647 to override the default. An explicit export_timeout wins over the environment variable. Invalid values produce a warning and keep the default.

Generation export auth modes

Auth is resolved for generation_export.

  • mode="none"
  • mode="tenant" (requires tenant_id, injects X-Scope-OrgID)
  • mode="bearer" (requires bearer_token, injects Authorization: Bearer <token>)
  • mode="basic" (requires basic_password + basic_user or tenant_id, injects Authorization: Basic <base64(user:password)>; also injects X-Scope-OrgID when tenant_id is set — for multi-tenant deployments only, not needed for Grafana Cloud)

Invalid mode/field combinations fail fast in resolve_config(...).

If explicit headers already include Authorization or X-Scope-OrgID, explicit headers win.

import os
from agento11y import ApiConfig, AuthConfig, ClientConfig, GenerationExportConfig

cfg = ClientConfig(
    generation_export=GenerationExportConfig(
        protocol="http",
        endpoint="https://agento11y-prod-<region>.grafana.net",
        auth=AuthConfig(
            mode="basic",
            tenant_id=os.environ["AGENTO11Y_AUTH_TENANT_ID"],
            basic_password=os.environ["AGENTO11Y_AUTH_TOKEN"],
        ),
    ),
    api=ApiConfig(endpoint="https://agento11y-prod-<region>.grafana.net"),
)

Grafana Cloud auth (basic)

For Grafana Cloud, use basic auth mode. The username is your Grafana Cloud instance/tenant ID and the password is your Grafana Cloud API key. See the Grafana Cloud Agent Observability getting started docs for full setup steps; for this SDK endpoint, copy the API URL from Observability → Agent Observability → Configuration. It looks like https://agento11y-prod-<region>.grafana.net.

import os
from agento11y import AuthConfig, ClientConfig, GenerationExportConfig

cfg = ClientConfig(
    generation_export=GenerationExportConfig(
        protocol="http",
        endpoint="https://agento11y-prod-<region>.grafana.net",
        auth=AuthConfig(
            mode="basic",
            tenant_id=os.environ["AGENTO11Y_AUTH_TENANT_ID"],
            basic_password=os.environ["AGENTO11Y_AUTH_TOKEN"],
        ),
    ),
)

If your deployment requires a distinct username, set basic_user explicitly:

auth=AuthConfig(
    mode="basic",
    tenant_id=os.environ["AGENTO11Y_AUTH_TENANT_ID"],
    basic_user=os.environ["AGENTO11Y_AUTH_TENANT_ID"],
    basic_password=os.environ["AGENTO11Y_AUTH_TOKEN"],
)

Wiring custom env vars

The SDK only auto-loads AGENTO11Y_* env vars (AGENTO11Y_ENDPOINT, AGENTO11Y_PROTOCOL, AGENTO11Y_AUTH_MODE, AGENTO11Y_AUTH_TOKEN, etc.) when you call Client(). For any other env var (for example one your secret manager exposes under a different name), read it in your app and pass the value into the config:

import os
from agento11y import AuthConfig, ClientConfig

cfg = ClientConfig()

gen_token = (os.getenv("MY_APP_AGENTO11Y_TOKEN") or "").strip()
if gen_token:
    cfg.generation_export.auth = AuthConfig(mode="bearer", bearer_token=gen_token)

Common topology:

  • Grafana Cloud: generation basic mode with instance ID and API key.
  • Self-hosted direct to the ingest API: generation tenant mode.
  • Traces/metrics via OTEL Collector/Alloy: configure exporters in your app OTEL SDK setup.
  • Enterprise proxy: generation bearer mode to proxy; proxy authenticates and forwards tenant header upstream.

Conversation Ratings

Use the SDK helper to submit user-facing ratings:

from agento11y import ConversationRatingInput, ConversationRatingValue

result = client.submit_conversation_rating(
    "conv-123",
    ConversationRatingInput(
        rating_id="rat-123",
        rating=ConversationRatingValue.BAD,
        comment="Answer ignored user context",
        metadata={"channel": "assistant-ui"},
        source="sdk-python",
    ),
)

print(result.rating.rating, result.summary.has_bad_rating)

submit_conversation_rating(...) sends requests to ClientConfig.api.endpoint, which should be the Grafana Cloud Agent Observability API URL from Agent Observability configuration, and uses the same generation-export auth headers already configured on the SDK client.

Instrumentation-only mode (no generation send)

Set generation_export.protocol="none" to keep generation/tool instrumentation and spans while disabling generation transport.

from agento11y import Client, ClientConfig, GenerationExportConfig

cfg = ClientConfig(
    generation_export=GenerationExportConfig(
        protocol="none",
    ),
)

client = Client(cfg)

Lifecycle and Error Semantics

  • flush() forces immediate export of queued generations.
  • shutdown() flushes pending generations, then closes generation exporters.
  • Always call shutdown() during process teardown to avoid dropped telemetry.
  • recorder.set_call_error(exc) marks provider-call failures on the generation payload and span status.
  • recorder.err() is for local SDK runtime errors only (validation, queue full, payload too large, shutdown).

SDK metrics

The SDK emits these OTel histograms through your configured OTEL meter provider:

  • gen_ai.client.operation.duration
  • gen_ai.client.token.usage
  • gen_ai.client.time_to_first_token
  • gen_ai.client.tool_calls_per_operation

Experiments

Run any agent over a dataset as an Agent Observability experiment (offline evaluation), grade its outputs, and publish scores you can compare in the Agent Observability UI. This is the framework-free path for Cloud users: one ingestion API key writes the run, trials, generations, scores, and final status.

from agento11y import experiments

suite = experiments.TestSuite(
    suite_id="smoke",
    name="Smoke",
    test_cases=[
        experiments.TestCase(test_case_id="capital-fr", input="Capital of France?", expected="Paris"),
    ],
)
verifier = experiments.Evaluator(evaluator_id="exact_match", version="2026-06-29")

with experiments.experiment(
    "PR 123",
    experiment_id="pr-123",
    suite=suite,
    planned_trial_count=len(suite.test_cases),
    tags=["ci"],
) as exp:
    for case in suite.test_cases:
        with exp.trial(case) as trial:
            answer = my_agent(case.input)

            # If your normal instrumentation already produced ids, bind them instead:
            # trial.bind_conversation(conversation_id)
            # trial.bind_generation(generation_id, conversation_id=conversation_id)
            trial.record_io(input=case.input, output=answer, model_provider="openai", model_name="gpt-4o-mini")

            passed = str(case.expected).lower() in answer.lower()
            trial.final_score(1.0 if passed else 0.0, passed=passed, evaluator=verifier)

    print(exp.url)  # deep link to the experiment in Agent Observability

experiment(...) creates the run (source="external"), each exp.trial(...) creates a typed trial, and the trial exports scores on exit. On normal exit the run finalizes as completed; on exception or Ctrl-C it finalizes as failed. Set planned_trial_count to the number of trials the runner intends to execute after filtering and attempt expansion. The SDK does not infer it from suite size because a runner may select cases or execute multiple attempts. Normal finalization lets Agent Observability count stored scores; pass score_count= directly to exp.finalize(...) only when you want the server to assert an exact count. A/B testing is two runs with different experiment_id/tags over the same suite and evaluators.

Experiment writes use the same Grafana Cloud ingestion API key as generation ingest. They do not require a control-plane URL or a separate eval API key. Experiment transcripts, string scores, explanations, metadata, and text artifacts redact recognized secrets by default. Pass redact_secrets=False to experiments.Client(...) only for an explicitly trusted destination. Experimental OTel eval spans/events are disabled by default; opt in with use_experimental_otel=True on experiments.experiment(...) or AGENTO11Y_USE_EXPERIMENTAL_OTEL=true.

Grading with an evaluator stored in Grafana Cloud

Experimental. Set AGENTO11Y_ENABLE_EXPERIMENTAL_FEATURES=true to use this. Without it, trial.evaluate(...), client.trigger_trial_evaluation(...), and client.get_trial_evaluation(...) raise agento11y.ExperimentalFeatureDisabledError without sending a request. Experimental features can change or be removed in any release.

When the grading prompt lives in Agent Observability instead of in the runner, bind the trial to the conversation id your normal instrumentation already produced and let that evaluator score it. trial.evaluate(...) persists the binding, triggers the evaluation, and polls until the worker reaches a terminal status. The evaluator reads every generation in that conversation, not only the last one:

from agento11y import experiments
from agento11y.errors import EvaluationExecutionError, EvaluationTimeoutError

with experiments.experiment("nightly", experiment_id="nightly-42") as exp:
    for case_id, question in CASES:
        with exp.trial(case_id) as trial:
            answer = my_agent(question)          # already instrumented
            trial.bind_conversation(answer.conversation_id)
            try:
                evaluation = trial.evaluate("helpfulness")
            except EvaluationExecutionError as exc:
                print(f"{case_id}: evaluation {exc.evaluation_id} failed: {exc.detail}")
                raise
            except EvaluationTimeoutError as exc:
                print(f"{case_id}: still pending: {exc.detail}")
                raise
            print(case_id, evaluation.status.value, evaluation.attempts)

The evaluator must already exist in Grafana Cloud; the SDK does not create it. A trial graded this way closes as completed with no local final_score.

Three consequences worth knowing before you go looking for the score:

  • The score is attached to the conversation and the trial, not to a generation, so a per-generation score lookup returns nothing. Read it from the experiment's scores or from each trial's scores in exp.report().
  • pass_rate in the report stays unset, because that verdict comes from a score stored under the final key and a stored evaluator writes under its own key.
  • Leave score_count unset when finalizing. The server counts every stored score for the run, cloud ones included, so a locally derived count raises ConflictError with ConflictKind.SCORE_COUNT_MISMATCH.

timeout bounds the polling loop, not the whole call, and a status request already in flight can push the call past it. When the deadline passes the evaluation keeps running server-side, so finalizing straight afterwards raises ConflictError with ConflictKind.PENDING_EVALUATIONS. That conflict is recoverable: poll to a terminal status, then finalize again.

trial.evaluate(...) waits for one evaluation before starting the next. To let them run at the same time, trigger without waiting and poll afterwards. Keep both loops inside the experiment block: leaving it while an evaluation is queued hits the conflict above, and the run stops before it polls anything.

import time

pending = []
with experiments.experiment("nightly", experiment_id="nightly-42") as exp:
    for case_id, question in CASES:
        with exp.trial(case_id) as trial:
            answer = my_agent(question)
            trial.bind_conversation(answer.conversation_id)
            exp.client.update_trial(exp.experiment_id, trial.trial_id, conversation_id=answer.conversation_id)
            queued = exp.client.trigger_trial_evaluation(exp.experiment_id, trial.trial_id, "helpfulness")
            pending.append((case_id, trial.trial_id, queued.evaluation_id))
            trial.succeed()  # no local final score, so mark it before close()

    for case_id, trial_id, evaluation_id in pending:
        evaluation = exp.client.get_trial_evaluation(exp.experiment_id, trial_id, evaluation_id)
        while not evaluation.status.terminal:
            time.sleep(2)
            evaluation = exp.client.get_trial_evaluation(exp.experiment_id, trial_id, evaluation_id)
        print(case_id, evaluation.status.value)

These two endpoints must be routed by the Grafana Cloud gateway for your stack. Where they are not, the request is rejected before it reaches Agent Observability and the SDK surfaces the gateway's answer, 401 invalid scope requested, which looks like a token problem but is not.

See examples/experiments/python/app/run_cloud_evaluator_experiment.py for a runnable version.

Local evaluator helpers

The tracking-first MVP boundary and deferred framework/OTel work are summarized in the Experiments Python SDK MVP design.

Final-output evaluators can be configured entirely in application code. They do not require a stored test suite or an evaluator created in Grafana. LLMJudge accepts any callable that invokes a model; trial.evaluate_output publishes the grader request and response as a linked generation automatically:

judge = experiments.LLMJudge(
    evaluator_id="judge.correctness",
    version="2026-07-21",
    invoke=lambda prompt: anthropic_model.invoke(prompt),
    model_provider="anthropic",
    model_name="claude-sonnet-4-5",
    prompt_template=(
        "Input: {input}\nExpected: {expected}\nOutput: {output}\n"
        'Return JSON: {"score": 0.0, "passed": false, "explanation": "reason"}'
    ),
)

with experiments.experiment("local eval", experiment_id="run-1") as exp:
    with exp.trial("case-1") as trial:
        answer = my_agent("Capital of France?")
        trial.evaluate_output(
            judge,
            input="Capital of France?",
            expected="Paris",
            output=answer,
        )

The score contains judge_provider and judge_model metadata plus the grader conversation and generation IDs. For deterministic checks, use the same path without a model call:

format_check = experiments.RegexJudge(
    evaluator_id="regex.currency",
    pattern=r"^\\$\\d+(?:\\.\\d{2})?$",
)
trial.evaluate_output(format_check, input=prompt, output=answer, score_key="format_valid")

evaluate_output grades only the values supplied by the caller; it does not fetch or normalize the trial's bound conversation. Frameworks and benchmark runners that already own a transcript should produce an EvaluationResult and pass it to trial.record_evaluation(...). The framework remains the authority for transcript rendering while Agent Observability records the result and its trace, conversation, generation, and artifact references.

Evaluator exceptions intentionally propagate after marking the current trial errored. To continue a batch after one judge failure, catch the exception around each with exp.trial(...) block.

See the Experiments v2 migration guide before upgrading a v1 runner.

Stored test suites use the Grafana control plane and therefore require a second, Grafana service-account credential. Set AGENTO11Y_CONTROL_ENDPOINT to either the Grafana stack URL, the Agent Observability app URL, or a full eval API prefix, and set AGENTO11Y_SERVICE_ACCOUNT_TOKEN to the service-account token. Then pull and run an exact published suite version in one call:

with experiments.experiment_from_suite(
    "dashboard-regression",
    version="latest_published",
    experiment_id="pr-123",
    candidate={"git_sha": "abc123", "model_name": "gpt-5"},
) as exp:
    for case in exp.suite.cases:
        with exp.trial(case) as trial:
            answer = my_agent(case.input)
            trial.final_score(answer == case.expected)

Use TestSuitesClient.pull_suite(...) for local editing and push_suite(..., publish=True) to create or update a draft and publish it. push_suite(..., prune=True) also removes remote-only cases from the draft; without prune, pushes are additive.

If you use a supported framework, prefer its adapter (e.g. agento11y-langgraph) — it can expose conversation or generation ids that you bind to the trial, so the experiment points at the same trace your agent already emits. See the agento11y-experiments skill (python/skills/agento11y-experiments/SKILL.md) and the runnable example at examples/experiments/python/ for grading patterns, including LLM-as-judge.

Public API Overview

Core client and lifecycle:

  • Client
  • Client.start_generation(...)
  • Client.start_streaming_generation(...)
  • Client.start_tool_execution(...)
  • Client.flush()
  • Client.shutdown()

Typed payloads:

  • GenerationStart, Generation, ModelRef
  • Message, Part, ToolDefinition, TokenUsage
  • ToolExecutionStart, ToolExecutionEnd
  • ContentCaptureMode

Helpers:

  • user_text_message(...), assistant_text_message(...)
  • with_conversation_id(...), with_agent_name(...), with_agent_version(...)
  • with_content_capture_mode(...)

Validation:

  • validate_generation(...)

Experiments:

  • agento11y.experiments.experiment(...)
  • agento11y.experiments.experiment_from_suite(...)
  • agento11y.experiments.Client
  • agento11y.experiments.TestSuitesClient, PushedSuite
  • agento11y.experiments.Experiment, Trial, TrialRef
  • agento11y.experiments.TestSuite, TestCase, Evaluator
  • agento11y.experiments.stable_id(...)

Provider Helper Packages

Provider wrappers are wrapper-first and mapper-explicit:

  • agento11y-openai
  • agento11y-anthropic
  • agento11y-gemini

Each package exposes sync + async wrappers and explicit mapper functions for custom integration points.

Metadata

Release files for agento11y 0.16.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for agento11y 0.16.0
File Size Uploaded
agento11y-0.16.0.tar.gz 244.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for agento11y 0.16.0
File Interpreter ABI Platform
agento11y-0.16.0-py3-none-any.whl Python 3 none any Details

Total release size: 388.9 kB

Release files / agento11y-0.16.0.tar.gz

Download URL agento11y-0.16.0.tar.gz
Size 244.2 kB
Tags Source
SHA-256 checksum
How to use checksums
6797a290410db8f2cad270095e4f22ee21a2464f5893645764fa8eb5bc49569c
BLAKE2b-256 checksum
How to use checksums
476d4e2d817a94317d94f2025edf737f846334365af9bfb1037ff77eef859823
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 18, 2026.

Transparency log

Release files / agento11y-0.16.0-py3-none-any.whl

Download URL agento11y-0.16.0-py3-none-any.whl
Size 144.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
bfd088d062b7711593752f6f42efa10da3be581f31c990992d20d26e27375592
BLAKE2b-256 checksum
How to use checksums
9127589d3eff18b7e8cb367f04656f2abecb089cab468e613562e6c193b55460
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 18, 2026.

Transparency log

Release history Release notifications | RSS feed

0.18.0

2 release files

0.17.0

2 release files

This release

0.16.0 This release

2 release files

0.15.0

2 release files

0.14.0

2 release files

0.12.0

2 release files

0.10.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page