Grafana Agent Observability Python SDK
agento11y records normalized LLM generation and tool-execution telemetry. It exports normalized generations to Agent Observability ingest and uses your OpenTelemetry tracer/meter setup for traces and metrics.
Use this package when you want:
- A provider-agnostic generation record (same schema for OpenAI, Anthropic, Gemini, or custom adapters).
- OTel-aligned tracing attributes for generation and tool spans.
- Async export with retry/backoff, queueing, batching, and explicit shutdown semantics.
Installation
pip install agento11y
For a Grafana Cloud setup walkthrough (where to find the endpoint URL, instance ID, and API token), refer to the Grafana Cloud setup guide.
Validation
Run the shared core conformance suite for the Python SDK from the repo root:
mise run test:py:sdk-conformance
Run the cross-language aggregate core conformance suite from the repo root:
mise run sdk:conformance
Optional provider helper packages:
pip install agento11y-openai
pip install agento11y-anthropic
pip install agento11y-gemini
Optional framework modules:
pip install agento11y-langchain
pip install agento11y-langgraph
pip install agento11y-openai-agents
pip install agento11y-llamaindex
pip install agento11y-google-adk
pip install agento11y-strands
pip install agento11y-claude-agent-sdk
pip install agento11y-litellm
pip install agento11y-pydantic-ai
Framework handler usage:
from agento11y import Client
from agento11y_langchain import with_agento11y_langchain_callbacks
from agento11y_langgraph import with_agento11y_langgraph_callbacks
from agento11y_openai_agents import with_agento11y_openai_agents_hooks
from agento11y_llamaindex import with_agento11y_llamaindex_callbacks
from agento11y_google_adk import with_agento11y_google_adk_callbacks
from agento11y_strands import with_agento11y_strands_hooks
from agento11y_claude_agent import with_agento11y_claude_agent_options
from agento11y_pydantic_ai import with_agento11y_pydantic_ai_capability
client = Client()
chain_config = with_agento11y_langchain_callbacks(None, client=client, provider_resolver="auto")
graph_config = with_agento11y_langgraph_callbacks(None, client=client, provider_resolver="auto")
openai_agents_run_options = with_agento11y_openai_agents_hooks(None, client=client, provider_resolver="auto")
llamaindex_config = with_agento11y_llamaindex_callbacks(None, client=client, provider_resolver="auto")
google_adk_agent_config = with_agento11y_google_adk_callbacks(None, client=client, provider_resolver="auto")
strands_agent_config = with_agento11y_strands_hooks(None, client=client, provider_resolver="auto")
claude_agent_options = with_agento11y_claude_agent_options(None, client=client)
pydantic_ai_capabilities = with_agento11y_pydantic_ai_capability(None, client=client, provider_resolver="auto")
LiteLLM uses a callback class instead of a with_agento11y_* helper:
import litellm
from agento11y import Client
from agento11y_litellm import Agento11yLiteLLMLogger
client = Client()
litellm.callbacks = [Agento11yLiteLLMLogger(client=client)]
Framework handlers use the Client instance you pass in. If that client is configured with
generation_sanitizer, the same redaction policy applies automatically to generations recorded
through LangChain, LangGraph, OpenAI Agents, LlamaIndex, Google ADK, Strands, Claude Agent SDK, LiteLLM, and Pydantic AI integrations.
Framework handlers inject framework tags/metadata on recorded generations:
agento11y.framework.name(langchain,langgraph,openai-agents,llamaindex,google-adk,strands,claude-agent-sdk,litellm, orpydantic-ai)agento11y.framework.source=handler(orhooksfor Strands Agents and Claude Agent SDK)agento11y.framework.language=pythonmetadata["agento11y.framework.run_id"]metadata["agento11y.framework.thread_id"](when present)metadata["agento11y.framework.parent_run_id"](when available)metadata["agento11y.framework.component_name"]metadata["agento11y.framework.run_type"]metadata["agento11y.framework.tags"]metadata["agento11y.framework.retry_attempt"](when available)metadata["agento11y.framework.event_id"](when available)metadata["agento11y.framework.langgraph.node"](LangGraph when available)
Conversation mapping is conversation-first:
conversation_id/session_id/group_idfrom framework context first- then
thread_id - deterministic fallback
agento11y:framework:<framework_name>:<run_id>
When present in generation metadata, low-cardinality framework keys are copied onto generation span attributes.
For LangGraph persistence, pass configurable.thread_id and reuse it across invocations:
thread_config = {
**with_agento11y_langgraph_callbacks(None, client=client, provider_resolver="auto"),
"configurable": {"thread_id": "customer-42"},
}
graph.invoke({"prompt": "Remember my timezone is UTC+1.", "answer": ""}, config=thread_config)
graph.invoke({"prompt": "What timezone did I give you?", "answer": ""}, config=thread_config)
Full framework examples:
- LangChain:
../python-frameworks/langchain/README.md - LangGraph:
../python-frameworks/langgraph/README.md - OpenAI Agents:
../python-frameworks/openai-agents/README.md - LlamaIndex:
../python-frameworks/llamaindex/README.md - Google ADK:
../python-frameworks/google-adk/README.md - Strands Agents:
../python-frameworks/strands/README.md - Claude Agent SDK:
../python-frameworks/claude-agent-sdk/README.md - LiteLLM:
../python-frameworks/litellm/README.md - Pydantic AI:
../python-frameworks/pydantic-ai/README.md
Quick Start (Sync Generation)
Client() reads AGENTO11Y_* env vars by default. See the Grafana Cloud setup guide for the variable names. Pass an explicit ClientConfig only when you need to override.
from agento11y import (
Client,
GenerationStart,
ModelRef,
assistant_text_message,
user_text_message,
)
client = Client() # reads AGENTO11Y_* env vars
with client.start_generation(
GenerationStart(
conversation_id="conv-1",
agent_name="my-service",
agent_version="1.0.0",
model=ModelRef(provider="openai", name="gpt-5"),
)
) as rec:
rec.set_result(
input=[user_text_message("What is the weather in Paris?")],
output=[assistant_text_message("It is 18C and sunny.")],
)
# Recorder errors are local SDK errors (validation/enqueue/shutdown),
# not provider call failures.
if rec.err() is not None:
raise rec.err()
client.shutdown()
Explicit configuration form:
import os
from agento11y import AuthConfig, Client, ClientConfig, GenerationExportConfig
client = Client(
ClientConfig(
generation_export=GenerationExportConfig(
protocol="http",
endpoint="https://agento11y-prod-<region>.grafana.net",
auth=AuthConfig(
mode="basic",
tenant_id=os.environ["AGENTO11Y_AUTH_TENANT_ID"],
basic_password=os.environ["AGENTO11Y_AUTH_TOKEN"],
),
),
)
)
Pre-Ingest Redaction
Use generation_sanitizer when you want to redact substrings from normalized generations before
validation, span sync, and export.
from agento11y import (
Client,
ClientConfig,
SecretRedactionOptions,
create_secret_redaction_sanitizer,
)
client = Client(
ClientConfig(
generation_sanitizer=create_secret_redaction_sanitizer(
SecretRedactionOptions(
redact_input_messages=False, # None falls back to AGENTO11Y_REDACT_INPUT_MESSAGES, then False
redact_email_addresses=True,
)
)
)
)
The built-in sanitizer:
- redacts high-confidence secret formats in assistant text and thinking
- redacts secret formats plus key/value secrets in system prompts, tool call inputs, and tool results
- redacts email addresses by default
- redacts the conversation title and call error
- redacts historic assistant turns and tool messages in input
- leaves user input unchanged unless input redaction is enabled
To preserve email addresses, opt out explicitly:
client = Client(
ClientConfig(
generation_sanitizer=create_secret_redaction_sanitizer(
SecretRedactionOptions(redact_email_addresses=False)
)
)
)
Configuring redaction via environment variables
create_secret_redaction_sanitizer() reads AGENTO11Y_REDACT_INPUT_MESSAGES (accepts 1/0, true/false, yes/no, on/off) when redact_input_messages is left None. Precedence is explicit option > env var > False. An unrecognised env value logs a warning through the agento11y logger and falls back to the next layer, so a typo cannot silently flip redaction.
from agento11y import (
Client,
ClientConfig,
create_secret_redaction_sanitizer,
)
# Leave redact_input_messages unset so AGENTO11Y_REDACT_INPUT_MESSAGES decides.
client = Client(
ClientConfig(
generation_sanitizer=create_secret_redaction_sanitizer(),
)
)
Hooks and Guards
Use hooks when you want Agent Observability guard rules to run before an LLM call. The SDK evaluates the hook on your request path; guard rules configured in Grafana Cloud decide whether to allow, deny, or transform the input.
Hooks are disabled by default. Enable them on the client and call evaluate_hook(...) before the provider request:
from agento11y import (
Client,
ClientConfig,
HookContext,
HookEvaluateRequest,
HookInput,
HookModel,
HookPhase,
HooksConfig,
Message,
MessageRole,
hook_denied_from_response,
text_part,
)
client = Client(ClientConfig(hooks=HooksConfig(enabled=True)))
messages = [
Message(role=MessageRole.USER, parts=[text_part("Summarize this customer note...")]),
]
response = client.evaluate_hook(
HookEvaluateRequest(
phase=HookPhase.PREFLIGHT.value,
context=HookContext(
agent_name="support-agent",
agent_version="1.0.0",
model=HookModel(provider="openai", name="gpt-5"),
conversation_id="support-case-42",
),
input=HookInput(
messages=messages,
system_prompt="You are a helpful support agent.",
conversation_preview="Summarize this customer note...",
),
)
)
denied = hook_denied_from_response(response)
if denied is not None:
raise denied
if response.transformed_input is not None:
messages = response.transformed_input.messages or messages
HooksConfig defaults to phases=["preflight"], timeout_seconds=15.0, and fail_open=True. With fail-open enabled, hook transport errors resolve to allow so an unavailable evaluator does not block production traffic. Set fail_open=False for strict paths that should fail closed.
evaluate_hook also takes a keyword-only hooks override that replaces the client's resolved hook configuration for that one call. It does not mutate the client, so later calls use the client configuration again:
from dataclasses import replace
response = client.evaluate_hook(request, hooks=replace(client.hooks_config, fail_open=False))
Use it when your caller has to tell a server allow from an allow the SDK synthesized after a failure. The response carries no marker separating the two, so a fail-open allow otherwise looks like a completed evaluation. agento11y-litellm calls fail-closed this way and applies the configured fail_open policy itself, after recording the failure as a guardrail verdict.
Set HookContext.conversation_id to the same ID used by start_generation(...). The SDK also reads with_conversation_id(...) and the active OpenTelemetry span when explicit correlation fields are omitted. This lets Agent Observability retain denied preflight attempts even though no LLM generation is created.
If you use transformed input, pass the transformed messages/system prompt to the provider and record those same values in start_generation(...). For a runnable example, see ../examples/getting-started/python-hooks/.
Configure OTEL exporters (traces/metrics) in your application OTEL SDK setup. You can optionally pass tracer and meter via ClientConfig.
Quick OTEL setup pattern before creating the agento11y client:
from opentelemetry import metrics, trace
from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.trace import TracerProvider
trace.set_tracer_provider(TracerProvider())
metrics.set_meter_provider(MeterProvider())
The providers above have no exporters attached, so nothing leaves the process. See OpenTelemetry Setup for the full wiring, including the OTLP exporters, and why analytics stays empty when this step is skipped.
Streaming Generation
Use start_streaming_generation(...) when the upstream provider call is streaming.
from agento11y import GenerationStart, ModelRef
with client.start_streaming_generation(
GenerationStart(
conversation_id="conv-stream",
model=ModelRef(provider="anthropic", name="claude-sonnet-4-5"),
)
) as rec:
rec.set_result(output=[assistant_text_message("partial stream summary")])
Embedding Observability
Use start_embedding(...) for embedding API calls. Embedding recording emits OTel spans and SDK metrics only, and does not enqueue generation exports.
from agento11y import EmbeddingResult, EmbeddingStart, ModelRef
with client.start_embedding(
EmbeddingStart(
agent_name="retrieval-worker",
agent_version="1.0.0",
model=ModelRef(provider="openai", name="text-embedding-3-small"),
)
) as rec:
response = openai.embeddings.create(model="text-embedding-3-small", input=["hello", "world"])
rec.set_result(
EmbeddingResult(
input_count=2,
input_tokens=response.usage.prompt_tokens,
input_texts=["hello", "world"], # captured only when embedding_capture.capture_input=True
response_model=response.model,
)
)
Input text capture is opt-in:
from agento11y import ClientConfig, EmbeddingCaptureConfig
cfg = ClientConfig(
embedding_capture=EmbeddingCaptureConfig(
capture_input=True,
max_input_items=20,
max_text_length=1024,
)
)
capture_input may expose PII/document content in spans. Keep it disabled by default and enable only for scoped debugging.
TraceQL examples:
traces{gen_ai.operation.name="embeddings"}traces{gen_ai.operation.name="embeddings" && gen_ai.request.model="text-embedding-3-small"}traces{gen_ai.operation.name="embeddings" && error.type!=""}
Tool Execution Span Recording
Tool spans are recorded independently of generation export.
from agento11y import ToolExecutionStart
with client.start_tool_execution(
ToolExecutionStart(
tool_name="weather",
tool_call_id="call_weather_1",
tool_type="function",
include_content=True,
)
) as rec:
rec.set_result(arguments={"city": "Paris"}, result={"temp_c": 18})
SDK identity attributes
- Generation and tool spans always include:
agento11y.sdk.name=sdk-python
- Normalized generation metadata always includes the same key.
- If caller metadata provides a conflicting value for this key, the SDK overwrites it.
Context Defaults
Use context helpers to set defaults once per request/task boundary.
from agento11y import with_agent_name, with_agent_version, with_conversation_id
with with_conversation_id("conv-ctx"), with_agent_name("planner"), with_agent_version("2026.02"):
with client.start_generation(
GenerationStart(model=ModelRef(provider="gemini", name="gemini-2.5-pro"))
) as rec:
rec.set_result(output=[assistant_text_message("ok")])
Content Capture Mode
ContentCaptureMode controls what content is included in exported generation payloads and OTel span attributes. Use it to prevent sensitive text (prompts, tool I/O, model responses) from leaving the process. See Content Capture Modes for the cross-SDK reference, including the per-surface behavior matrix.
| Mode | Generation export | Generation span | Tool spans | Embedding span |
|---|---|---|---|---|
FULL |
Full content | Content attributes included | Arguments and results included | Input texts included when capture is on |
NO_TOOL_CONTENT (SDK default) |
Full content | Content attributes included | Arguments and results excluded | Input texts included when capture is on |
METADATA_ONLY |
Structure only; text and tool I/O stripped | Content attributes omitted | Arguments and results excluded | Input texts omitted |
FULL_WITH_METADATA_SPANS |
Full content | Content attributes omitted | Arguments and results excluded | Input texts omitted |
DEFAULT is a placeholder for "inherit from the next layer"; at the client level it resolves to NO_TOOL_CONTENT. The SDK default is NO_TOOL_CONTENT, which matches the SDK's behavior before this feature was added.
FULL_WITH_METADATA_SPANS is the right mode when the gRPC ingest destination is private but the OTel trace/metric destination is shared and must not receive any content. Tool execution and embedding spans behave like METADATA_ONLY under this mode because they have no separate gRPC export.
User-provided metadata and tags are not stripped by any capture mode; callers must avoid putting sensitive content in those dicts when using METADATA_ONLY or FULL_WITH_METADATA_SPANS. SDK-internal metadata keys that carry content (e.g. call_error, agento11y.conversation.title) are stripped along with the matching content. See Tags and Metadata for where client tags, per-generation tags, metadata, and user_id each show up (export vs spans vs metrics).
Client-level default
from agento11y import Client, ClientConfig, ContentCaptureMode
client = Client(ClientConfig(
content_capture=ContentCaptureMode.METADATA_ONLY,
))
Per-generation override
from agento11y import ContentCaptureMode, GenerationStart, ModelRef
with client.start_generation(
GenerationStart(
model=ModelRef(provider="openai", name="gpt-5"),
content_capture=ContentCaptureMode.FULL,
)
) as rec:
rec.set_result(
input=[user_text_message("What is the weather?")],
output=[assistant_text_message("18C and sunny.")],
)
Context propagation
Child tool executions inherit the active capture mode from the parent generation via ContextVar. You can also set it explicitly for a block:
from agento11y import ContentCaptureMode, with_content_capture_mode
with with_content_capture_mode(ContentCaptureMode.METADATA_ONLY):
with client.start_tool_execution(
ToolExecutionStart(tool_name="search")
) as rec:
rec.set_result(arguments={"q": "weather"}, result={"temp_c": 18})
Dynamic resolution via resolver
A callback on ClientConfig that resolves the capture mode per-recording at runtime. Useful for feature flags, per-tenant policies, or context-dependent decisions:
from agento11y import Client, ClientConfig, ContentCaptureMode
def resolve_capture(metadata: dict) -> ContentCaptureMode:
if metadata.get("tenant") == "healthcare":
return ContentCaptureMode.METADATA_ONLY
return ContentCaptureMode.DEFAULT # fall through to client default
client = Client(ClientConfig(
content_capture_resolver=resolve_capture,
))
Resolution precedence
For generations, highest to lowest:
GenerationStart.content_capturewith_content_capture_mode(...)when setcontent_capture_resolverreturn valueClientConfig.content_capture(defaults toNO_TOOL_CONTENT;DEFAULTat the client level resolves toNO_TOOL_CONTENT)
For tool executions, highest to lowest:
ToolExecutionStart.content_capture- Parent generation's resolved mode, or
with_content_capture_mode(...)when set content_capture_resolverreturn valueClientConfig.content_capture(defaults toNO_TOOL_CONTENT;DEFAULTat the client level resolves toNO_TOOL_CONTENT)
Exceptions in the resolver are caught and treated as METADATA_ONLY (fail-closed).
Export Configuration
HTTP generation export
import os
from agento11y import ApiConfig, AuthConfig, ClientConfig, GenerationExportConfig
cfg = ClientConfig(
generation_export=GenerationExportConfig(
protocol="http",
endpoint="https://agento11y-prod-<region>.grafana.net",
auth=AuthConfig(
mode="basic",
tenant_id=os.environ["AGENTO11Y_AUTH_TENANT_ID"],
basic_password=os.environ["AGENTO11Y_AUTH_TOKEN"],
),
),
api=ApiConfig(endpoint="https://agento11y-prod-<region>.grafana.net"),
)
generation_export.export_timeout bounds each HTTP or gRPC generation and workflow-step request. It defaults to 30 seconds.
Set AGENTO11Y_EXPORT_TIMEOUT_MS to a base-10 integer from 1 through 2147483647 to override the default. An explicit export_timeout wins over the environment variable. Invalid values produce a warning and keep the default.
Generation export auth modes
Auth is resolved for generation_export.
mode="none"mode="tenant"(requirestenant_id, injectsX-Scope-OrgID)mode="bearer"(requiresbearer_token, injectsAuthorization: Bearer <token>)mode="basic"(requiresbasic_password+basic_userortenant_id, injectsAuthorization: Basic <base64(user:password)>; also injectsX-Scope-OrgIDwhentenant_idis set — for multi-tenant deployments only, not needed for Grafana Cloud)
Invalid mode/field combinations fail fast in resolve_config(...).
If explicit headers already include Authorization or X-Scope-OrgID, explicit headers win.
import os
from agento11y import ApiConfig, AuthConfig, ClientConfig, GenerationExportConfig
cfg = ClientConfig(
generation_export=GenerationExportConfig(
protocol="http",
endpoint="https://agento11y-prod-<region>.grafana.net",
auth=AuthConfig(
mode="basic",
tenant_id=os.environ["AGENTO11Y_AUTH_TENANT_ID"],
basic_password=os.environ["AGENTO11Y_AUTH_TOKEN"],
),
),
api=ApiConfig(endpoint="https://agento11y-prod-<region>.grafana.net"),
)
Grafana Cloud auth (basic)
For Grafana Cloud, use basic auth mode. The username is your Grafana Cloud instance/tenant ID and the password is your Grafana Cloud API key. See the Grafana Cloud Agent Observability getting started docs for full setup steps; for this SDK endpoint, copy the API URL from Observability → Agent Observability → Configuration. It looks like https://agento11y-prod-<region>.grafana.net.
import os
from agento11y import AuthConfig, ClientConfig, GenerationExportConfig
cfg = ClientConfig(
generation_export=GenerationExportConfig(
protocol="http",
endpoint="https://agento11y-prod-<region>.grafana.net",
auth=AuthConfig(
mode="basic",
tenant_id=os.environ["AGENTO11Y_AUTH_TENANT_ID"],
basic_password=os.environ["AGENTO11Y_AUTH_TOKEN"],
),
),
)
If your deployment requires a distinct username, set basic_user explicitly:
auth=AuthConfig(
mode="basic",
tenant_id=os.environ["AGENTO11Y_AUTH_TENANT_ID"],
basic_user=os.environ["AGENTO11Y_AUTH_TENANT_ID"],
basic_password=os.environ["AGENTO11Y_AUTH_TOKEN"],
)
Wiring custom env vars
The SDK only auto-loads AGENTO11Y_* env vars (AGENTO11Y_ENDPOINT, AGENTO11Y_PROTOCOL, AGENTO11Y_AUTH_MODE, AGENTO11Y_AUTH_TOKEN, etc.) when you call Client(). For any other env var (for example one your secret manager exposes under a different name), read it in your app and pass the value into the config:
import os
from agento11y import AuthConfig, ClientConfig
cfg = ClientConfig()
gen_token = (os.getenv("MY_APP_AGENTO11Y_TOKEN") or "").strip()
if gen_token:
cfg.generation_export.auth = AuthConfig(mode="bearer", bearer_token=gen_token)
Common topology:
- Grafana Cloud: generation
basicmode with instance ID and API key. - Self-hosted direct to the ingest API: generation
tenantmode. - Traces/metrics via OTEL Collector/Alloy: configure exporters in your app OTEL SDK setup.
- Enterprise proxy: generation
bearermode to proxy; proxy authenticates and forwards tenant header upstream.
Conversation Ratings
Use the SDK helper to submit user-facing ratings:
from agento11y import ConversationRatingInput, ConversationRatingValue
result = client.submit_conversation_rating(
"conv-123",
ConversationRatingInput(
rating_id="rat-123",
rating=ConversationRatingValue.BAD,
comment="Answer ignored user context",
metadata={"channel": "assistant-ui"},
source="sdk-python",
),
)
print(result.rating.rating, result.summary.has_bad_rating)
submit_conversation_rating(...) sends requests to ClientConfig.api.endpoint, which should be the Grafana Cloud Agent Observability API URL from Agent Observability configuration, and uses the same generation-export auth headers already configured on the SDK client.
Instrumentation-only mode (no generation send)
Set generation_export.protocol="none" to keep generation/tool instrumentation and spans while disabling generation transport.
from agento11y import Client, ClientConfig, GenerationExportConfig
cfg = ClientConfig(
generation_export=GenerationExportConfig(
protocol="none",
),
)
client = Client(cfg)
Lifecycle and Error Semantics
flush()forces immediate export of queued generations.shutdown()flushes pending generations, then closes generation exporters.- Always call
shutdown()during process teardown to avoid dropped telemetry. recorder.set_call_error(exc)marks provider-call failures on the generation payload and span status.recorder.err()is for local SDK runtime errors only (validation, queue full, payload too large, shutdown).
SDK metrics
The SDK emits these OTel histograms through your configured OTEL meter provider:
gen_ai.client.operation.durationgen_ai.client.token.usagegen_ai.client.time_to_first_tokengen_ai.client.tool_calls_per_operation
Experiments
Run any agent over a dataset as an Agent Observability experiment (offline evaluation), grade its outputs, and publish scores you can compare in the Agent Observability UI. This is the framework-free path for Cloud users: one ingestion API key writes the run, trials, generations, scores, and final status.
from agento11y import experiments
suite = experiments.TestSuite(
suite_id="smoke",
name="Smoke",
test_cases=[
experiments.TestCase(test_case_id="capital-fr", input="Capital of France?", expected="Paris"),
],
)
verifier = experiments.Evaluator(evaluator_id="exact_match", version="2026-06-29")
with experiments.experiment(
"PR 123",
experiment_id="pr-123",
suite=suite,
planned_trial_count=len(suite.test_cases),
tags=["ci"],
) as exp:
for case in suite.test_cases:
with exp.trial(case) as trial:
answer = my_agent(case.input)
# If your normal instrumentation already produced ids, bind them instead:
# trial.bind_conversation(conversation_id)
# trial.bind_generation(generation_id, conversation_id=conversation_id)
trial.record_io(input=case.input, output=answer, model_provider="openai", model_name="gpt-4o-mini")
passed = str(case.expected).lower() in answer.lower()
trial.final_score(1.0 if passed else 0.0, passed=passed, evaluator=verifier)
print(exp.url) # deep link to the experiment in Agent Observability
experiment(...) creates the run (source="external"), each exp.trial(...)
creates a typed trial, and the trial exports scores on exit. On normal exit the
run finalizes as completed; on exception or Ctrl-C it finalizes as failed.
Set planned_trial_count to the number of trials the runner intends to execute
after filtering and attempt expansion. The SDK does not infer it from suite
size because a runner may select cases or execute multiple attempts. Normal
finalization lets Agent Observability count stored scores; pass score_count= directly
to exp.finalize(...) only when you want the server to assert an exact count.
A/B testing is two runs with different experiment_id/tags over the same
suite and evaluators.
Experiment writes use the same Grafana Cloud ingestion API key as generation
ingest. They do not require a control-plane URL or a separate eval API key.
Experiment transcripts, string scores, explanations, metadata, and text
artifacts redact recognized secrets by default. Pass redact_secrets=False to
experiments.Client(...) only for an explicitly trusted destination.
Experimental OTel eval spans/events are disabled by default; opt in with
use_experimental_otel=True on experiments.experiment(...) or
AGENTO11Y_USE_EXPERIMENTAL_OTEL=true.
Grading with an evaluator stored in Grafana Cloud
Experimental. Set
AGENTO11Y_ENABLE_EXPERIMENTAL_FEATURES=trueto use this. Without it,trial.evaluate(...),client.trigger_trial_evaluation(...), andclient.get_trial_evaluation(...)raiseagento11y.ExperimentalFeatureDisabledErrorwithout sending a request. Experimental features can change or be removed in any release.
When the grading prompt lives in Agent Observability instead of in the runner,
bind the trial to the conversation id your normal instrumentation already
produced and let that evaluator score it. trial.evaluate(...) persists the
binding, triggers the evaluation, and polls until the worker reaches a terminal
status. The evaluator reads every generation in that conversation, not only the
last one:
from agento11y import experiments
from agento11y.errors import EvaluationExecutionError, EvaluationTimeoutError
with experiments.experiment("nightly", experiment_id="nightly-42") as exp:
for case_id, question in CASES:
with exp.trial(case_id) as trial:
answer = my_agent(question) # already instrumented
trial.bind_conversation(answer.conversation_id)
try:
evaluation = trial.evaluate("helpfulness")
except EvaluationExecutionError as exc:
print(f"{case_id}: evaluation {exc.evaluation_id} failed: {exc.detail}")
raise
except EvaluationTimeoutError as exc:
print(f"{case_id}: still pending: {exc.detail}")
raise
print(case_id, evaluation.status.value, evaluation.attempts)
The evaluator must already exist in Grafana Cloud; the SDK does not create it. A
trial graded this way closes as completed with no local final_score.
Three consequences worth knowing before you go looking for the score:
- The score is attached to the conversation and the trial, not to a generation,
so a per-generation score lookup returns nothing. Read it from the
experiment's scores or from each trial's
scoresinexp.report(). pass_ratein the report stays unset, because that verdict comes from a score stored under thefinalkey and a stored evaluator writes under its own key.- Leave
score_countunset when finalizing. The server counts every stored score for the run, cloud ones included, so a locally derived count raisesConflictErrorwithConflictKind.SCORE_COUNT_MISMATCH.
timeout bounds the polling loop, not the whole call, and a status request
already in flight can push the call past it. When the deadline passes the
evaluation keeps running server-side, so finalizing straight afterwards raises
ConflictError with ConflictKind.PENDING_EVALUATIONS. That conflict is
recoverable: poll to a terminal status, then finalize again.
trial.evaluate(...) waits for one evaluation before starting the next. To let
them run at the same time, trigger without waiting and poll afterwards. Keep both
loops inside the experiment block: leaving it while an evaluation is queued hits
the conflict above, and the run stops before it polls anything.
import time
pending = []
with experiments.experiment("nightly", experiment_id="nightly-42") as exp:
for case_id, question in CASES:
with exp.trial(case_id) as trial:
answer = my_agent(question)
trial.bind_conversation(answer.conversation_id)
exp.client.update_trial(exp.experiment_id, trial.trial_id, conversation_id=answer.conversation_id)
queued = exp.client.trigger_trial_evaluation(exp.experiment_id, trial.trial_id, "helpfulness")
pending.append((case_id, trial.trial_id, queued.evaluation_id))
trial.succeed() # no local final score, so mark it before close()
for case_id, trial_id, evaluation_id in pending:
evaluation = exp.client.get_trial_evaluation(exp.experiment_id, trial_id, evaluation_id)
while not evaluation.status.terminal:
time.sleep(2)
evaluation = exp.client.get_trial_evaluation(exp.experiment_id, trial_id, evaluation_id)
print(case_id, evaluation.status.value)
These two endpoints must be routed by the Grafana Cloud gateway for your stack.
Where they are not, the request is rejected before it reaches Agent
Observability and the SDK surfaces the gateway's answer,
401 invalid scope requested, which looks like a token problem but is not.
See examples/experiments/python/app/run_cloud_evaluator_experiment.py
for a runnable version.
Local evaluator helpers
The tracking-first MVP boundary and deferred framework/OTel work are summarized in the Experiments Python SDK MVP design.
Final-output evaluators can be configured entirely in application code. They do
not require a stored test suite or an evaluator created in Grafana. LLMJudge
accepts any callable that invokes a model; trial.evaluate_output publishes
the grader request and response as a linked generation automatically:
judge = experiments.LLMJudge(
evaluator_id="judge.correctness",
version="2026-07-21",
invoke=lambda prompt: anthropic_model.invoke(prompt),
model_provider="anthropic",
model_name="claude-sonnet-4-5",
prompt_template=(
"Input: {input}\nExpected: {expected}\nOutput: {output}\n"
'Return JSON: {"score": 0.0, "passed": false, "explanation": "reason"}'
),
)
with experiments.experiment("local eval", experiment_id="run-1") as exp:
with exp.trial("case-1") as trial:
answer = my_agent("Capital of France?")
trial.evaluate_output(
judge,
input="Capital of France?",
expected="Paris",
output=answer,
)
The score contains judge_provider and judge_model metadata plus the
grader conversation and generation IDs. For deterministic checks, use the same
path without a model call:
format_check = experiments.RegexJudge(
evaluator_id="regex.currency",
pattern=r"^\\$\\d+(?:\\.\\d{2})?$",
)
trial.evaluate_output(format_check, input=prompt, output=answer, score_key="format_valid")
evaluate_output grades only the values supplied by the caller; it does not
fetch or normalize the trial's bound conversation. Frameworks and benchmark
runners that already own a transcript should produce an EvaluationResult
and pass it to trial.record_evaluation(...). The framework remains the
authority for transcript rendering while Agent Observability records the result and its
trace, conversation, generation, and artifact references.
Evaluator exceptions intentionally propagate after marking the current trial
errored. To continue a batch after one judge failure, catch the exception around
each with exp.trial(...) block.
See the Experiments v2 migration guide before upgrading a v1 runner.
Stored test suites use the Grafana control plane and therefore require a second,
Grafana service-account credential. Set AGENTO11Y_CONTROL_ENDPOINT to either
the Grafana stack URL, the Agent Observability app URL, or a full eval API prefix,
and set AGENTO11Y_SERVICE_ACCOUNT_TOKEN to the service-account token. Then pull
and run an exact published suite version in one call:
with experiments.experiment_from_suite(
"dashboard-regression",
version="latest_published",
experiment_id="pr-123",
candidate={"git_sha": "abc123", "model_name": "gpt-5"},
) as exp:
for case in exp.suite.cases:
with exp.trial(case) as trial:
answer = my_agent(case.input)
trial.final_score(answer == case.expected)
Use TestSuitesClient.pull_suite(...) for local editing and
push_suite(..., publish=True) to create or update a draft and publish it.
push_suite(..., prune=True) also removes remote-only cases from the draft;
without prune, pushes are additive.
If you use a supported framework, prefer its adapter (e.g. agento11y-langgraph)
— it can expose conversation or generation ids that you bind to the trial, so
the experiment points at the same trace your agent already emits. See the
agento11y-experiments skill
(python/skills/agento11y-experiments/SKILL.md) and the runnable example at
examples/experiments/python/ for grading patterns, including LLM-as-judge.
Public API Overview
Core client and lifecycle:
ClientClient.start_generation(...)Client.start_streaming_generation(...)Client.start_tool_execution(...)Client.flush()Client.shutdown()
Typed payloads:
GenerationStart,Generation,ModelRefMessage,Part,ToolDefinition,TokenUsageToolExecutionStart,ToolExecutionEndContentCaptureMode
Helpers:
user_text_message(...),assistant_text_message(...)with_conversation_id(...),with_agent_name(...),with_agent_version(...)with_content_capture_mode(...)
Validation:
validate_generation(...)
Experiments:
agento11y.experiments.experiment(...)agento11y.experiments.experiment_from_suite(...)agento11y.experiments.Clientagento11y.experiments.TestSuitesClient,PushedSuiteagento11y.experiments.Experiment,Trial,TrialRefagento11y.experiments.TestSuite,TestCase,Evaluatoragento11y.experiments.stable_id(...)
Provider Helper Packages
Provider wrappers are wrapper-first and mapper-explicit:
agento11y-openaiagento11y-anthropicagento11y-gemini
Each package exposes sync + async wrappers and explicit mapper functions for custom integration points.
Metadata
Release files for agento11y 0.16.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| agento11y-0.16.0.tar.gz | 244.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| agento11y-0.16.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 388.9 kB
Release files / agento11y-0.16.0.tar.gz
| Download URL | agento11y-0.16.0.tar.gz |
|---|---|
| Size | 244.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
6797a290410db8f2cad270095e4f22ee21a2464f5893645764fa8eb5bc49569c
|
|
BLAKE2b-256 checksum How to use checksums |
476d4e2d817a94317d94f2025edf737f846334365af9bfb1037ff77eef859823
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.13
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 18, 2026.
Transparency logRelease files / agento11y-0.16.0-py3-none-any.whl
| Download URL | agento11y-0.16.0-py3-none-any.whl |
|---|---|
| Size | 144.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
bfd088d062b7711593752f6f42efa10da3be581f31c990992d20d26e27375592
|
|
BLAKE2b-256 checksum How to use checksums |
9127589d3eff18b7e8cb367f04656f2abecb089cab468e613562e6c193b55460
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.13
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 18, 2026.
Transparency log