Skip to main content

Agent Reliability

AI agents can finish successfully while still doing the wrong thing. Agent Reliability provides vendor-neutral Python primitives for measuring whether agents meet explicit reliability objectives.

It brings evaluations, SLOs, error budgets, burn rates, and measurement provenance to agent applications—and refuses to produce a misleading number when evaluation methodologies are incompatible.

Status: GA (1.2.1). Public APIs documented as stable in GA_CONTRACT.md follow Semantic Versioning. See compatibility.

Why this exists

Traces explain what an agent did. Reliability answers whether it consistently achieved a defined outcome. Agent Reliability connects one logical execution to an explicit evaluation method, then calculates exact local reliability against an SLO.

The OSS package works offline and without a hosted service. It does not automatically capture prompts, responses, tool arguments, credentials, or arbitrary application payloads. The base install has no runtime dependencies and sends nothing over the network.

30-second example

python -m pip install agent-reliability

This standalone example instruments four agent-like calls, evaluates task_success, records attributable results, and applies a 75% SLO:

from fractions import Fraction

from agent_reliability.domain import ObjectiveDirection, Slo, UnknownPolicy
from agent_reliability.evaluation import (
    EqualityEvaluator,
    EvaluationResult,
    EvaluatorIdentity,
)
from agent_reliability.reliability import (
    AggregationConflict,
    ReliabilityObservation,
    evaluate_reliability,
)
from agent_reliability.sdk import AgentReliability, EvaluatorRunner

sdk = AgentReliability()
runner = EvaluatorRunner()
evaluator = EqualityEvaluator(EvaluatorIdentity("expected-answer", "1"), "approved")
observations = []

for actual in ("approved", "approved", "needs-review", "approved"):
    with sdk.run(agent_id="approval-agent", name="Approval Agent", version="1") as run:
        result = runner.evaluate(evaluator, actual)
        if not isinstance(result, EvaluationResult):
            raise RuntimeError("evaluation did not produce an observation")
        run.record_evaluation(indicator="task_success", result=result)
        observations.append(
            ReliabilityObservation.from_evaluation(
                indicator="task_success", result=result
            )
        )

report = evaluate_reliability(
    indicator="task_success",
    observations=observations,
    slo=Slo("task-success", Fraction(3, 4), ObjectiveDirection.AT_LEAST),
    unknown_policy=UnknownPolicy.EXCLUDE,
)
if isinstance(report, AggregationConflict):
    raise RuntimeError("incompatible measurement methodologies")

print(f"Reliability: {float(report.ratio.pass_ratio):.2%}")
print(f"SLO status: {report.slo_evaluation.status.value.upper()}")

Output:

Reliability: 75.00%
SLO status: MET

This block runs in CI. The canonical example also shows the error budget. Follow the 5–10 minute quickstart for interpretation and next steps.

What it measures

  • An indicator says what is measured, such as task_success.
  • An evaluator says how it is judged and returns PASS, FAIL, or UNKNOWN.
  • Provenance records evaluator name, behavior version, configuration, and determinism.
  • An SLI is the observed ratio; an SLO is the desired target.
  • The error budget is permitted unreliability; burn rate compares an observed bad-event rate with that allowance.

UNKNOWN means evaluation completed but was indeterminate. An EvaluationExecutionFailure means the evaluator or its timestamping failed; it is not an agent failure and creates no observation.

Reliability answers “how often is the agent behaving correctly?” Measurement health answers “do we have enough trustworthy evidence to make that claim?” They remain independent, and applications—not the SDK—decide how degraded evidence affects an action. See the measurement-health guide. Runnable fail-open, fail-closed, and bounded-degradation examples are in the application policy guide.

If evaluator v1 and v2 measured the same indicator, the engine returns an AggregationConflict instead of averaging them. A changed measurement method is not automatically comparable. See Core concepts.

Installation

Python 3.11–3.13 is supported. The distribution and import names differ:

pip install agent-reliability
import agent_reliability

The only optional runtime extra is the OpenTelemetry API bridge:

python -m pip install "agent-reliability[otel]"

Framework compatibility

Any Python agent can use the explicit sync or async context manager. Wrap one logical task execution, evaluate the relevant output, and retain observations for the window your application chooses. No framework adapter, monkey patch, API key, storage layer, or network service is required. See Integrations and the async example.

The local engine calculates one supplied collection at a time; it does not retain history or select rolling windows.

OpenTelemetry

OpenTelemetryRunContextBridge activates the agent span in an existing host trace. Your application owns the TracerProvider, sampling, processors, propagation, exporter, collector, and backend. Agent Reliability configures none of them and exports nothing by itself. See the OTel example and mapping reference.

Project status and scope

Agent Reliability is a stable, local-first Python SDK for measuring and testing the reliability of AI agents.

The SDK provides sync and async agent instrumentation, deterministic evaluators, evaluator provenance, reliability aggregation, SLO and error-budget semantics, measurement-health tracking, optional OpenTelemetry interoperability, human-readable local reports, machine-readable reliability results, and framework-neutral SLO assertions for tests and CI.

The SDK is designed to work standalone. It requires no hosted account, API key, or network service, and has zero mandatory runtime dependencies.

Current release: 1.2.1

The public API follows semantic versioning, with compatibility and installed-artifact verification included in the release process.

No remote ingestion, dashboard, LLM judge, persistence, auto-instrumentation, or framework-specific adapter is included. See the roadmap.

Documentation

Development and contributing

See CONTRIBUTING.md for setup and quality gates. Security vulnerabilities belong in the private process in SECURITY.md, not a public issue.

License

Apache License 2.0. See LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

agent_reliability-1.2.1.tar.gz (217.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

agent_reliability-1.2.1-py3-none-any.whl (65.6 kB view details)

Uploaded Python 3

File details

Details for the file agent_reliability-1.2.1.tar.gz.

File metadata

  • Download URL: agent_reliability-1.2.1.tar.gz
  • Upload date:
  • Size: 217.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for agent_reliability-1.2.1.tar.gz
Algorithm Hash digest
SHA256 56282aefcf228060f1594a0ee205382111dcab03b7dba4dec763ed704c7b16fc
MD5 205b7db6314bbb6ffb837f8903c82ce3
BLAKE2b-256 2f1f1d13ff0b134d282db50f74d65e1910d292ca57a78ebbed5948fa905e0bdc

See more details on using hashes here.

Provenance

The following attestation bundles were made for agent_reliability-1.2.1.tar.gz:

Publisher: release.yml on kamleshd07/agent-reliability

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file agent_reliability-1.2.1-py3-none-any.whl.

File metadata

File hashes

Hashes for agent_reliability-1.2.1-py3-none-any.whl
Algorithm Hash digest
SHA256 a2dd5687cee143c57481bc3118e53acf0d868714ed940bd68d93d001188ad15f
MD5 d0b99b874f98a0e9946e0147c77ae206
BLAKE2b-256 aaf3d8fdf11b1c5ac425138be9752557de6a37d4c8522a8128044a5986930ba5

See more details on using hashes here.

Provenance

The following attestation bundles were made for agent_reliability-1.2.1-py3-none-any.whl:

Publisher: release.yml on kamleshd07/agent-reliability

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

1.2.1 This release

2 files

1.2.0

2 files

1.1.1

2 files

1.1.0

2 files

1.0.1

2 files

1.0.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page