Measure the real-world error rate and dollar cost of an AI agent's decisions. OpenTelemetry-native.

These details have been verified by PyPI

Project links

Repository

GitHub Statistics

Maintainers

sanjayshukla

These details have not been verified by PyPI

Project links

Project description

agentloss

Your eval tool tells you your AI agent's hallucination rate. agentloss tells you what it costs. An OpenTelemetry-native SDK that measures the real-world error rate and dollar loss of an AI agent's decisions — by capturing its consequential actions in-process and joining them to ground truth (real resolved outcomes, not an offline labeled set).

Every eval/observability tool scores quality proxies — LLM-judge, hallucination rate, task completion. agentloss answers the question the market keeps asking and no tool measures: what are my agent's mistakes costing, and is it safe to trust with more autonomy?

Part of ADMT (Automated Decision-Making Technology) — admt.ai.

Install

pip install agentloss

Quickstart

Instrument only the consequential action — the tool call that moves money or commits the business — not every LLM call.

from agentloss import decision, report_outcome, Decision

@decision                                     # bare decorator; the returned Decision is recorded
def approve_payment(invoice):
    action = run_matching(invoice)            # "approve" | "hold" | "reject"
    return Decision(action=action, value_at_risk_usd=invoice.total,
                    business_key=invoice.number, use_case="ap_3way_match")

# when the outcome resolves (correction, dispute, chargeback, audit, human review):
report_outcome(business_key="INV-1", ground_truth="duplicate-should-block",
               source="recovery_audit", realized_loss_usd=14200)

You already have the ground truth? (the common case — a disputes / chargebacks table). That's the default: each reported outcome is a census observation that counts toward the number, no flags needed. Join the whole table in one line:

from agentloss import record_outcomes

record_outcomes([
    {"business_key": "INV-1", "ground_truth": "reject", "source": "chargeback",
     "realized_loss_usd": 80.0},
    {"business_key": "INV-2", "ground_truth": "approve", "source": "dispute"},  # a CORRECT one
])

Report the outcomes that agreed with the agent too, not only the disputes — the rate's denominator is reported approvals, so reporting only errors makes it read ~100%. source is one of recovery_audit | dispute | chargeback | refund | human_queue | verification_agent.

It computes the error rate by segment (with confidence intervals), realized + expected dollar loss, and the agent's incremental risk vs. a baseline. Raw prompts/records stay in your boundary; only derived metrics leave.

Confirm the wiring — agentloss.doctor() inspects the store and catches the silent failures in plain language (outcomes reported but none counted, only-errors reported, a loss source that won't be summed). Or from a shell: agentloss doctor --json.

Works with your existing traces (Phoenix / Langfuse / Braintrust / OTel)

Already tracing your agent with OpenInference/OpenTelemetry? Don't re-instrument. Add a few agentloss.* attributes to the consequential span, point agentloss at your spans, and it adds the loss/outcome layer on top of what your tracer already emits:

from agentloss import ingest_spans, sample_and_verify, print_report

ingest_spans(your_spans)       # OTel/OpenInference spans carrying agentloss.* attributes
sample_and_verify(verify_fn)   # Tier A: get a number with no external labels wired
print_report()                 # error rate by segment + dollar loss

See examples/from_spans.py.

How it works

Instrument consequential actions, not the whole agent. The costly events are the handful of tool calls that move money or commit state.
Ground truth arrives late, from outside the agent — a correction, dispute, audit result, or human review. Capture it via report_outcome, the human-review queue, and active sampling
- a verification agent. This is real resolved outcomes, not an offline dataset.
Honest statistics. Monetary-unit sampling with a target verifier budget; two-phase calibration corrects a fallible verifier's bias back to truth (with confidence intervals).

See docs/SDK-SPEC.md for the full API, agentloss.* semantic conventions, and the pack/adapter model.

Try the demo

An oracle-validated harness that seeds an accounts-payable environment with known errors and checks that agentloss recovers the true error rate and dollar loss:

python -m dogfood.run                                  # deterministic mock, no deps
AGENTLOSS_VERIFIER_LLM=claude ANTHROPIC_API_KEY=... python -m dogfood.run

For AI coding agents

agentloss is built to be discovered and wired by coding agents: llms.txt, the instrument-agent-reliability skill, the AGENTS.md rule, and an MCP server (how_to_instrument, explain_attribute, validate_integration).

License

Apache-2.0.

Project details

These details have been verified by PyPI

Project links

Repository

GitHub Statistics

Maintainers

sanjayshukla

These details have not been verified by PyPI

Project links

Release history Release notifications | RSS feed

This version

0.0.11

Jul 1, 2026

0.0.10

Jul 1, 2026

0.0.9

Jul 1, 2026

0.0.8

Jul 1, 2026

0.0.7

Jul 1, 2026

0.0.6

Jul 1, 2026

0.0.5

Jul 1, 2026

0.0.4

Jul 1, 2026

0.0.3

Jul 1, 2026

0.0.2

Jul 1, 2026

0.0.1

Jul 1, 2026

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

agentloss-0.0.11.tar.gz (34.6 kB view details)

Uploaded Jul 1, 2026 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

agentloss-0.0.11-py3-none-any.whl (41.1 kB view details)

Uploaded Jul 1, 2026 Python 3

File details

Details for the file agentloss-0.0.11.tar.gz.

File metadata

Download URL: agentloss-0.0.11.tar.gz
Upload date: Jul 1, 2026
Size: 34.6 kB
Tags: Source
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for agentloss-0.0.11.tar.gz
Algorithm	Hash digest
SHA256	`7bd3b38ecdaa19afa81291a3356053eb3da3a521f74215280504bcf792c26427`
MD5	`163fdaa4fbc894958a5ffd0b03767fb6`
BLAKE2b-256	`2dbedea9b91016c62594934fd93da310165b5cc70381e9cafb4b193c280c47ae`

See more details on using hashes here.

Provenance

The following attestation bundles were made for agentloss-0.0.11.tar.gz:

Publisher: publish.yml on ADMT-ai/agentloss

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: agentloss-0.0.11.tar.gz
- Subject digest: 7bd3b38ecdaa19afa81291a3356053eb3da3a521f74215280504bcf792c26427
- Sigstore transparency entry: 2040220959
- Sigstore integration time: Jul 1, 2026
Source repository:
- Permalink: ADMT-ai/agentloss@584e13eb1c06905ea420569e07c2be55bcb392f8
- Branch / Tag: refs/tags/v0.0.11
- Owner: https://github.com/ADMT-ai
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: publish.yml@584e13eb1c06905ea420569e07c2be55bcb392f8
- Trigger Event: release

File details

Details for the file agentloss-0.0.11-py3-none-any.whl.

File metadata

Download URL: agentloss-0.0.11-py3-none-any.whl
Upload date: Jul 1, 2026
Size: 41.1 kB
Tags: Python 3
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for agentloss-0.0.11-py3-none-any.whl
Algorithm	Hash digest
SHA256	`6f1f84ab29123b55670673bea85738d9bd18ffebcc338b81a3f00726fc38a30d`
MD5	`7081fa8330b178d4a74a1cf594039e4c`
BLAKE2b-256	`d8b7006ad098057fc2f421145130d97f63db8cad3435d12f98767fcdb42b3db8`

See more details on using hashes here.

Provenance

The following attestation bundles were made for agentloss-0.0.11-py3-none-any.whl:

Publisher: publish.yml on ADMT-ai/agentloss

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: agentloss-0.0.11-py3-none-any.whl
- Subject digest: 6f1f84ab29123b55670673bea85738d9bd18ffebcc338b81a3f00726fc38a30d
- Sigstore transparency entry: 2040221198
- Sigstore integration time: Jul 1, 2026
Source repository:
- Permalink: ADMT-ai/agentloss@584e13eb1c06905ea420569e07c2be55bcb392f8
- Branch / Tag: refs/tags/v0.0.11
- Owner: https://github.com/ADMT-ai
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: publish.yml@584e13eb1c06905ea420569e07c2be55bcb392f8
- Trigger Event: release

agentloss 0.0.11

Navigation

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

agentloss

Install

Quickstart

Works with your existing traces (Phoenix / Langfuse / Braintrust / OTel)

How it works

Try the demo

For AI coding agents

License

Project details

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

Provenance

File details

File metadata

File hashes

Provenance