agentloss
Your eval tool tells you your AI agent's hallucination rate. agentloss tells you what it
costs. An OpenTelemetry-native SDK that measures the real-world error rate and dollar
loss of an AI agent's decisions — by capturing its consequential actions in-process and
joining them to ground truth (real resolved outcomes, not an offline labeled set).
Every eval/observability tool scores quality proxies — LLM-judge, hallucination rate, task
completion. agentloss answers the question the market keeps asking and no tool measures:
what are my agent's mistakes costing, and is it safe to trust with more autonomy?
Part of ADMT (Automated Decision-Making Technology) — admt.ai.
Install
pip install agentloss
Quickstart
Instrument only the consequential action — the tool call that moves money or commits the business — not every LLM call.
from agentloss import decision, report_outcome, Decision
@decision # bare decorator; the returned Decision is recorded
def approve_payment(invoice):
action = run_matching(invoice) # "approve" | "hold" | "reject"
return Decision(action=action, value_at_risk_usd=invoice.total,
business_key=invoice.number, use_case="ap_3way_match")
# when the outcome resolves (correction, dispute, chargeback, audit, human review):
report_outcome(business_key="INV-1", ground_truth="duplicate-should-block",
source="recovery_audit", realized_loss_usd=14200)
You already have the ground truth? (the common case — a disputes / chargebacks table). That's the default: each reported outcome is a census observation that counts toward the number, no flags needed. Join the whole table in one line:
from agentloss import record_outcomes
record_outcomes([
{"business_key": "INV-1", "ground_truth": "reject", "source": "chargeback",
"realized_loss_usd": 80.0},
{"business_key": "INV-2", "ground_truth": "approve", "source": "dispute"}, # a CORRECT one
])
Report the outcomes that agreed with the agent too, not only the disputes — the rate's
denominator is reported approvals, so reporting only errors makes it read ~100%. source
is one of recovery_audit | dispute | chargeback | refund | human_queue | verification_agent.
It computes the error rate by segment (with confidence intervals), realized + expected dollar loss, and the agent's incremental risk vs. a baseline. Raw prompts/records stay in your boundary; only derived metrics leave.
Confirm the wiring — agentloss.doctor() inspects the store and catches the silent
failures in plain language (outcomes reported but none counted, only-errors reported, a loss
source that won't be summed). Or from a shell: agentloss doctor --json.
Works with your existing traces (Phoenix / Langfuse / Braintrust / OTel)
Already tracing your agent with OpenInference/OpenTelemetry? Don't re-instrument. Add a few
agentloss.* attributes to the consequential span, point agentloss at your spans, and it adds
the loss/outcome layer on top of what your tracer already emits:
from agentloss import ingest_spans, sample_and_verify, print_report
ingest_spans(your_spans) # OTel/OpenInference spans carrying agentloss.* attributes
sample_and_verify(verify_fn) # Tier A: get a number with no external labels wired
print_report() # error rate by segment + dollar loss
How it works
- Instrument consequential actions, not the whole agent. The costly events are the handful of tool calls that move money or commit state.
- Ground truth arrives late, from outside the agent — a correction, dispute, audit result,
or human review. Capture it via
report_outcome, the human-review queue, and active sampling- a verification agent. This is real resolved outcomes, not an offline dataset.
- Honest statistics. Monetary-unit sampling with a target verifier budget; two-phase calibration corrects a fallible verifier's bias back to truth (with confidence intervals).
See docs/SDK-SPEC.md for the full API, agentloss.* semantic conventions,
and the pack/adapter model.
Try the demo
An oracle-validated harness that seeds an accounts-payable environment with known errors and
checks that agentloss recovers the true error rate and dollar loss:
python -m dogfood.run # deterministic mock, no deps
AGENTLOSS_VERIFIER_LLM=claude ANTHROPIC_API_KEY=... python -m dogfood.run
For AI coding agents
agentloss is built to be discovered and wired by coding agents:
llms.txt, the instrument-agent-reliability
skill, the AGENTS.md rule, and an MCP server
(how_to_instrument, explain_attribute, validate_integration).
License
Apache-2.0.
Metadata
Release files for agentloss 0.0.11
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| agentloss-0.0.11.tar.gz | 34.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| agentloss-0.0.11-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 75.7 kB
Release files / agentloss-0.0.11.tar.gz
| Download URL | agentloss-0.0.11.tar.gz |
|---|---|
| Size | 34.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7bd3b38ecdaa19afa81291a3356053eb3da3a521f74215280504bcf792c26427
|
|
BLAKE2b-256 checksum How to use checksums |
2dbedea9b91016c62594934fd93da310165b5cc70381e9cafb4b193c280c47ae
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 1, 2026.
Transparency logRelease files / agentloss-0.0.11-py3-none-any.whl
| Download URL | agentloss-0.0.11-py3-none-any.whl |
|---|---|
| Size | 41.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6f1f84ab29123b55670673bea85738d9bd18ffebcc338b81a3f00726fc38a30d
|
|
BLAKE2b-256 checksum How to use checksums |
d8b7006ad098057fc2f421145130d97f63db8cad3435d12f98767fcdb42b3db8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 1, 2026.
Transparency log