Attestry
Verifiable trust infrastructure for AI agents. Four things that keep going wrong in agent systems, fixed by one mechanism underneath all four.
No server. No account. No API keys required. No runtime dependencies.
pip install attestry
The four gaps
| The problem | Attestry's answer | |
|---|---|---|
| Model drift | Providers update models behind stable names. Nothing errors; behaviour just changes. You find out from a customer. | Record what the behaviour was, re-check it nightly, and report movement separately from failure. |
| Tool schema chaos | MCP wants inputSchema. Anthropic wants input_schema. OpenAI wants parameters. LangChain wants a pydantic model. CrewAI wants a BaseTool subclass. |
Define the tool once; generate all ten shapes, and report what each translation loses. |
| Invisible access | An agent reads your mail, calendar and files. The audit trail, if any, is JSON nobody reads. | Receipts in plain English, checked against consent grants that can actually refuse. |
| Unverifiable skills | Sharing agent skills means trusting a registry server to keep telling the truth. | Content-addressed packages, signed provenance, and a ledger you can sync over git. |
The one mechanism
All four write to the same thing: an append-only, hash-chained, signed ledger — one JSONL file you commit alongside your code.
That is what makes this one tool instead of four scripts in a repository. A drift run, a mailbox access, a tool schema change and a package publication are all entries in the same chain, so one command verifies every claim any subsystem has ever made:
$ attestry ledger verify
OK 12 entries, 12 signed, chain intact
head 0edf10907b3ea58671f46413ddd43818d612cf4af019ec2c86154b244a5d4f56
merkle 638c02e307e766c681e1473d6f4200770a2c6736e9fa46e86c5dba1b51f2859e
$ attestry ledger stats
12 entries
drift 3
ledger 3
receipt 2
registry 2
schema 2
Edit any byte of any entry and every hash after it stops matching. Signatures mean appending needs a key, not just write access to the file.
Sixty seconds, no API key
The bundled echo provider is a deterministic stand-in, so the whole loop works
offline and free. Set ATTESTRY_ECHO_VARIANT=b and it answers differently —
which is how you can watch drift detection actually fire.
mkdir demo && cd demo
attestry init .
attestry drift run example # first run records the baseline
ATTESTRY_ECHO_VARIANT=b attestry drift run example --diff
The second run:
suite example
model echo:demo
!! model identity changed: 'demo' was serving 'demo-variant-a', now serves
'demo-variant-b' (3 case(s)). Any movement below is explained by this.
case status score detail
--------------------------------------------------------------------------
refund-window drift 0.535 behaviour moved (similarity 0.535)
tone-check drift 0.530 behaviour moved (similarity 0.530)
length-guard drift 0.485 behaviour moved (similarity 0.485)
DRIFT (3 drift)
--- refund-window ---
--- baseline
+++ current
-Answer: How long does a customer have to request a refund?. Confidence: high.
+I think the answer here is how long does a customer have to request a refund?,
+though it depends on context.
Exit code 1. In CI that fails the build, and the first line already tells you why the outputs moved.
Then point it at a real model — same suite, same commands:
attestry drift run example --provider ollama --model llama3 # free, local
attestry drift run example --provider anthropic --model claude-sonnet-5
1. Model-drift detection
A suite is a JSON file you commit: prompts, and what counts as a correct answer.
{
"name": "support",
"provider": "anthropic",
"model": "claude-sonnet-5",
"params": { "temperature": 0 },
"drift_threshold": 0.95,
"cases": [
{
"id": "refund-window",
"prompt": "How long does a customer have to request a refund?",
"system": "You are a concise support agent.",
"expect": { "mode": "contains_all", "value": ["refund", "14 days"] }
},
{
"id": "no-apology-boilerplate",
"prompt": "Explain our uptime guarantee to an annoyed customer.",
"expect": { "mode": "never_contains", "value": ["sorry for the inconvenience"] }
},
{
"id": "structured-output",
"prompt": "Classify this ticket. Reply with JSON.",
"expect": { "mode": "json_shape", "value": { "category": "string", "urgency": "integer" } }
}
]
}
Expectation modes: exact, normalized, contains_all, contains_any,
never_contains, regex, json_shape, numeric, similarity.
Why two verdicts per case, not one
Every case is graded twice, and keeping these apart is the whole point:
- Did it meet the expectation? A pass/fail assertion. Fails loudly.
- Did the behaviour move? Compared against the last accepted output for the same prompt.
A case can keep passing while the answer changes character completely — longer, differently formatted, suddenly hedged. That is drift, and it is the early warning. By the time your assertions start failing you have usually been shipping different behaviour for weeks.
So the exit codes differ: 1 for moved, 2 for broken.
The model-identity check
Every provider response carries the model string the provider actually served. Attestry compares it against what the baseline saw — not against the alias you asked for.
That distinction matters. gpt-4o resolving to gpt-4o-2024-08-06 is normal.
That snapshot changing between runs is the event this whole subsystem exists
to catch, and comparing only against the alias would miss it entirely. OpenAI's
system_fingerprint is captured for the same reason.
What goes where
| Contents | Committed? | |
|---|---|---|
| the ledger | digests, scores, statuses | yes — signed, tamper-evident |
drift/baselines/ |
the accepted outputs | yes — this is what a reviewer reads |
drift/runs/ |
full outputs of every run | no — gitignored |
Model outputs echo whatever was in the prompt. Writing them into a signed append-only log that you then push to a shared repository would take a privacy problem and make it permanent. The digest proves what was seen without disclosing it.
attestry drift show support refund-window # when exactly did this change?
entry when output digest status served model
#2 2026-09-25T23:54:03Z f3c959bfa51e ok claude-sonnet-5
#4 2026-09-26T04:11:44Z 560a1b58ecf8 drift claude-sonnet-5 <- changed here
2. One tool definition, ten frameworks
attestry schema ingest their-mcp-tools.json # start from tools you already have
attestry schema emit send_invoice --all --out ./adapters
Targets: mcp, openai, ollama, anthropic, gemini, json-schema,
langchain, autogen, crewai, pydantic.
The last four emit Python source, because those frameworks do not consume a
JSON description of a tool — they want a pydantic model, an annotated signature,
a BaseTool subclass:
# Generated by Attestry from the neutral tool schema for 'send_invoice'.
# NTS interface digest: d2699c38f63837e9
# Declared effects: writes billing records; uses the network; can destroy data
class SendInvoiceInput(BaseModel):
'Email an invoice to a customer.'
customer_id: str = Field(..., description='Account id.')
amount: float = Field(..., description='Amount in USD.', ge=0)
cc: Optional[List[str]] = Field(None, description='Extra recipients.')
class SendInvoiceTool(BaseTool):
name: str = 'send_invoice'
description: str = 'Email an invoice to a customer.'
args_schema: Type[BaseModel] = SendInvoiceInput
def _run(self, **kwargs: Any) -> Any:
raise NotImplementedError('implement SendInvoiceTool._run')
Note the header. The declared effects survive into the framework that has nowhere to put them, so "this tool can destroy data" does not get lost in translation.
Every emitter reports what it drops
$ attestry schema emit send_invoice --target openai --strict
note (openai): strict mode: cc became nullable-but-required, so the model will
pass null rather than omitting it
note (openai): no place to declare side effects; a caller cannot tell from this
payload that the tool writes or destroys data
Those losses are exactly where moving a tool between frameworks silently changes its meaning. Round trips are tested: emit to MCP, OpenAI (including strict mode) and Anthropic, parse back, and names, types and requiredness all survive.
Tool schemas drift too
attestry schema register send_invoice
send_invoice: interface changed -- escalation
escalation (4):
escalation consent.data_classes now touches financial, contact_info
escalation consent.egress data can now leave to mail.example.com
escalation effects.destructive can now destroy data
escalation effects.writes now writes customer billing records
breaking (1):
breaking force new required parameter
Four severities, and escalation is the one no other schema tool reports:
the call signature did not break, but the tool now does more to the world than
it used to. That is the severity that matters when you are about to install
somebody else's skill.
breaking stops existing callers. behaviour still works but may act
differently. cosmetic is prose. Rewording a description is not a change —
the digest covers the interface, not the help text.
3. Consent receipts a human can read
attestry receipts record --actor agent:inbox-triage --action read \
--resource email:gmail/INBOX --purpose "draft daily priorities" \
--items 42 --fields subject,from --data-classes message_content
On 25 Sep 2026 at 23:54 UTC, the inbox-triage agent read 42 messages in your
Gmail inbox (subject and from only), in order to draft daily priorities. No data
left your machine.
ledger entry #7
actor agent:inbox-triage
action read on email:gmail/INBOX
data touched message_content
consent permitted by grant inbox-triage
unusual first recorded access by agent:inbox-triage
Structured audit logs already exist and nobody reads them, because reading one requires knowing the schema. If the person whose inbox it was can read that sentence and think "it did what?", the system worked.
The renderer holds to two rules. It never invents specificity — an unrecognised resource keeps its raw identifier rather than being dressed up in prose that might be wrong. And it always states the egress position in both directions, because "no data left your machine" is the sentence people most want, and silence does not convey it.
Grants can refuse, not just record
.attestry/receipts/consent.json, committed:
{
"default_deny": true,
"grants": [
{
"id": "inbox-triage",
"actor": "agent:inbox-triage",
"resources": ["email:*/INBOX"],
"actions": ["read", "list", "search"],
"purposes": ["draft daily priorities", "flag urgent messages"],
"data_classes": ["message_content", "contact_info"],
"egress": [],
"max_items": 200
}
]
}
Purpose is part of the permission. "May read your calendar" and "may read your calendar in order to schedule meetings" are different grants, and the second is the one people think they gave. Most access-control systems leave this out entirely.
In Python, with enforce=True, an ungranted access is stopped before your code
touches anything:
from attestry.receipts import Recorder, ConsentPolicy
recorder = Recorder(ledger, ConsentPolicy.load(ws.consent_path),
actor="agent:inbox-triage", enforce=True)
with recorder.access("read", "email:gmail/INBOX",
purpose="draft daily priorities",
fields=["subject", "from"],
data_classes=["message_content"]) as event:
messages = mailbox.fetch(limit=50)
event.items = len(messages) # filled in after the fact
The consent check runs on entry. The receipt is written on exit either way — including when the body raises, because "the agent tried to read your mailbox and crashed" is exactly the event an investigator wants and the one a naive implementation drops.
It flags what is out of character
Compared against recorded history, not statistics:
unusual agent:inbox-triage has only ever read before; this is its first
write-type action (send)
unusual touched 900 items, more than three times the previous high of 42
unusual data sent to api.example.com for the first time
unusual first access to financial by this actor
Denied attempts are counted, never credited
attestry receipts digest --period "18-25 September 2026"
5 accesses by 2 agents, touching 42 items in total.
Nothing was changed, sent or deleted.
No data left your machine.
3 attempts refused by policy (one of which would have sent data to api.example.com).
the inbox-triage agent
your Gmail inbox -- 42 items
refused: send on your Gmail Sent folder (reply to customers)
refused: read on your Gmail inbox (train a model)
stated purposes: draft daily priorities
A refused read touched nothing, so its item count and its intended egress stay out of every total. A receipt that overstates what happened is worse than no receipt.
Also: attestry receipts html --out receipts.html for a self-contained page
with no scripts and no external assets, and attestry receipts check to
re-check recorded history against the policy as it stands today — which catches
grants that were later narrowed or expired.
4. A skill registry with no server
attestry registry pack ./my-skill
attestry registry publish ./my-skill --source https://example.com/my-skill-1.0.0.tar.gz
attestry registry install my-skill@^1.0.0
The ledger is the index. A package's digest is a Merkle root over its file tree — not over the archive bytes, because tar and zip embed timestamps and ordering, and a content address that depends on who ran the build is not a content address.
What that substitution buys, cryptographically rather than socially:
- provenance — publications are signed, so "who published this" is a fact about the entry, not a claim about an account;
- immutability — republishing a version with different content is refused, because the earlier entry is still in the chain saying what that version was;
- transparency — you cannot show one person one version and somebody else a different one without producing two ledgers that visibly fork;
- no operator — sync over git, a URL, a shared drive, a USB stick.
$ attestry registry publish ./hello-web
REFUSED: hello-web@1.0.0 was already published as 9013e3f93e80 at entry #1, but
this package hashes to 8a2c1f0b4e91. Published versions are immutable -- release
a new version instead.
Installing verifies entirely in memory before a single byte reaches the filesystem, and keeps working afterwards:
$ attestry registry install hello-web@1.0.0 --from ./tampered.tar.gz
attestry: hello-web@1.0.0 does not match the ledger: expected 9013e3f93e80, the
fetched archive hashes to d0c17f2c8281. The source has been altered since
publication.
$ attestry registry verify ./installed hello-web@1.0.0
MISMATCH: on disk hashes to d0c17f2c8281, ledger says 9013e3f93e80
Archive members that try to escape the extraction directory are rejected rather than extracted, because installing a stranger's skill is exactly where that matters.
What it deliberately does not solve
Namespace arbitration. Two people can publish pdf-tools from different keys,
and no amount of hashing decides which one deserves the name.
$ attestry registry conflicts
pdf-tools@1.0.0 was published 2 times with different content:
#14 2026-09-20T09:12:03Z digest 9013e3f93e80 by 2185f517b640c9c6
#31 2026-09-24T17:40:55Z digest 61ba0c7d4e12 by af5a0aef8f5ad858
The same version exists with different content. Nothing can decide this
automatically -- trust only the publisher you meant to trust.
Trust settles it: you install from keys you chose, which is roughly how you already decide whose code to run.
Verification, in more detail
Trust modes
| Mode | Behaviour | Use it for |
|---|---|---|
strict |
only key ids you added by hand | CI, consuming other people's skills, audits |
tofu |
trust on first use, then pinned — SSH host keys | a single developer, the default |
open |
signatures checked, trust not enforced | local development only |
Public keys travel inside the ledger in ledger.key-trusted entries, so the
file is self-describing: hand somebody the ledger and one trusted key id and
they can check the whole thing.
Revocation asks a question most tools skip
attestry ledger revoke <keyid> --reason "rotated" --invalidates after
attestry ledger revoke <keyid> --reason "leaked" --invalidates all
Rotation means everything signed beforehand is still good. Compromise means you cannot believe anything it ever signed, because whoever held the key could have backdated entries. Getting this wrong in either direction is harmful, so Attestry refuses to guess and makes the flag required.
Proving one entry without revealing the others
attestry ledger checkpoint # publish a Merkle root
attestry ledger prove 7 --out proof.json # ~log2(n) hashes
attestry ledger prove --check proof.json # anyone, offline
You can prove to an auditor that a particular access was recorded when you say it was, without disclosing everything else the agent touched that week. The proof re-hashes the entry from its own contents, so it cannot claim inclusion for contents that were swapped afterwards.
Syncing and forks
attestry ledger export shared.jsonl
attestry ledger sync https://example.com/ledger.jsonl
Merging is only ever a fast-forward. Incoming entries are verified before they are merged, never after — appending somebody else's file to your own signed log and checking it later would mean your ledger had already vouched for content you had not examined. Genuine divergence is reported, not resolved:
FORK: ledgers agree up to #3 then diverge: local a3ded7794f4d, incoming cd0bbeff8044
Histories diverge, so no automatic merge is possible. Compare the two
ledgers and decide which history is the real one.
Which history is the true one is a question about the world, not about the data.
Cryptography
Ed25519 throughout, with two interchangeable backends:
- pure Python, bundled, following the RFC 8032 reference implementation. Tested against the RFC's own vectors.
cryptography, used automatically when importable (pip install attestry[fast]): constant-time and roughly 100× faster.
Same key format, same 64-byte signatures. A test signs with each and verifies with the other.
The bundled implementation ships because a registry whose signatures only verify after you install a native dependency is not much of a decentralised registry. It is not constant-time — point multiplication branches on secret scalar bits. For signing your own ledger on your own machine that is an acceptable trade, and verification touches only public data. On shared hardware, install the fast backend.
Other choices worth knowing about:
- Canonical JSON (close to RFC 8785) before anything is hashed or signed,
with UTF-16 code-unit key ordering so a JavaScript implementation agrees.
NaN,Infinityand non-string keys are rejected rather than coerced. - Merkle trees follow RFC 6962 with domain separation — leaves prefixed
0x00, nodes0x01— so an internal node cannot be replayed as a leaf. Proofs are tested exhaustively for every tree size from 1 to 33, including truncated, padded and wrong-index proofs. - Signatures cover the canonical payload, not the hex hash, so a verifier
never has to trust the
hashfield it was handed. - Chain verification compares recomputed hashes, so editing a body and
leaving the
hashfield alone still breaks the links after it.
In CI
- run: pip install attestry
- run: attestry ledger verify --require-signatures # 4 if the log was touched
- run: attestry schema register # 1 on escalation/breaking
- run: attestry drift run support # 1 moved, 2 broken
| Code | Meaning |
|---|---|
| 0 | fine |
| 1 | something moved — drift, or a schema changed |
| 2 | something failed — an expectation broke, a check did not pass |
| 3 | something could not run — provider unreachable, bad arguments |
| 4 | something is untrustworthy — broken chain, bad signature, untrusted key |
Code 4 is deliberately separate. A failing test and a ledger that does not verify are different emergencies.
Python API
from attestry import Workspace, Ledger, Keyring, TrustStore
ws = Workspace.init(".")
key = Keyring(ws.keys_dir).create("ci")
ledger = Ledger(ws.ledger_path, key=key, trust=TrustStore(ws.trust_path))
# 1. drift
from attestry.drift import Suite, run_suite, get_provider, record_run
run = run_suite(Suite.load("support.suite.json"), get_provider("ollama", "llama3"))
record_run(ledger, run)
# 2. schema
from attestry.schema import ToolSchema, Param, emit, from_callable
payload, notes = emit(from_callable(my_function), "mcp")
# 3. receipts
from attestry.receipts import Recorder, ConsentPolicy
recorder = Recorder(ledger, ConsentPolicy.load(ws.consent_path), enforce=True)
# 4. registry
from attestry.registry import Package, Registry
Registry(ledger, trust, ws.cache_dir).publish(Package.from_dir("./my-skill"))
assert ledger.verify().ok # covers all four
Layout
.attestry/
ledger.jsonl the log: one canonical JSON entry per line
trust.json keys this machine accepts signatures from
config.json workspace settings
keys/<keyid>.json private keys, owner-readable, gitignored
drift/suites/*.json prompts and expectations
drift/baselines/*.json accepted behaviour <- commit this
drift/runs/*.json full outputs, gitignored
schemas/*.nts.json neutral tool schemas
receipts/consent.json what each agent may touch, and why
registry/cache/ verified packages, gitignored
Commit .attestry. The private keys, package cache and raw run logs are
machine-local and the directory carries its own .gitignore for them.
The ledger is a plain append-only text file on purpose. You can read it with
tail, review it in a pull request, and merge it by concatenation. A database
would have been easier to query and much harder to trust.
Documentation
| docs/ARCHITECTURE.md | why one ledger, and how the four subsystems compose |
| docs/LEDGER.md | the on-disk format, precisely enough to reimplement |
| docs/NTS.md | the Neutral Tool Schema spec and per-target mappings |
| docs/DRIFT.md | suite format, expectation modes, CI patterns |
| docs/RECEIPTS.md | the access model, grants, and the rendering rules |
| docs/REGISTRY.md | package format, digests, trust, syncing |
Limitations, stated plainly
- Drift similarity is lexical, not semantic. It will not notice that two differently worded answers mean the same thing. For a regression detector that is the correct trade: the question is whether the output changed, not whether it is still true. It also means no embedding model, no network, and identical results on every machine.
- Receipts are only as honest as their instrumentation. Attestry records
what your code tells it. It is a transparency and consent layer, not a sandbox
— it cannot see an access that never called it.
enforce=Truerefuses accesses that go through the recorder, and nothing else. - No consensus. Two people appending offline produce two valid ledgers. Forks are detected, never silently resolved.
- No namespace authority in the registry, by design. Trust decides.
- The keyring is developer-grade. Owner-only file permissions where the platform supports it; not a hardware token.
Status
0.1.0. The ledger format is versioned and documented; changes to it will come with a migration path. 269 tests, no runtime dependencies, Python 3.9+, tested on Linux, macOS and Windows with both crypto backends.
License
MIT — Nulfied
Metadata
Release files for attestry 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| attestry-0.1.0.tar.gz | 142.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| attestry-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 272.2 kB
Release files / attestry-0.1.0.tar.gz
| Download URL | attestry-0.1.0.tar.gz |
|---|---|
| Size | 142.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d977a334087989ad68e9af41e2fd7d5404c7ae93681a9398d98d09dddee61767
|
|
BLAKE2b-256 checksum How to use checksums |
da110edd8168419c87bc5970879f91fff94e2af6111296a70fb170ea5049d9a7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency logRelease files / attestry-0.1.0-py3-none-any.whl
| Download URL | attestry-0.1.0-py3-none-any.whl |
|---|---|
| Size | 129.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7c51697b9ec98558c3d0ea8f2e0bda0abe66156ad730e8f609bb1d14bfeb6708
|
|
BLAKE2b-256 checksum How to use checksums |
aa0577aadfce17247df6bec1136faf834190e3cef238e8cf8244dc639f297f48
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency log