Skip to main content

Haldir

Scoped permissions, spend caps, an encrypted vault, and an audit log that can prove it wasn't edited — for AI agents that call tools, move money, and read secrets.

tests PyPI License: MIT GitHub Stars

A live Haldir audit log being tampered with: a past entry is rewritten, the inclusion proof stops matching the live Merkle root, and the verdict flips to 'Tamper detected'

That loop is the whole idea, running live. Someone rewrites a row in the audit log — silently, straight in the database. The entry's inclusion proof no longer matches the live Merkle root, and the verdict flips. Not caught by monitoring, not caught by a diff: caught by arithmetic, because the root is a hash of what the log actually contains and the earlier Signed Tree Head is already pinned somewhere you don't control.

→ Try it yourself — no install, runs in your browser

What you get

  • Scoped sessions — permissions and spend caps per agent, revocable the moment something looks wrong.
  • Encrypted vault — AES-256-GCM. Your agent asks for a secret; the model never sees it.
  • Tamper-evident audit — every call logged into an RFC 6962 Merkle tree with signed tree heads, so history can be proven, not just trusted.
  • Human approvals — pause a run on a spend threshold and get a webhook.
pip install haldir && haldir overview

Works with Claude Code, Cursor, LangChain, CrewAI, AutoGen, LlamaIndex and the Vercel AI SDK — anything that can make an HTTP call or speak MCP. MIT licensed: self-host it, or point at haldir.xyz (free tier, no signup).

See it in action

Here's what Haldir actually looks like — no diagrams, no spec sheets, just screenshots of the real thing.

Without Haldir vs with Haldir: no oversight vs scoped sessions, spend limits, secrets hidden, immutable audit trail

Without Haldir, an agent calls whatever API it wants, spends whatever it wants, and accesses whatever secret it finds — with zero oversight and zero audit trail. With Haldir, every action is scoped, spend-limited, logged immutably, and secrets never leave the vault.

Here's the three things you'd see as a new visitor, in order:

Quick tour: landing page, cloud dashboard, audit trail

  1. Landing page — dark mode, live terminal animation at the top, four product cards (Gate, Vault, Watch, Proxy), a self-host vs cloud comparison, and a call to claim a design partner spot. One page, everything a first-time visitor needs.

  2. Cloud dashboard — this is what you see after signing in. A sidebar on the left takes you to any page — account, quotas, sessions, audit, webhooks, approvals, compliance, or settings. The account view shows your tenant, tier, live counts, and API keys by prefix (the full key is never shown again after it's minted, and revoking one never involves a database shell).

  3. Audit trail — the killer feature. Filter by session, agent, or tool. Click any row to see the full MCP call details: what tool was called, what upstream API it hit, how long it took, what arguments it sent, and what it returned. This is the one thing that makes the whole product click — you can see exactly what every agent did, when, and with what.

Here's the dashboard with the important parts labeled:

Cloud dashboard with annotations: sidebar, stat cards, sessions table, audit table

The sidebar on the left takes you anywhere. The numbered markers point at the parts you'll actually use: your tenant and tier, the live counts, and your API keys by prefix — with the revoke button right there, so ending an agent's access never means opening a database shell.

Play with it yourself

Three things run live, no signup, straight from these links:

→ The tamper demo — the one in the GIF above. Rewrite a real log row and watch the inclusion proof stop matching the Merkle root. Nothing is simulated; it is the same Merkle code the API ships.

→ The playground — walks you through minting a key, opening a scoped session, checking a permission and writing to the audit trail, against a sandbox tenant of your own.

→ The gallery — every screenshot on this page in one place, if you'd rather look than read.

The rest of the site

The docs, pricing page, quickstart, compliance evidence pack, and every other page are linked from the nav bar on every page. Below: the full API reference, Python quickstart, performance numbers, and compliance mapping.

Try the real thing at haldir.xyz — free tier, no signup, point at it from any agent and go.


Two ways to run

Same product either way.

Self-host Cloud (haldir.xyz)
Price Free forever Free tier + paid plans
You run API + Postgres Nothing
Best for Regulated, air-gapped, "must own data" "Just make it work"

The cloud tier is free to start and needs no signup. We're taking 5 design partners — 30 days, full access, direct line to the founder: sterling@haldir.xyz.

Self-host in 5 minutes

git clone https://github.com/ExposureGuard/haldir.git
cd haldir
cp .env.example .env
python3 -c 'import base64, os; print(base64.urlsafe_b64encode(os.urandom(32)).decode())'
# paste the output into .env as HALDIR_ENCRYPTION_KEY, then:
docker compose up -d
curl http://localhost:8000/healthz

Full self-hosting guide: SELF_HOSTING.md

Cloud (no setup)

pip install haldir

That's it — point at https://haldir.xyz, no signup for the free tier.


CLI

Install once, drive the whole platform from the terminal:

$ haldir overview

  Haldir tenant overview
  acct_xyz123  ·  tier pro  ·  2026-04-19T18:42:11+00:00

  Status     ● ok
  Actions      4,217 / 50,000   ████░░░░░░░░░░░░░░░░    8.4%
  Spend      $ 47.30 this month
  Sessions        12 active  ·  3/10 agents
  Vault            8 secrets  ·  62 accesses this month
  Audit        1,847 entries  ·  0 flagged (7d)  ·  chain ✓
  Webhooks         2 registered  ·  541 deliveries (24h)  ·  99.82% success
  Approvals        1 pending
pip install haldir
haldir login                           # one-time; stashes API key
haldir overview --watch                # top-style live dashboard
haldir status                          # green/yellow/red component pills
haldir ready                           # exits 0/1, perfect for CI
haldir audit trail --agent my-bot      # the last N entries
haldir audit export --format=jsonl --out audit-2026-04.jsonl
haldir audit verify                    # hash chain integrity check
haldir webhooks deliveries             # last 20 retry attempts
haldir migrate up                      # apply pending schema migrations

Every command takes --json for scripts. haldir --help for the full surface.


Why Haldir

AI agents are calling APIs, spending money, and accessing credentials with zero oversight. Haldir is the missing layer:

Without Haldir With Haldir
Agent has unlimited access Scoped sessions with permissions
Secrets in plaintext env vars AES-256-GCM encrypted vault
No spend limits Per-session budget enforcement
No record of what happened Immutable, tamper-evident audit
No human oversight Approval workflows with webhooks
Agent talks to tools directly Proxy intercepts + enforces

Everything on the right is one process in front of your tools. Your agent keeps its existing tool calls; Haldir answers first:

Haldir architecture: Agent → Proxy → (Gate/Vault/Watch/Policy) → Upstream APIs


Quick Start (Python)

from sdk.client import HaldirClient

h = HaldirClient(api_key="hld_xxx", base_url="https://haldir.xyz")

# Create a governed agent session
session = h.create_session("my-agent", scopes=["read", "spend:50"])

# Store secrets agents never see directly
h.store_secret("stripe_key", "sk_live_xxx")

# Retrieve with scope enforcement
key = h.get_secret("stripe_key", session_id=session["session_id"])

# Authorize payments against budget
h.authorize_payment(session["session_id"], 29.99)

# Every action is logged
h.log_action(session["session_id"], tool="stripe", action="charge", cost_usd=29.99)

# Revoke when done
h.revoke_session(session["session_id"])

Under the hood that's four HTTP calls — mint a key, open a session, check a permission, write to the audit chain:

Haldir quickstart: install, create a scoped session, check permission, log the action to the hash-chained audit trail


Products

Gate — Agent Identity & Auth

Scoped sessions with permissions, spend limits, and TTL. No session = no access.

curl -X POST https://haldir.xyz/v1/sessions \
  -H "Authorization: Bearer hld_xxx" \
  -H "Content-Type: application/json" \
  -d '{"agent_id": "my-bot", "scopes": ["read", "browse", "spend:50"], "ttl": 3600}'

Vault — Encrypted Secrets & Payments

AES-encrypted storage. Agents request access; Vault checks session scope. Payment authorization with per-session budgets.

curl -X POST https://haldir.xyz/v1/secrets \
  -H "Authorization: Bearer hld_xxx" \
  -H "Content-Type: application/json" \
  -d '{"name": "api_key", "value": "sk_live_xxx", "scope_required": "read"}'

Watch — Audit Trail & Compliance

Immutable log for every action. Anomaly detection. Cost tracking. Compliance exports.

curl https://haldir.xyz/v1/audit?agent_id=my-bot \
  -H "Authorization: Bearer hld_xxx"

Proxy — Enforcement Layer

Sits between agents and MCP servers. Every tool call is intercepted, authorized, and logged. Supports policy enforcement: allow lists, deny lists, spend limits, rate limits, time windows.

# Register an upstream MCP server
curl -X POST https://haldir.xyz/v1/proxy/upstreams \
  -H "Authorization: Bearer hld_xxx" \
  -H "Content-Type: application/json" \
  -d '{"name": "myserver", "url": "https://my-mcp-server.com/mcp"}'

# Call through the proxy — governance enforced
curl -X POST https://haldir.xyz/v1/proxy/call \
  -H "Authorization: Bearer hld_xxx" \
  -H "Content-Type: application/json" \
  -d '{"tool": "scan_domain", "arguments": {"domain": "example.com"}, "session_id": "ses_xxx"}'

Approvals — Human-in-the-Loop

Pause agent execution for human review. Webhook notifications. Approve or deny from dashboard or API.

# Require approval for spend over $100
curl -X POST https://haldir.xyz/v1/approvals/rules \
  -H "Authorization: Bearer hld_xxx" \
  -H "Content-Type: application/json" \
  -d '{"type": "spend_over", "threshold": 100}'

MCP Server

Haldir is available as an MCP server with 19 tools for Claude, Cursor, Windsurf, and any MCP-compatible AI:

{
  "mcpServers": {
    "haldir": {
      "command": "haldir-mcp",
      "env": {
        "HALDIR_API_KEY": "hld_xxx"
      }
    }
  }
}

MCP Tools (the process above registers all 19):

Governance Tamper-evidence Approvals & compliance
haldir_create_session haldir_verify_audit_chain haldir_request_approval
haldir_get_session haldir_get_tree_head haldir_get_approval_status
haldir_check_permission haldir_get_inclusion_proof haldir_compliance_score
haldir_revoke_session haldir_get_consistency_proof haldir_build_evidence_pack
haldir_store_secret haldir_log_audit_action haldir_authorize_payment
haldir_get_secret haldir_query_audit_trail
haldir_list_secrets haldir_get_spend

There is one tool catalog. The stdio server (haldir-mcp, or haldir mcp serve) registers all 19; the hosted POST /mcp endpoint implements a 10-tool subset under the same names. A name means the same thing on both surfaces, so a client written against one works against the other.

MCP HTTP Endpoint: POST https://haldir.xyz/mcp


Performance

Haldir is fast enough to sit in the hot path of every agent tool call without becoming the bottleneck.

Single-box HTTP throughput (gunicorn 4 workers, 32 concurrent clients, tuned SQLite backend, every request goes through the full middleware stack — auth, validation, idempotency, metrics, structured logging):

Endpoint RPS p50 p95 p99
GET /healthz 1,638 19.1 ms 32.5 ms 41.6 ms
GET /v1/status 1,382 22.2 ms 30.8 ms 45.4 ms
GET /v1/sessions/:id 903 29.2 ms 95.5 ms 172.1 ms
POST /v1/sessions (create) 1,142 27.7 ms 35.2 ms 39.9 ms
POST /v1/audit (hash-chain) 1,092 28.7 ms 37.6 ms 52.6 ms

Hardware: 12th-gen Intel Core i3-1215U (8 cores, 8 GB RAM). SQLite is configured with WAL + synchronous=NORMAL + 256 MiB mmap + in-memory temp store — the session-lookup p99 dropped by 52 % versus the untuned path. Postgres deployments (configurable pool via HALDIR_PG_POOL_MIN/MAX) flatten the p99 further still; enable via DATABASE_URL=postgresql://....

Primitive cost (pure-Python, no I/O):

Primitive p50 Notes
Vault.store_secret (AES-256-GCM encrypt + AAD) < 10 µs in-memory, no DB write
Vault.get_secret (AES-256-GCM decrypt + AAD) < 10 µs in-memory
AuditEntry.compute_hash (SHA-256 over payload) < 10 µs
Gate.check_permission over REST ~50-120 ms network + DB round-trip, Cloudflare-fronted
Watch.log_action over REST ~50-150 ms includes chain lookup + DB write
Full governed-tool envelope (check + log) ~100-250 ms

Agents typically wait 500-3000 ms for an LLM completion and 100-1000 ms for an upstream API call, so Haldir's overhead sits inside the noise. Reproduce locally:

# Concurrent HTTP throughput (launches a local gunicorn, ~60s total)
python bench/bench_http.py --duration 10 --concurrency 32 --workers 4

# Primitive cost only (no API key needed)
python bench/bench_primitives.py --local

# End-to-end against the hosted service
export HALDIR_API_KEY=hld_...
python bench/bench_primitives.py

Compliance

One endpoint produces an auditor-ready proof-of-control pack covering eight sections, each anchored to a SOC2 trust services criterion:

haldir compliance evidence --since 2026-01-01 --out evidence-q1-2026.md
# Section SOC2
1 Identity (tenant, subscription, period) —
2 Access control (API keys + per-key scopes) CC6.1
3 Encryption (AES-256-GCM, AAD binding) CC6.7
4 Audit trail (entry count, hash chain) CC7.2
5 Spend governance (per-session caps) CC5.2
6 Human approvals (request/decision lifecycle) CC8.1
7 Outbound alerting (webhook delivery rate) CC7.3
8 Document signature (SHA-256 self-hash) —

The pack signs itself: a SHA-256 over the canonical JSON of sections 1-7. An auditor receiving an archived pack can re-call /v1/compliance/evidence/manifest and confirm the digest matches — proof the document was not modified after issuance.

JSON for evidence-locker upload, Markdown for the "show this to the auditor" moment, both from the same /v1/compliance/evidence endpoint.


Retention and deletion

Audit data is kept forever by default. When a policy requires otherwise, you can set a window and prune to it — and the prune stays provable:

haldir retention set 90      # keep 90 days (0 = forever)
haldir retention show        # what a prune would remove, before running it
haldir retention prune --yes

The audit log is a hash chain, so deleting old entries naively leaves the surviving chain pointing at a hash that no longer exists — which would turn a working audit trail into one that fails verification. Instead, a Signed Tree Head is taken over the log before anything is removed, and the hash of the last deleted entry is recorded as the link across the boundary.

The result is that pruning is not silent. haldir audit verify still passes, and reports what was removed along with the signed Merkle root that commits to it — so the honest answer to an auditor is "entries before this point were deleted under a retention policy, and here is the root they produced at the time." If that commitment cannot be produced, nothing is deleted.


API Reference

Full docs at haldir.xyz/docs — the complete OpenAPI 3.1 spec is at haldir.xyz/openapi.json.

Key endpoints (see the spec for the full surface):

Endpoint Method Description
/v1/keys POST Create API key
/v1/sessions POST Create agent session
/v1/sessions/:id GET/DEL Get / revoke session
/v1/sessions/:id/check POST Check permission
/v1/secrets POST/GET/DEL Store / list / delete secrets
/v1/payments/authorize POST Authorize payment
/v1/audit POST/GET Log / query actions
/v1/audit/spend GET Spend summary
/v1/audit/retention GET/PUT Read / set the retention window
/v1/audit/retention/prune POST Prune to the window (needs confirm)
/v1/audit/retention/checkpoints GET Prune history + signed commitments
/v1/approvals/rules POST Add approval rule
/v1/approvals/request POST Request approval
/v1/approvals/:id/approve POST Approve
/v1/approvals/:id/deny POST Deny
/v1/webhooks POST/GET Register / list webhooks
/v1/proxy/upstreams POST Register upstream MCP server
/v1/proxy/call POST Call through the proxy
/v1/usage GET Usage stats
/v1/metrics GET Platform metrics

Agent Discovery

Haldir is discoverable through every major protocol:

URL Protocol
haldir.xyz/openapi.json OpenAPI 3.1
haldir.xyz/llms.txt LLM-readable docs
haldir.xyz/.well-known/ai-plugin.json ChatGPT plugins
haldir.xyz/.well-known/mcp/server-card.json MCP discovery
haldir.xyz/mcp MCP JSON-RPC
smithery.ai/server/haldir/haldir Smithery registry
pypi.org/project/haldir PyPI

Design partners wanted

Live now: haldir.xyz · API Docs · OpenAPI Spec · Smithery

We're taking 5 design partners — 30 days free, full access, direct line to the founder. If you're shipping AI agents to production, email sterling@haldir.xyz.


License

MIT


Metadata

Release files for haldir 0.3.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for haldir 0.3.1
File Size Uploaded
haldir-0.3.1.tar.gz 243.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for haldir 0.3.1
File Interpreter ABI Platform
haldir-0.3.1-py3-none-any.whl Python 3 none any Details

Total release size: 488.3 kB

Release files / haldir-0.3.1.tar.gz

Download URL haldir-0.3.1.tar.gz
Size 243.3 kB
Tags Source
SHA-256 checksum
How to use checksums
21abb68cbb5469834f81729a6ae240181c207bb97090a9922b1f8bb58decda70
BLAKE2b-256 checksum
How to use checksums
7d18582c768d61a738513b24d84fcb64e741e9b6bb83d60d62d76987f9141b9c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 20, 2026.

Transparency log

Release files / haldir-0.3.1-py3-none-any.whl

Download URL haldir-0.3.1-py3-none-any.whl
Size 245.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
68ff94462b9ad24ef68d13888c6a6db883f71b3e03447cb6c997982aa0ab6947
BLAKE2b-256 checksum
How to use checksums
ad045ebe77c95f9f39cb42b102a054ea5885ae0011593d603321b424235905e8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 20, 2026.

Transparency log

Release history Release notifications | RSS feed

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.2

2 release files

This release

0.3.1 This release

2 release files

0.3.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page