Skip to main content

Double Check Harness

Stop your AI agent from double-charging, double-emailing, or double-doing anything.

AI agents call tools that cause real-world side effects — charging a card, sending an email, creating a ticket. Networks are unreliable, so sometimes a call succeeds but the response never comes back. The agent doesn't know if it worked, so it retries — and the action happens twice.

Double Check Harness sits between your agent and the tool it's calling. Every action gets an idempotency key. If we've already seen that exact action, we hand back the original result instead of letting it fire again.

It works underneath whatever you're already using — LangChain, CrewAI, a raw function-calling loop — because it wraps the tool call, not the agent's reasoning.

Status

Early / pre-release. Three pieces are implemented and tested:

  • Python SDK — at the repo root (Development)
  • TypeScript SDK — under ts/ (Development)
  • Hosted API + dashboard — under server/, a multi-tenant HTTP service the SDK's HttpStore talks to, with a minimal live dashboard (see server/README.md)

Both SDKs implement: idempotency keys, an in-memory store, a Redis-backed distributed store with real concurrent-race handling, a tamper-evident hash-chained audit log, and pre-built connectors for Stripe, Twilio, SendGrid, and generic webhooks (see Connectors). The Python SDK also ships native middleware for LangChain and CrewAI tool lists (see Framework middleware). The server persists the same guarantees (atomic per-tenant key acquisition, per-tenant hash-chained audit log) over HTTP, behind a pluggable Backend interface with two implementations: SQLite for local dev/single-instance, and Postgres for real multi-instance production (see server/README.md for how the concurrency guarantees are enforced and proven under each). Not yet published to PyPI/npm (publish-ready — see Publishing); no billing or public signup flow yet. See PRD.md for the full product plan.

How it works (Python)

from double_check_harness import Guard, derive_key, guard

g = Guard()  # defaults to an in-memory store; pass a RedisStore for multi-process use

@guard(
    key=lambda customer_id, amount: derive_key(customer_id, amount, namespace="charge"),
    guard_instance=g,
)
def charge_customer(customer_id: str, amount: int):
    return stripe.charges.create(customer=customer_id, amount=amount)

The first call executes normally. If the agent retries with the same arguments — because of a timeout, a crash, or a confused retry loop — the wrapped function is not called again; the original result is returned instead.

Two callers racing on the same key at the same instant (not a retry-after-completion, a genuine concurrent collision) are handled explicitly too: by default the second caller gets a DuplicateInProgressError rather than silently double-executing; pass on_race="wait" to Guard to have it poll for the first caller's result instead.

Every attempt, block, and completion is recorded in a hash-chained AuditLog, so the history can be independently verified (g.audit_log.verify()) rather than trusted on faith — the basis for the audit-grade export in the compliance tier described in the PRD.

For production/multi-process deployments, back the guard with Redis instead of the default in-memory store:

import redis
from double_check_harness import Guard
from double_check_harness.store.redis_store import RedisStore

g = Guard(store=RedisStore(redis.Redis(host="localhost", port=6379)))

Or run against the hosted API (server/) instead of managing Redis yourself:

from double_check_harness import Guard
from double_check_harness.store.http_store import HttpStore

g = Guard(store=HttpStore(base_url="https://your-server.example.com", api_key="dch_live_..."))

How it works (TypeScript)

import { Guard, deriveKey, guard } from "double-check-harness";

const g = new Guard(); // defaults to an in-memory store; pass a RedisStore for multi-process use

const chargeCustomer = guard<[string, number], string>(
  (customerId, amount) => deriveKey([customerId, amount], "charge"),
  g,
)(async (customerId, amount) => {
  return stripe.charges.create({ customer: customerId, amount });
});

Same semantics as the Python SDK: a second call with the same derived key returns the original result instead of re-executing; a genuine concurrent race raises DuplicateInProgressError by default (or waits, with onRace: "wait"); every attempt/block/completion is recorded in a hash-chained AuditLog verifiable via g.auditLog.verify(). For multi-process deployments, back it with RedisStore from double-check-harness/redis (any ioredis-compatible client).

Connectors

Pre-built wrappers for the tool calls that most commonly get double-fired by agent retries. Each one derives a stable key from its arguments, runs the underlying client call through a Guard, and — where the provider itself accepts an idempotency key (Stripe does; Twilio and SendGrid don't) — forwards the same key so you get the provider's own dedup as a second line of defense on top of the Guard's.

Connector Python TypeScript
Stripe (charge, create_payment_intent) double_check_harness.connectors.stripe double-check-harness/connectors/stripe
Twilio (send_sms) double_check_harness.connectors.twilio double-check-harness/connectors/twilio
SendGrid (send_email) double_check_harness.connectors.sendgrid double-check-harness/connectors/sendgrid
Generic webhook (post) double_check_harness.connectors.webhook double-check-harness/connectors/webhook

None of them import their provider's SDK — they take an already-constructed client object and just need it to expose that provider's normal call shape, so installing double-check-harness never pulls in stripe/twilio/sendgrid/requests as a dependency.

import stripe
from double_check_harness import Guard
from double_check_harness.connectors.stripe import charge

g = Guard()
stripe.api_key = "sk_live_..."

charge(g, stripe, customer="cus_123", amount=500)  # a retry with the same args never double-charges
import Stripe from "stripe";
import { Guard } from "double-check-harness";
import { charge } from "double-check-harness/connectors/stripe";

const g = new Guard();
const stripeClient = new Stripe("sk_live_...");

await charge(g, stripeClient, { customer: "cus_123", amount: 500 });

The generic webhook connector (post) covers anything with no dedicated connector yet: it POSTs a JSON payload through any client shaped like requests/fetch, and sets an Idempotency-Key header (the convention Stripe/GitHub/etc. use) alongside the local Guard.

Framework middleware (Python)

Connectors wrap a specific third-party API; this wraps any tool in an agent's own tool list — LangChain or CrewAI — so a step the agent retries doesn't re-fire the underlying action, without changing how the agent calls the tool.

from langchain_core.tools import tool
from double_check_harness import Guard
from double_check_harness.integrations.langchain import guard_tools

g = Guard()

@tool
def send_invoice(customer_id: str, amount: int) -> str:
    """Send an invoice to a customer."""
    return billing_api.send_invoice(customer_id, amount)

tools = guard_tools([send_invoice, lookup_customer, refund_charge], g)
agent_executor = AgentExecutor(agent=agent, tools=tools)
from crewai.tools import tool
from double_check_harness import Guard
from double_check_harness.integrations.crewai import guard_tools

g = Guard()

@tool("Send Invoice")
def send_invoice(customer_id: str, amount: int) -> str:
    """Send an invoice to a customer."""
    return billing_api.send_invoice(customer_id, amount)

researcher = Agent(role="...", tools=guard_tools([send_invoice], g))

Both guard_tool/guard_tools work against whatever shape the tool actually is — a Tool/StructuredTool (.func/.coroutine), a custom BaseTool subclass instance (._run/._arun), or a plain function — and mutate it in place, so the object you pass to the agent keeps its identity. Async tools (.coroutine/._arun) are supported too: the guard's own store I/O runs off the event loop in a worker thread. Neither langchain nor crewai is a dependency of this package — the wrapping is duck-typed against each framework's public tool interface.

Why this exists

  • Idempotency is a well-understood pattern in distributed systems, but almost no agent framework gives it to you by default.
  • The one heavyweight existing answer (Temporal) requires restructuring your whole app around its workflow engine.
  • This is meant to be the lightweight version: wrap the risky call, get the guarantee, keep everything else the same.

Deployment

  • Hosted: point the SDK's HttpStore (TypeScript or Python) at a running instance of server/ and get the dashboard for free — every call and every duplicate blocked, live.
  • Self-hosted (library-only): skip the server entirely and back Guard with MemoryStore or RedisStore directly, as shown above.
  • Self-hosted / VPC platform (planned): run server/ itself inside a customer's own infrastructure for compliance-sensitive workloads, with only non-sensitive metadata reported to a central dashboard.

Pricing (planned)

Usage-based. Free tier for low volume, paid tiers scale with call volume, and a compliance/enterprise tier adds self-hosting, audit-grade exportable logs, and SLAs. Full detail in PRD.md.

Development (Python)

python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
pytest tests/ -v

Development (TypeScript)

cd ts
npm install
npm test        # vitest run
npm run build   # tsc -> dist/

Both suites cover the same guarantees: duplicate suppression on retry, independent execution for distinct keys, correct surfacing (not silent replay) of a previously-failed action, concurrent-race handling under both "raise" and "wait" modes, the decorator/wrapper form, the Redis-backed store's atomic acquire/done/release semantics under real concurrency, and audit-log chain integrity (including a test that deliberately tampers with a historical entry and confirms verify() detects it). Both SDKs also ship an HttpStore and an end-to-end integration test that spawns the real server/ as a subprocess and drives it over HTTP through Guard — not a mock of the server's contract.

Development (server)

cd server
npm install
npm test         # vitest run
npm run dev       # tsx watch src/index.ts, then open http://localhost:8787

See server/README.md for the full API reference, multi-tenancy model, and dashboard notes.

Publishing

Both packages are publish-ready as double-check-harness (name available, unclaimed on both registries as of this writing) but not yet published — this is a manual, deliberate step, not part of CI.

Python (PyPI), from the repo root:

python3 -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
rm -rf dist build src/*.egg-info
python -m build
twine check dist/*
twine upload dist/*          # prompts for your PyPI credentials/token

TypeScript (npm), from ts/:

cd ts
npm install
npm run build && npm test    # also runs automatically via prepublishOnly
npm login                    # if not already authenticated
npm publish

Bump version in both pyproject.toml and ts/package.json before each release — they're versioned independently but kept in sync by convention so far.

Roadmap

  • Core SDK (Python): idempotency key derivation, Guard/guard() decorator
  • Core SDK (TypeScript): mirrors the Python SDK's guarantees
  • In-memory store (dev/single-process) and Redis store (distributed, race-safe) — both languages
  • Tamper-evident hash-chained audit log — both languages
  • Hosted API (multi-tenant, pluggable persistence) + minimal live dashboard
  • HttpStore client (Python and TypeScript) so either SDK can run in hosted mode
  • Postgres backend for the server, safe across multiple instances (SQLite remains the local-dev default) — concurrency guarantees proven under real parallel load, not just asserted
  • Public signup + billing (the server currently only supports admin-token tenant creation)
  • Pre-built connectors: Stripe, Twilio, SendGrid, generic webhook — both languages
  • Audit-log export format for compliance/auditor consumption
  • Self-hosted/VPC deployment mode
  • LangChain / CrewAI native middleware (Python)
  • PyPI publish (Python) / npm publish (TypeScript)

Contributing

Not yet open for external contributions — this is in active spec/design. Watch this repo for updates.

License

MIT — see LICENSE.

Release files for double-check-harness 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for double-check-harness 0.1.0
File Size Uploaded
double_check_harness-0.1.0.tar.gz 30.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for double-check-harness 0.1.0
File Interpreter ABI Platform
double_check_harness-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 57.5 kB

Release files / double_check_harness-0.1.0.tar.gz

Download URL double_check_harness-0.1.0.tar.gz
Size 30.6 kB
Tags Source
SHA-256 checksum
How to use checksums
dad21b8b693058ba54d08f742f66d4f5e150b6c5b53239558ed0476e25f408dc
BLAKE2b-256 checksum
How to use checksums
67f505ebc25db4b0ed258256747ed99e886aa10c196a0e6190954a788da75465
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.13

Release files / double_check_harness-0.1.0-py3-none-any.whl

Download URL double_check_harness-0.1.0-py3-none-any.whl
Size 26.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
85b853477de8d5b4e52872155bc808abbfad9f8d50994ced6be4d01d49af2379
BLAKE2b-256 checksum
How to use checksums
88d25bd838cc11179279c178935ae76a9128d89d0b879a7e76d0699e60ae6e9e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.13

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page