ACEL — Agent Contract Enforcement Layer (core)
Runtime verification for AI agent tool calls. Declare temporal ordering contracts and Hoare-style pre/postconditions in plain Python, and have them enforced live against the stream of tool calls an agent makes — halting the agent the moment a rule is broken.
acel-coreships the transport-independent monitor (Phase 1) plus a live MCP proxy (Phase 2): the same contracts enforced against a real MCP server, via the official MCP Python SDK's request-middleware pipeline.
This is runtime verification — checking each concrete execution against a specification as it happens. It does not prove the agent correct in general; it guarantees that this run did not violate the rules you declared.
Why
Statistical agent-eval tools answer "how often does this agent behave well, on average?" ACEL answers the production question they can't: "did this execution just violate a rule we cannot allow to be violated?" — before the bad tool call lands.
New here? Read Install → Quickstart → The eight temporal templates and you'll have a working contract in a few minutes. Everything past that is reference material for specific needs (MCP, multi-tenant servers, config files, evidence/signing, metrics) — jump to whichever section matches what you're trying to do:
- Install · Quickstart · The eight temporal templates · Pre/postconditions
- Using ACEL outside MCP (LangChain / OpenAI function calling / any Python callable)
- Guarding Claude Code itself (block Bash/Edit/Write via hooks) · Content-aware matching (match on call content, not just tool name)
- Live MCP proxy · Multiple simultaneous clients (multi-tenant servers)
- Shadow mode (safe rollout) · Config-driven contracts (no-code rules files)
- Verifying evidence for tampering · Security notes (read this before production use)
- Metrics · Concurrency · Correctness / Performance (the numbers behind the claims)
Install
pip install acel-core
Requires Python 3.10+. No required dependencies for the core monitor — mcp,
pyyaml (config files), cryptography (Ed25519 signing), and
langchain-core are optional, installed as needed (see the sections that use
them below, or grab everything with pip install "acel-core[mcp,config,sign,langchain]").
Working from a clone of this repo instead (e.g. to run the tests or examples)?
git clone https://github.com/Purav-Kanda/Acel
cd Acel/acel-core
pip install -e ".[dev,mcp,config,sign,langchain]"
Quickstart
from acel import Session, must_precede, at_most_n_times
session = Session(state={"authenticated": False})
# Temporal ordering rules (no logic syntax required):
session.add_contract(must_precede("validate_record", "delete_record"))
session.add_contract(at_most_n_times("send_payment", n=1))
# State-based gate: reads only allowed once authenticated.
session.register_tool(
"read_user_data",
precondition=lambda s: s.get("authenticated") is True,
)
# Authentication commits trusted info into session state.
session.register_tool(
"authenticate",
commit=lambda s, args, result: s.set("authenticated", result["ok"]),
)
session.call("authenticate", {"user": "p"}, result={"ok": True})
session.call("read_user_data", {"query": "SELECT ..."}, result={"rows": []})
# This halts: delete before validate.
session.call("delete_record", {"id": "r_42"}) # raises ContractViolation
On violation, a ContractViolation is raised carrying a Violation record:
from acel import ContractViolation
try:
session.call("delete_record", {"id": "r_42"})
except ContractViolation as exc:
v = exc.violation
print(v.kind) # "temporal"
print(v.spec) # "must_precede(validate_record, delete_record)"
print(v.step) # index of the offending call
print(v.trace) # every call up to the violation
print(v.state_snapshot) # symbolic state at the moment it broke
The eight temporal templates
| Template | Meaning |
|---|---|
must_precede(a, b) |
every b must be preceded by some a |
at_most_n_times(a, n) |
a occurs at most n times per session |
at_most_total(a, field, limit) |
the sum of args[field] across all calls to a must not exceed limit |
never_after(a, b) |
a must never occur after b |
required_before_session_end(a) |
a must occur at least once before the session ends |
cannot_follow_without(a, b) |
a may not occur unless b occurred earlier |
mutually_exclusive(a, b) |
a and b must not both occur in one session |
rate_limit(a, n, window_seconds) |
a occurs at most n times in any rolling window_seconds-second window |
Each template is a deterministic automaton advanced in O(1) (amortized, for
rate_limit) per tool call, with a three-valued verdict (SATISFIED /
VIOLATED / UNKNOWN) over the finite trace. at_most_total and
rate_limit are the two templates that read more than just the tool
name: at_most_total sums a numeric argument field across calls (e.g. a
payment amount) and fails closed — a call missing the field, or with a
non-numeric value there, is treated as a violation rather than silently let
through, since silently ignoring an unreadable amount would be the actually
dangerous failure mode for a spend cap. rate_limit tracks wall-clock
timestamps of recent matching calls instead of a running session total, so
it caps bursts (calls per minute) rather than a per-session budget — the
two compose if you want both.
a/b/tool in every row above can also be a content-aware matcher
instead of a plain name — see Content-aware matching
for how to key a rule off what a call's arguments actually contain, not
just which tool it called.
from acel import Session, at_most_total, rate_limit
session = Session()
session.add_contract(at_most_total("send_payment", "amount", limit=500))
session.add_contract(rate_limit("send_payment", n=3, window_seconds=60))
session.call("send_payment", {"amount": 300}, result={"sent": True})
session.call("send_payment", {"amount": 250}, result={"sent": True}) # raises: 550 > 500
Pre/postconditions with decorators
from acel import Session, precondition, postcondition
@precondition(lambda s: s.get("authenticated") is True)
@postcondition(lambda s, r: r["tenant_id"] == s.get("current_tenant"))
def search_database(query): ...
session = Session(state={"authenticated": True, "current_tenant": "t_9"})
session.register(search_database)
A precondition can also inspect the current call's own arguments, not
just accumulated session state — write lambda s, args: ... instead of
lambda s: ... and ACEL detects the difference automatically (by
inspecting the function's signature) and calls it the right way; the
original 1-argument form keeps working exactly as before, no changes
needed to anything already written this way:
session.register_tool(
"run_shell",
precondition=lambda s, args: "rm -rf" not in args.get("command", ""),
)
This is genuinely different from content-aware temporal matching: a temporal contract answers "did a call matching this pattern happen (possibly zero times allowed)", while a 2-argument precondition can run any logic at all over the specific call being gated right now — combine it with state, check multiple fields together, whatever the rule actually needs. The tradeoff is the same one preconditions have always had: this is Python-only, not expressible in a JSON/YAML rules file, for the same reason arbitrary logic never has been (see Config-driven contracts).
Offline analysis / CI mode
Session.replay runs a recorded trace and returns every violation without
executing anything — the basis for the coming acel replay trace.json CLI and
for testing contracts against known-bad traces.
violations = session.replay([
{"tool": "delete_record", "args": {"id": "1"}},
])
Using ACEL outside MCP: LangChain, OpenAI function calling, or anything else
Session.call() is already framework-agnostic — MCP is just one caller of
it, not a requirement. acel.adapters has two small helpers that save you
the boilerplate of matching a specific framework's tool-call shape onto it:
from acel import Session, must_precede
from acel.adapters import guard
session = Session()
session.add_contract(must_precede("validate_record", "delete_record"))
# Wrap a plain function (a LangChain tool's `func=`, a dispatch-table
# callable, anything called as `the_tool(**kwargs)`) so every invocation
# is gated first — same enforcement, same evidence log, no MCP involved.
guarded_delete = guard(session, "delete_record", delete_record)
For a hand-rolled agent loop around an OpenAI-compatible chat completions
API, guard_openai_tool_call gates a raw tool_calls[i] entry directly:
from acel.adapters import guard_openai_tool_call
for tool_call in response.choices[0].message.tool_calls:
result = guard_openai_tool_call(session, tool_call, TOOLS[tool_call.function.name])
Full runnable demos: examples/langchain_agent_example.py (requires
pip install "acel-core[langchain]") and
examples/openai_function_calling_example.py (no extra dependency, no API
key needed to run it — uses hand-built tool_calls dicts shaped like the
real API's response).
Guarding Claude Code itself
Everything above wires ACEL into an agent you build. This section is
different: it plugs ACEL into Claude Code's own built-in tools (Bash,
Edit, Write, ...) via Claude Code's PreToolUse/PostToolUse hooks, so
it can block a real coding-agent tool call before it runs — using the exact
same temporal contracts as everywhere else.
Each hook fires as a fresh subprocess per tool call, so there's no
long-lived process to hold a live Session in memory. acel hook-pretooluse
handles this by reconstructing state on every invocation: it replays a
small trace file it maintains itself (one per Claude Code session) through
a fresh Session, then checks the new call against that reconstructed
state. acel hook-posttooluse appends each completed call to that trace
file so the next invocation sees it.
-
Write a rules file describing what to guard, e.g.
.claude/acel_rules.yaml:contracts: - template: rate_limit args: [Bash] kwargs: {n: 10, window_seconds: 60} - template: must_precede args: - {tool: Bash, matches: "pytest|npm (run )?test"} - {tool: Bash, matches: "git commit"}
The first rule catches a runaway retry loop (more than 10
Bashcalls in any 60-second window). The second is the more interesting one:BashandBashare the same tool name, so a plain tool-name rule can't tell a commit apart from a test run — the{tool: ..., matches: ...}form matches on the call's content instead (here, thecommandargument against a regex), so this actually blocks agit committhat wasn't preceded by a test run. See Content-aware matching for the full explanation, and the eight temporal templates for what else you can express this way. -
Add the hooks to
.claude/settings.jsonin your project (find your Python path with(Get-Command python).Sourceon Windows orwhich python3on macOS/Linux — using it directly, rather than relying onacelbeing onPATHin whatever environment Claude Code spawns hooks in, avoids the most common setup failure):{ "hooks": { "PreToolUse": [ { "matcher": "*", "hooks": [ { "type": "command", "command": "/path/to/python", "args": ["-m", "acel.cli", "hook-pretooluse", "--rules", "${CLAUDE_PROJECT_DIR}/.claude/acel_rules.yaml"] } ] } ], "PostToolUse": [ { "matcher": "*", "hooks": [ { "type": "command", "command": "/path/to/python", "args": ["-m", "acel.cli", "hook-posttooluse"] } ] } ] } }
-
Restart Claude Code. A blocked call surfaces to the model as a denied tool call with ACEL's reason attached, the same way any other permission denial does.
Unconditional bans work today too, declaratively. at_most_n_times(matcher, n=0)
blocks a specific dangerous call outright on its very first occurrence — a
hard "this is never allowed" rule, expressed with an ordinary template, no
Python required:
contracts:
- template: at_most_n_times
args: [{tool: Bash, matches: "rm -rf"}]
kwargs: {n: 0}
Scope, honestly: that covers "ban this pattern outright," but it's
still one regex against one field — "block rm -rf unless it's inside
/tmp" needs real conditional logic, which is what
argument-aware preconditions are for
(lambda s, args: ...). Those work identically through this same hook
pipeline when you wire the Session in Python — but preconditions have
never been YAML-file-representable, for the same reason arbitrary code
never is, so a hook-pretooluse --rules foo.yaml invocation specifically
can't reach one; that path only builds temporal contracts from the file.
If you need argument-aware preconditions with the hooks integration, build
your own thin wrapper around acel.hooks.pretooluse()/posttooluse()
that constructs the Session directly instead of going through
--rules.
See acel.hooks for the implementation and acel hook-pretooluse --help /
acel hook-posttooluse --help for the CLI reference.
Content-aware matching: matching on what a call actually does
Every temporal template — must_precede(a, b), at_most_n_times(a, n),
all eight — takes a/b/tool as either
a plain tool-name string (exact match, the original behavior) or a
content-aware matcher: something that also looks at the call's
arguments, not just its name. This is what makes the "commit before test"
rule above possible — git commit and pytest are both just Bash calls,
indistinguishable by name alone.
In Python, build one with matching():
from acel import Session, matching, must_precede
test_run = matching("Bash", r"pytest|npm (run )?test")
commit = matching("Bash", r"git commit")
session = Session()
session.add_contract(must_precede(test_run, commit))
matching(tool, pattern, field="command") matches a call to tool whose
args[field] (stringified) contains a case-insensitive regex match for
pattern. field defaults to "command" — the argument Claude Code's
Bash/PowerShell tools use — but works for any tool/field, e.g.
matching("Write", r"\.env$", field="file_path") to catch a write to a
.env file specifically, as opposed to any write.
In a JSON/YAML rules file, the same thing is a small dict wherever a plain tool name would otherwise go:
contracts:
- template: must_precede
args:
- {tool: Bash, matches: "pytest|npm (run )?test"}
- {tool: Bash, matches: "git commit"}
Still safe, still no code execution. A {tool, matches, field} dict is
just data — three strings — parsed with yaml.safe_load/json.loads same
as everything else in a rules file; there's no eval involved, and the
regex only ever runs against one argument value, never against your rules
file or anything else. This is a meaningfully different (and safe) thing
from a full precondition, which is why it's allowed in config files while
preconditions still aren't — see Config-driven contracts
for that reasoning in full.
If you need something a plain regex genuinely can't express — real
conditional logic, multiple fields combined, anything stateful — build a
ContentMatch directly with your own predicate function in Python instead
of matching(); that part, like preconditions, can't come from a data file.
Live MCP proxy (Phase 2)
ACEL can gate a real MCP server's tool calls, live, via the official MCP
Python SDK's ServerMiddleware hook. Every tools/call request passes
through ACEL's gate before the real tool handler runs — a blocked call has
zero side effects.
pip install "acel-core[mcp]"
from mcp.server.mcpserver import MCPServer
from acel import Session, must_precede
from acel.mcp_middleware import ACELMiddleware
session = Session()
session.add_contract(must_precede("validate_record", "delete_record"))
server = MCPServer("my-server", middleware=[ACELMiddleware(session)])
@server.tool()
def delete_record(record_id: str) -> dict: ...
See examples/toy_server.py for a complete toy server (5 tools, 3 contracts)
and tests/test_mcp_proxy.py for an end-to-end demo: a real ClientSession
talking to this server, with ACEL catching an ordering violation, a
cardinality violation, and a state-precondition violation — each one halted
before the tool it would have run.
Multiple simultaneous clients
ACELMiddleware(session) above wires one fixed Session shared by every
client that connects — fine for local testing or a server that only ever
has one client at a time, but unsafe once more than one client can connect
at once: every connection would read and write the same state, contracts,
and trace, so one client's calls could trip another client's rules or one
client's authentication could leak into another's session.
For a server meant to serve more than one client at once, pass
session_factory instead of session: ACEL builds a brand-new, fully
isolated Session the first time each connection is seen, and reuses it
for the rest of that connection's requests. Different connections never
share state, contracts, or trace.
from acel.mcp_middleware import ACELMiddleware
def build_session() -> Session:
session = Session(state={"authenticated": False})
session.add_contract(must_precede("authenticate", "read_user_data"))
return session
middleware = ACELMiddleware(session_factory=build_session)
server = MCPServer("my-server", middleware=[middleware])
Sessions are tracked internally by the MCP SDK's own per-connection
Connection object via a weak-reference map, so a session is released as
soon as its connection closes rather than accumulating forever on a
long-running server. See examples/multi_tenant_server.py for a complete
runnable example and tests/test_multi_tenant.py for an end-to-end proof
against two real, simultaneous ClientSession connections: one client
authenticates, its own follow-up call succeeds, and the other client's
identical call is still blocked, proving zero state leakage between them.
Shadow mode
The recommended way to roll out a new set of contracts: shadow mode detects and records every violation exactly as enforce mode does — same evidence log, same hash chain — but never blocks a call. Run it against real traffic first, see what it would have caught, then switch to enforce once you trust the rules.
session = Session(mode="shadow") # default is "enforce"
acel serve examples/toy_server.py --shadow
Session.call(), .precheck()/.postcheck() (the MCP proxy path), and the
CLI all respect mode. Session.replay() does not — it's a retrospective
CI-gate tool ("would this recorded trace have been blocked"), not a live
session, so it always reports every violation regardless of mode.
Config-driven contracts (no code required)
Temporal contracts can be declared in a plain JSON or YAML file instead of Python — useful for trying ACEL against your own tools without writing any code, or for keeping the rule set separate from your server implementation:
acel init-config rules.yaml # writes a starter file
acel validate rules.yaml # parses it, prints the contracts it declares
state:
authenticated: false
contracts:
- template: must_precede
args: [validate_record, delete_record]
- template: at_most_n_times
args: [send_payment]
kwargs: {n: 1}
Layer a rules file on top of a live server (--contracts adds to whatever
build_server() already sets up, and merges the state block in):
acel serve examples/toy_server.py --contracts rules.yaml
Or check a recorded trace against a rules file directly (the same format
acel replay has always used, now also parseable as YAML):
acel replay trace.json --rules rules.yaml
Why preconditions/postconditions aren't in the config file: they
evaluate real logic over state (lambda s: s.get("authenticated") is True),
and there's no safe way to deserialize arbitrary logic from a data file
without either an eval-style security hole or a bespoke expression
language. Temporal contracts have no such problem — every template is fully
described by tool names and simple parameters, so building one from a config
file is just constructing an object from validated data, no code execution
involved. Pre/postconditions stay in Python, wired directly to your tools —
install YAML support with pip install "acel-core[config]".
Naming a bundle of contracts as a group
Purely organizational — a group isn't a new kind of contract or a change to enforcement, it's a name for a bundle of contracts you keep referring to together. Worth it once a server has enough rules that "these three are the refund policy" is worth saying out loud:
session = Session()
session.add_contract_group("refund_policy", [
must_precede("verify_customer", "issue_refund"),
at_most_total("issue_refund", "amount", limit=500),
])
session.groups # {"refund_policy": [...]}
session.contracts_in_group("refund_policy")
Or declare it once in a rules file and pull it into contracts wherever
it's needed with {group: name}:
groups:
refund_policy:
- template: must_precede
args: [verify_customer, issue_refund]
- template: at_most_total
args: [issue_refund, amount]
kwargs: {limit: 500}
contracts:
- template: must_precede
args: [open_ticket, close_ticket]
- group: refund_policy
acel validate shows group membership alongside the flat contract list. A
group declared but never referenced from contracts has no effect — it's
inert until something pulls it in.
Verifying evidence for tampering
Every violation is recorded as a tamper-evident, hash-chained bundle. Save
one to disk and check it later — from a completely fresh process, with no
in-memory state — with acel verify:
acel replay trace.json --rules rules.json --save-evidence evidence.json
acel verify evidence.json
OK — 3 bundle(s) verified. Hash chain is intact, no tampering detected.
That checks hash-chain consistency — SHA-256 is unkeyed, so on its own it
can't prove authenticity against someone who can edit the file (they can
just recompute the hashes too). If you signed the log (ed25519_signer,
see Security notes below), always pass the public key to actually check
the signature:
acel verify evidence.json --public-key 4f2e...c19a
OK — 3 bundle(s) verified. Hash chain is intact and every signature checks out, no tampering detected.
If any field in any bundle was altered after the fact, acel verify fails
and reports the exact bundle index where the chain first breaks — everything
from that point onward is untrustworthy, but pinpointing where it broke is
what actually helps you investigate:
FAIL — tampering detected. Bundle 1 (of 5) is the first to break the chain...
To actually look at what's in an evidence log, rather than just check its
integrity, use acel show — a human-readable timeline instead of raw JSON:
acel show evidence.json --trace
ACEL Evidence Log — evidence.json
1 bundle(s), chain OK
[0] 2026-08-06T18:07:34.466187+00:00 step 1
kind: temporal
contract: must_precede(validate_record, delete_record)
tool: delete_record
args: {"id": "1"}
trace (0 call(s) leading up to this):
hash: fab359c33c… (unsigned)
--trace prints the full call history leading up to each violation, not
just the offending call; drop it for a shorter summary. If the chain is
broken, acel show marks the exact bundle where it happened the same way
acel verify does.
Metrics
Opt in to Prometheus-style metrics by passing a Metrics instance to Session:
from acel import Session, Metrics, must_precede
metrics = Metrics()
session = Session(metrics=metrics)
session.add_contract(must_precede("validate_record", "delete_record"))
# ... handle real traffic ...
print(metrics.render_prometheus())
Tracks call volume, gate latency (the same thing benchmarks/latency.py
measures offline, but live from your own traffic), and violation counts —
both overall by kind and broken down per contract, so you can see which
rule is actually tripping in production, not just that something did:
acel_calls_total 142
acel_violations_total{kind="temporal"} 3
acel_violations_total{kind="precondition"} 1
acel_violations_total{kind="postcondition"} 0
acel_contract_violations_total{contract="must_precede(validate_record, delete_record)"} 3
acel_gate_latency_seconds_count 142
acel_gate_latency_seconds_sum 0.000312
render_prometheus() just returns a string — serve it however fits your
deployment (a /metrics route on whatever web framework fronts your
server, a sidecar, a log line). If you don't already have an HTTP server to
hang a route off of, serve_metrics_http(metrics, port=9090) starts a
minimal stdlib-only one for you. Entirely opt-in: leave metrics unset and
none of this bookkeeping runs.
Concurrency
A Session is safe to call from more than one thread or async task at
once. Every method that mutates shared state (call, precheck,
postcheck, replay, end_session) is guarded by an internal
threading.RLock, so concurrent callers can't corrupt a contract's
internal counters, the step counter, or the trace — verified with real
ThreadPoolExecutor-driven tests firing dozens of concurrent calls at a
shared at_most_n_times/at_most_total/rate_limit contract and checking
for lost updates (tests/test_concurrency.py). The lock is never held
across an await: the MCP middleware's precheck() → (real tool
runs, unguarded) → postcheck() split exists specifically so a slow tool
call doesn't serialize every other in-flight request on the same
connection.
This is about safety within one Session, not about sharing one
Session across multiple clients — for that, see multi-tenant
session_factory support above, which gives each connection its own
fully isolated Session in the first place.
Security notes
-
Evidence bundles embed full call arguments, results, and state snapshots by default. That's what makes them useful evidence, but it also means anything sensitive passed as a tool argument (a password, a raw token, a secret) ends up persisted verbatim if you save an evidence log to disk or share it — unless you opt into redaction. Pass
redact_fields=toSession(or directly toEvidenceLog) with the dict-key names you consider sensitive, and any matching value anywhere in a violation's args/result/trace/state — nested dicts and per-call trace entries included — is replaced with a short, non-reversible hash marker before the bundle is hashed or signed:session = Session(redact_fields={"password", "api_key", "ssn"})
Two redacted entries with the same original value still produce the same marker (so "this session reused the same token twice" stays visible to an auditor). The marker is an HMAC-SHA256 keyed with a random, in-memory-only key that
EvidenceLoggenerates once and never writes to the log — recovering the original value from the marker requires that key, so an attacker who only has the evidence log (the threat this feature defends against) can't dictionary-attack it offline, even for a short/guessable value like a PIN. (If you callredact_violation()directly without going throughSession/EvidenceLog, it defaults to an unkeyed SHA-256 marker instead — safe for correlation and for high-entropy secrets, but brute-forceable for low-entropy ones; pass your ownkey=bytes if you need the same protection outsideEvidenceLog.) Fields you don't list are left alone, so still prefer keeping secrets out of tool arguments entirely where you can — pass a reference/ID and resolve the real secret inside your own tool implementation instead. -
ed25519_signer()can persist its key, or stay ephemeral — your choice. Called with no arguments, it generates a fresh, unpersisted key every time (fine for signing within one process's lifetime, but restart and old signatures stop matching the new public key). Called with a path (ed25519_signer("~/.acel/signing_key.bin")), it generates the key once, writes it to that file with owner-only permissions (0o600), and reuses it on every future call with the same path — so the public key, and every signature made against it, stays verifiable across restarts:sign, public_key_hex = ed25519_signer("~/.acel/signing_key.bin") session = Session(signer=sign)
Treat that key file exactly like an SSH private key: back it up if you need old signatures to keep verifying, and never commit it to a repo or evidence log.
-
The hash chain alone proves consistency, not authenticity — always verify with the public key if signing is enabled.
EvidenceLog.verify()andacel verify/acel showrecompute SHA-256 hash links, which is enough to catch accidental corruption, but SHA-256 is an unkeyed function: anyone who can edit the evidence-log file can also recompute every hash from an edited bundle onward, and the chain will still "verify" with no key required. If you're signing evidence (see above), always pass the public key so the signature is actually checked, not just its presence:acel verify evidence.json --public-key <hex from ed25519_signer> acel show evidence.json --public-key <hex from ed25519_signer>
Without
--public-key, both commands still run and print a warning — useful for a quick corruption check, but it is not tamper-evidence against a party who can write to the file. -
State-based preconditions are not safe against concurrent/pipelined calls to the same tool — use temporal contracts for anything that needs to hold under concurrency. A precondition only reads
session.state; the state isn't updated untilpostcheckcommits it, after the real tool has run. If a client has two calls to the same tool in flight at once (which MCP allows, and whichSessionsupports via itsprecheck/postchecksplit), both can read the same pre-commit state and both pass — so a precondition likelambda s: s["balance"] >= amountcan let two concurrent calls both pass against a balance that should only cover one of them. Temporal contracts (at_most_total,rate_limit,at_most_n_times) don't have this gap, because their counters are mutated synchronously under the lock duringprecheckitself — express spend caps, quantity limits, and rate limits as temporal contracts, not as a hand-written precondition, if concurrent calls are possible. -
Config files (
--rules,--contracts) are parsed withyaml.safe_loadandjson.loadsonly — neveryaml.loadoreval. There is no code execution path from a rules file; that's exactly why pre/postconditions can't be declared there (see above) — only tool names, counts, and plain values are ever deserialized.
Correctness
python benchmarks/correctness.py
A labeled dataset of 67 synthetic tool-call traces spanning all 8 temporal
templates (valid sequences, violating sequences, and edge cases like empty
traces and multiple simultaneous contracts) — measured at 100% precision
and 100% recall. Since the monitor is deterministic automaton checking, not
statistical detection, that's the expected result; the suite exists to prove
it and to catch any future regression (it's also wired into pytest as
tests/test_correctness_suite.py, so a miss fails CI directly).
Performance
python benchmarks/latency.py
Measured on the reference dev machine, 20,000 iterations, discarding a 1,000-call warmup: added p95 latency per tool call is ~0.005ms at 1 active contract and ~0.04ms at 50 concurrently active contracts — well under the <5ms target. Each temporal contract is a deterministic automaton advanced in O(1) per event, so overhead scales linearly with the number of active contracts, not with session length.
Tests
pip install pytest
pytest # core monitor + evidence (no extra deps)
pip install "acel-core[mcp]"
pytest tests/test_mcp_proxy.py tests/test_cli_serve.py # live MCP proxy + CLI
Testing against a real agent, not a script
Everything above proves ACEL works against scripted tool calls. For the
stronger version — a real LLM in Claude Desktop or Claude Code actually
driving the tool calls, and ACEL blocking a mistake the model made itself —
see docs/TESTING_WITH_REAL_AGENTS.md.
It walks through wiring up examples/support_agent_server.py (a realistic
customer-support/refund scenario) and gives adversarial prompts designed to
actually trigger each contract.
License
MIT
Metadata
Release files for acel-core 0.1.13
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| acel_core-0.1.13.tar.gz | 108.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| acel_core-0.1.13-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 171.9 kB
Release files / acel_core-0.1.13.tar.gz
| Download URL | acel_core-0.1.13.tar.gz |
|---|---|
| Size | 108.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5ef49cc3341888c47cef861e393ee50e9fce4de9b0e7123574928c03c00f5ec1
|
|
BLAKE2b-256 checksum How to use checksums |
6b65b19c6d9773d50dbb9a9f6aea20922cf6242a2edad8906c177a8529a1fce9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.3
|
Release files / acel_core-0.1.13-py3-none-any.whl
| Download URL | acel_core-0.1.13-py3-none-any.whl |
|---|---|
| Size | 63.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e817e5217fa921eb83e99265c0671ccfe0f51f7c4f5b4adf723e71b50c5c4a8d
|
|
BLAKE2b-256 checksum How to use checksums |
979c2b19ed6fd3a30c73a9e3b72eebe86c40ea4fd33e39ed0ffc2c3290177589
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.3
|