Execution assurance framework for tool-using AI agents
Project description
Bracket
English | 한국어
A framework for verifying that AI agents actually did what they claim when they say "done."
Contract (what to do) + Evidence (what was done) = Verdict (was it really done)
Problem
LLM-based agents claim to read files without reading them, and claim to run tests without running them. Agent frameworks (LangGraph, Google ADK, etc.) are good at executing, but they do not verify that the required steps actually happened.
Bracket collects evidence during an agent's execution and mechanically decides pass/fail against a pre-defined contract.
When to use it
- When an agent claims to have modified code, and you want to confirm it didn't patch a file it never read.
- When you want to plug a verification step like "tests passed" into an agent pipeline.
- When you want to persist agent execution logs and re-judge them later under the same conditions.
- When you use multiple agent frameworks and want a single verification layer across them.
Install
pip install bracket-harness
The import name stays bracket (e.g. from bracket import Harness). Python 3.12+. No external dependencies for the core.
30-second example
from bracket import Harness, ExecutionContract
# 1. Define the contract: "modify code and make the tests pass"
contract = ExecutionContract.code_change(
goal="Fix failing test and verify it passes",
)
# 2. Collect evidence during execution
harness = Harness(app_name="my-agent", artifact_dir=".bracket")
run = harness.start_run(contract)
run.record_file_read("app.py", byte_count=1842)
run.record_file_changed("app.py")
run.record_command("pytest tests/", exit_code=0, kind="verification")
# 3. Judge
result = harness.finish_run_sync(run, final_output="Fixed the bug.")
print(result.verdict.outcome) # VerdictOutcome.VERIFIED
print(result.verdict.missing_requirement_ids) # []
The code_change contract is VERIFIED only when all of the following hold:
- The goal is resolved in the final output (intent resolved)
- The file was read before it was modified (read-before-write)
- At least one file was changed
- At least one command or tool was executed
- A verification command (pytest, etc.) was executed
- No hard failures occurred (tool permission denied, approval denied, etc.)
If any of these is missing, the outcome becomes BLOCKED or PARTIAL. The rules are defined in code; no LLM judges the result.
Layout
src/bracket/
core/ contracts, evidence, verdict, policy, approval
probes/ host-side verification (file, command, HTTP, git, pytest)
replay/ re-run the verdict from saved logs
adapters/ LangChain, LangGraph, Google ADK adapters
Built-in profiles
| Profile | Core requirements |
|---|---|
code_change |
read-before-write, file changed, verification command |
research |
file read, web fetch, grounding evidence |
file_task |
file changed, artifact emitted |
text_answer |
grounding evidence |
Every profile includes "intent resolved" and "no hard failure" conditions.
Probes
A probe checks the host environment directly before the verdict is computed. When an agent says "I created the file," a probe actually checks the disk.
from bracket.probes import PytestProbe, FilesystemProbe
result = harness.finish_run_sync(
run,
final_output="Done.",
probes=[
PytestProbe("tests/"),
FilesystemProbe("output.json", contains='"status": "ok"'),
],
)
| Probe | Verifies |
|---|---|
FilesystemProbe |
File existence, contains / does-not-contain |
CommandProbe |
Shell command exit code, stdout contents |
HTTPProbe |
HTTP response status, body contents |
GitDiffProbe |
Whether specific files appear in git diff |
PytestProbe |
pytest run result |
CustomProbe |
Arbitrary callable |
A failing probe is recorded as a hard failure in the verdict.
Adapters
Drop-in integrations for existing agent frameworks. Evidence is collected without modifying agent code.
LangChain
from bracket.adapters.langchain import BracketCallbackHandler
handler = BracketCallbackHandler(run)
agent.invoke(query, config={"callbacks": [handler]})
The callback observes tool calls and converts them to canonical evidence like file reads, web fetches, and shell commands.
LangGraph
from bracket.adapters.langgraph import BracketGraphHandler
handler = BracketGraphHandler(harness, contract)
result = graph.invoke({"input": "fix the bug"}, config={"callbacks": [handler.callback]})
artifact = handler.finish(final_output=result["output"])
Individual nodes can also be wrapped with a decorator:
@handler.node("code_writer")
def write_code(state):
...
return state
Google ADK
from bracket.adapters.google_adk import BracketADKHandler
handler = BracketADKHandler(harness, contract)
wrapped_tools = handler.wrap_tools([search_web, read_file])
Optional installs per adapter:
pip install bracket-harness[langchain]
pip install bracket-harness[langgraph]
pip install bracket-harness[google-adk]
pip install bracket-harness[all]
For frameworks without an adapter, call run.record_* methods directly or use GenericAdapter.
Artifacts
Each run is stored under .bracket/runs/<run_id>/:
contract.json contract definition
events.jsonl canonical event log collected during execution
summary.json event aggregates
probes.json probe results
verdict.json final verdict
replay.json replay manifest
metadata.json metadata such as app name
Everything is JSON, so other tools can read it. On POSIX, files are stored with 0600 permissions so they are isolated from other users on the same host.
Replay
Recomputes a verdict from a saved artifact. The external environment is not re-executed.
from bracket.replay import TraceReplay
verdict = TraceReplay(".bracket/runs/run_20260408_abc123").replay()
Useful for re-judging past runs in bulk when the requirement definitions change.
To also record LLM calls in LangGraph/LangChain, pass record_llm=True. On finish, llm_calls.json is written next to the other artifacts.
handler = BracketGraphHandler(harness, contract, record_llm=True)
# ... run graph ...
artifact = handler.finish(final_output=output)
# .bracket/runs/<run_id>/llm_calls.json holds the LLM request/response pairs
llm_calls.json stores prompts and responses verbatim, so it may contain sensitive data. The file is written with 0600 permissions and .bracket/ is in .gitignore, but leaks via backups, CI logs, or shared hosts are the caller's responsibility.
Policy & Approval
Risky actions are gated by a policy that allows, asks, or denies.
from bracket import Harness
from bracket.core.policy import PolicyRule, PolicyDecision, ActionKind, RiskLevel
harness = Harness(
app_name="my-agent",
artifact_dir=".bracket",
policy_rules=[
PolicyRule(ActionKind.SHELL, pattern="pytest", decision=PolicyDecision.ALLOW, risk_level=RiskLevel.LOW),
],
)
Defaults: file reads are ALLOW; shell, network, and file writes are ASK or DENY depending on risk level. Calling check_policy() evaluates the policy, and a DENY emits an approval_resolved event into the evidence stream, which the policy.no_hard_failures requirement detects. So even if the agent ignores the DENY and proceeds, the verdict becomes BLOCKED.
Security notes
- Artifact files (
.bracket/runs/*) are stored with0600on POSIX. On Windows they inherit the OS default ACL. FilesystemProbe'scontains/not_containscheck reads only the first 10 MB of the file. An infinite stream like/dev/zeroor a very large file cannot induce OOM.PolicyRule.patternuses substring matching. Use"*"to match all resources; the empty string is rejected.llm_calls.jsonwritten byrecord_llm=Trueholds prompts and responses in plaintext. Redacting sensitive data and controlling access to shared storage are the caller's responsibility.
What Bracket is not
- Not an agent framework (use LangGraph, OpenAI Agents SDK, etc. for that).
- Not guardrails (not input/output filtering, but execution verification).
- Not an observability tool (does not show logs, it decides pass/fail).
- Not an eval platform (not about response quality, but about execution completeness).
License
Apache-2.0
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file bracket_harness-0.1.0.tar.gz.
File metadata
- Download URL: bracket_harness-0.1.0.tar.gz
- Upload date:
- Size: 230.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
23d1941ae404a295484e15f27801d5af7992b65e9b4ab4e3677f790db56acf32
|
|
| MD5 |
19174b2ae8f21da18f74d8202a0ea53b
|
|
| BLAKE2b-256 |
b48e54a4a53382f94b40d1b2276dd7cb4fc1cc0aec401b132703a6e3e839c117
|
Provenance
The following attestation bundles were made for bracket_harness-0.1.0.tar.gz:
Publisher:
publish.yml on dybala-21/bracket
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
bracket_harness-0.1.0.tar.gz -
Subject digest:
23d1941ae404a295484e15f27801d5af7992b65e9b4ab4e3677f790db56acf32 - Sigstore transparency entry: 1339105880
- Sigstore integration time:
-
Permalink:
dybala-21/bracket@d918b8d65e467085650e86c00a402553a7eee4fe -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/dybala-21
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@d918b8d65e467085650e86c00a402553a7eee4fe -
Trigger Event:
push
-
Statement type:
File details
Details for the file bracket_harness-0.1.0-py3-none-any.whl.
File metadata
- Download URL: bracket_harness-0.1.0-py3-none-any.whl
- Upload date:
- Size: 47.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e436b48d69ea7dc2a1029a6c4faefa136f949dd0d33f074852a37a1a9807778d
|
|
| MD5 |
09927ff60007aca82e2ab311a9b86c90
|
|
| BLAKE2b-256 |
3db521250e5799d61a9d0c2973ecc92d087babe35effe6dfb3b54f42160e4d19
|
Provenance
The following attestation bundles were made for bracket_harness-0.1.0-py3-none-any.whl:
Publisher:
publish.yml on dybala-21/bracket
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
bracket_harness-0.1.0-py3-none-any.whl -
Subject digest:
e436b48d69ea7dc2a1029a6c4faefa136f949dd0d33f074852a37a1a9807778d - Sigstore transparency entry: 1339106100
- Sigstore integration time:
-
Permalink:
dybala-21/bracket@d918b8d65e467085650e86c00a402553a7eee4fe -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/dybala-21
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@d918b8d65e467085650e86c00a402553a7eee4fe -
Trigger Event:
push
-
Statement type: