benchpress-agent
The reliability layer for AI agents with write access. Benchpress is a task-agnostic control loop you wrap around any tool layer. It reads the workspace's rules before deciding what "done" means. It picks the right record among look-alikes and locks the rest in a deny-list enforced in code. It plans only the writes the definition of done needs, gates every one, and reads every write back. Status comes from provider state, never from an HTTP 200. Every run ends with an auditable receipt.
pip install benchpress-agent # or: uv add benchpress-agent
Python 3.12+. Runtime dependencies: httpx, pydantic.
Quickstart
import asyncio
import benchpress
async def execute_tool(tool_name: str, tool_input: dict) -> dict:
# tool_name == "provider_api"
# tool_input == {"provider": "stripe", "method": "POST", "path": "/v1/customers/cus_123",
# "query": {...}, "body": {...}}
# Call your real API (or a sandbox/twin) and return:
return {"ok": True, "status_code": 200, "body": {...}}
agent = benchpress.wrap("deepseek-v4-pro", execute_tool, providers=["hubspot", "stripe", "slack", "gmail"])
result = asyncio.run(agent.run(
"Acme asked for renewal notices to go to ap@acme.example. Update billing and the CRM, "
"and keep the account owner in the loop. Do not send external mail.",
trace_dir="runs/acme", # writes receipt.json
))
print(result.status) # "completed" | "partial" | "escalated", from read-back evidence only
print(result.context.refusals) # every write the gate refused, with the rule that refused it
modelis a model id (DEEPSEEK_API_KEY/BENCHPRESS_API_KEY+BENCHPRESS_API_BASEfor any OpenAI-compatible endpoint,ANTHROPIC_API_KEYforclaude-*), abenchpress.ModelConfig, orNoneto readBENCHPRESS_MODEL.executoris an asyncexecute_tool(tool_name, tool_input)function or any object exposing one.- Built-in playbooks cover Slack, Gmail, HubSpot and Stripe; pass
playbooks=for your own systems. transport=swaps the model wire protocol (bring your own client, or a scripted one in tests).agent.run_sync(...)for synchronous callers.
A real-app gateway for Slack, Gmail, HubSpot and Stripe ships in benchpress.realapp
(gateway_from_env(providers)), and the CLI runs a request end to end:
benchpress run --providers hubspot,stripe --prompt "..." --trace-dir runs/demo
benchpress receipt runs/demo --html
The loop
P0 orient → P1 policy sweep → P2 resolve (lock look-alikes) → P3 definition of done → P4 plan → P5 execute through the gate + read-back → P6 verify (+ one repair round) → P7 deliver
| Principle | What it means in code |
|---|---|
| Code beats prompt for safety | The gate refuses writes to protected ids, prohibited verbs, sent (vs drafted) external mail, control-plane paths. The prompt only explains. |
| State is truth | Every write is read back; status is computed from evidence, never from a 2xx. |
| Task-agnostic | No code keyed to task ids, seeded names or domains. CI greps for them. |
| Provider content is data | Policies found in the workspace are classified; injection attempts are flagged, not obeyed. |
| Honest endings | P7 always runs, with whatever evidence exists. Timeouts and budget exhaustion end in a receipt, not a crash. |
Evidence
Benchpress was built for the Multi-App AI Agent Hackathon (2026-09-13) and is measured against ArgaBench's hardest published scenario, graded by ArgaBench's verifier on the published seed rebuilt locally and on real Slack, Gmail, HubSpot and Stripe (test mode). It makes no claim about the official leaderboard. Methods, results, failures and cost are in the repository: README · reports · disclosure.
Status
Alpha (0.x). The public API is benchpress.wrap, Benchpress.run, run_trial, TrialResult,
ModelConfig, Ablations, Gate. Expect additions (MCP guard, policy packs, Rehearse) in minor
releases; see the roadmap.
Apache-2.0 · Built by Raj Karia
Metadata
Release files for benchpress-agent 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| benchpress_agent-0.1.0.tar.gz | 194.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| benchpress_agent-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 313.8 kB
Release files / benchpress_agent-0.1.0.tar.gz
| Download URL | benchpress_agent-0.1.0.tar.gz |
|---|---|
| Size | 194.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
afd8796575d77442ad6d22db23f0adae9fd6c494a753cffd4e6b96c59a041ca4
|
|
BLAKE2b-256 checksum How to use checksums |
8142700ae57ce292f1c4bf2bad09756105f59bdd5d05977ba757c1cd12714f21
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.6 {"installer":{"name":"uv","version":"0.11.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / benchpress_agent-0.1.0-py3-none-any.whl
| Download URL | benchpress_agent-0.1.0-py3-none-any.whl |
|---|---|
| Size | 119.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
bd27502d868c2d9ce8e968d75dfb41bfb652639a5d683624d08a70711e1ddb78
|
|
BLAKE2b-256 checksum How to use checksums |
f7ed8f564c3785de57e6da500ad6f1f10fe822e62c5883fb8d03adcaeca29ca4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.6 {"installer":{"name":"uv","version":"0.11.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|