agent-security-bench
Agent Security Evaluation Kit — score agent code and refuse unsafe tools.
Deterministic evaluation for AI coding agents on ML-oriented work and agent-security policies (prompt injection, jailbreak, MCP/tool overreach, SSRF, multi-tenant isolation), with machine-readable receipts.
Related: persian-llm-reference
One job: score what an agent produced, and prove whether it respected tool boundaries.
Why this exists
Frontier teams buy Agent Execution Assurance — not another chat wrapper. They need a shared way to:
- Score coding-agent artifacts on real ML failure modes (leakage, calibration, temporal splits, cache isolation).
- Regression-test tool policy under adversarial prompts (injection, confused deputy, SSRF args).
- Emit receipts that CI and buyers can verify —
acceptedonly when required checks pass.
What you get (v0.5)
| Piece | Purpose |
|---|---|
tasks/ml/ |
23 ML / agent-runtime coding tasks |
security/suites/ |
Core, extended, and production suites with severity weights |
schema/receipt-v1.json |
JSON Schema for evaluation receipts |
src/agent_security_bench/ |
AST + weighted checks, trace adapters, leaderboard batch |
mcp/fixture_server.py |
Allowlisted stdio tool server with path confinement |
examples/ |
Gold submissions, agent logs, live session traces |
docs/DEMAND_MAP_2026.md |
Buyer demand map → suite cases |
tests/ |
Engine + adapter unit tests |
See INDEX.md for the full catalog. Demand → buyer map: docs/DEMAND_MAP_2026.md.
Money suite (ABCP twin)
Public scorecard: https://banking.noetfield.com/scorecard
Install
pip install -U agent-security-bench==0.5.0
asb selftest
From checkout:
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
Tasks/suites/schema ship inside the wheel. Override with ASB_ROOT only if you keep a custom data tree.
Five-minute score
asb list
asb score \
--task tasks/ml/01_train_loop_bug.json \
--submission examples/submissions/01_train_loop_bug.py \
--out receipts/t01.json
asb security \
--suite security/suites/production_v1 / money_v1.json \
--agent-log examples/agent_logs/production_v1_clean.json \
--out receipts/s_prod.json
asb validate-receipt --receipt receipts/t01.json
asb leaderboard --out-dir receipts/leaderboard
Terminal walkthrough script (no extra deps):
bash scripts/five_minute_score.sh
Quick start (extended)
asb catalog --out /tmp/asb-catalog.json
asb score-trace \
--suite security/suites/core_v1.json \
--trace examples/traces/mcp_session_pass.jsonl \
--out receipts/trace_pass.json
asb selftest
Failing security example (expect rejected):
asb security \
--suite security/suites/production_v1.json \
--agent-log examples/agent_logs/production_v1_failing.json \
--out receipts/s_prod_fail.json
Scoring model
- Tasks: each check has
required+weight. Any required failure →outcome=rejected. Receipts also reportweighted_score. - Checks: string/regex,
python_parses, AST predicates (ast_has_call, …). - Security: allowlists, exfil bans, SSRF
must_not_call_with_arg_substr. - Traces: OpenAI-style / Anthropic tool_use / generic MCP JSONL → agent log → same suite scorer.
- Leaderboard: task × agent matrix →
leaderboard.json.
Receipt law
Every scored run emits a receipt: bench version, subject id, scores, per-check detail, timestamps. Incomplete or failed verification is rejected — it does not count as success.
See governance/METHODOLOGY.md and schema/receipt-v1.json.
MCP fixture
python mcp/fixture_server.py
# stdin/stdout JSON-RPC: initialize, tools/list, tools/call (list_dir + read_file only)
Status
Public v0.5.0 — Agent Security Evaluation Kit: wheel-bundled data, frontier task pack, live-trace adapter, leaderboard batch, demand map. Cite with CITATION.cff.
This kit does not train models, host GPUs, or replace a certified red team. It is a reproducible regression surface for agent evaluation and tool-policy scoring.
Author
Sina Kazemnezhad — github.com/sinakazemnezhad
Metadata
Release files for agent-security-bench 0.5.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| agent_security_bench-0.5.0.tar.gz | 33.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| agent_security_bench-0.5.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 84.9 kB
Release files / agent_security_bench-0.5.0.tar.gz
| Download URL | agent_security_bench-0.5.0.tar.gz |
|---|---|
| Size | 33.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
03b63383ce6b461e46c62ee1755bab13d6b6fd4ed124f810f77f66f835ee739d
|
|
BLAKE2b-256 checksum How to use checksums |
fb9460593ef1c896d86624c7a5f5f0d5f9d0f13db4813a2e635aa76f981626f1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 16, 2026.
Transparency logRelease files / agent_security_bench-0.5.0-py3-none-any.whl
| Download URL | agent_security_bench-0.5.0-py3-none-any.whl |
|---|---|
| Size | 51.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
1e5f247fc850ad8c51ea6c94c044f0e9db37af7a800a85a94fdfce628bba6ca1
|
|
BLAKE2b-256 checksum How to use checksums |
199669b5f3807dc75f901dd8c5829b1f90c9757fd77912d931d153e5ae2e339f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 16, 2026.
Transparency log