Wallbreaker Hermes: AI Red-Team Harness
██╗ ██╗ █████╗ ██╗ ██╗ ██████╗ ██████╗ ███████╗ █████╗ ██╗ ██╗███████╗██████╗
██║ ██║██╔══██╗██║ ██║ ██╔══██╗██╔══██╗██╔════╝██╔══██╗██║ ██╔╝██╔════╝██╔══██╗
██║ █╗ ██║███████║██║ ██║ ██████╔╝██████╔╝█████╗ ███████║█████╔╝ █████╗ ██████╔╝
██║███╗██║██╔══██║██║ ██║ ██╔══██╗██╔══██╗██╔══╝ ██╔══██║██╔═██╗ ██╔══╝ ██╔══██╗
╚███╔███╔╝██║ ██║███████╗███████╗██████╔╝██║ ██║███████╗██║ ██║██║ ██╗███████╗██║ ██║
╚══╝╚══╝ ╚═╝ ╚═╝╚══════╝╚══════╝╚═════╝ ╚═╝ ╚═╝╚══════╝╚═╝ ╚═╝╚═╝ ╚═╝╚══════╝╚═╝ ╚═╝
break the wall · not the rules of engagement ⚔ authorized testing only
Wallbreaker Hermes is an AGPL-licensed fork of
Wallbreaker. It preserves the standard
Wallbreaker harness and adds an opt-in native laboratory target for Hermes Agent runtimes. The
laboratory uses an ephemeral home but does not provide operating-system sandboxing. Wallbreaker
provides a Claude-Code-style terminal (CLI command: wallbreaker)
that reasons and calls tools in a loop. The backend is fully configurable, so
it runs on OpenRouter, the Z.AI GLM coding plan, the local Claude Code CLI, a
local server, or any OpenAI-/Anthropic-compatible API (including third-party proxies via
bearer-auth). It ships with a deep red-team toolkit: the Parseltongue transform engine,
the L1B3RT4S jailbreak library, the HarmBench behavior benchmark, automated attack
loops (PAIR/TAP, Crescendo, best-of-N), a from-scratch persona author, native-format
target mimicry from a leaked system-prompt corpus, a multimodal image-edit attack channel,
an LLM judge, and reliability validation.
For authorized security testing only.
Project status
The standard Wallbreaker CLI, TUI, dashboard, MCP server, attack tools, judge, and reliability
validation work as documented below. The Hermes laboratory adapter targets the fixed Hermes Agent
v2026.8.13 revision. It supports one text turn with a clean home or selected SOUL, memory, and
working-directory rules. It rejects tools, MCP, custom prompt layers, profiles, prefill,
continuation, and multimodal input. See Hermes Native Laboratory. The
campaign API and wallbreaker hermes run|review|verify CLI load versioned YAML suites, run isolated
repetitions through the existing autonomous loop, and write HMAC-scoped JSON evidence for manual
review. The optional operator skill lives under integrations/hermes/.
The PyPI distribution is wallbreaker-hermes. The import package and commands remain
wallbreaker and wb for upstream compatibility.
Highlights
- Dual-protocol provider layer: OpenAI Chat Completions + Anthropic Messages, any
base_url/model. Captures reasoning/thinking channels; converts network errors to clean failures (no crashes on timeout). - Autonomous attack loop: keeps mutating/re-firing until it succeeds (
finish()exits the tool) or needs you (ask_operator()). - Standardized, unbiased prompts: pulls test batteries from HarmBench (400 behaviors, 7 categories) instead of hand-picked examples.
- Reliability-first:
validatere-fires N times for the real success rate; a one-shot COMPLIED is never called a "bypass". Pin the OpenRouter backend for reproducibility. - Parseltongue: 59 chainable transforms (encodings, unicode fonts, stego, homoglyph,
zero-width, tag smuggling, bijection, gibberish…) plus
mutate(LLM anti-classifier). - P4RS3LT0NGV3 over MCP: an optional MCP server wraps elder-plinius's upstream
Parseltongue: all 222 transforms (45 ciphers, runic/braille/symbol scripts, every
encoding, steganography) + a universal decoder, exposed as
parsel_*tools the agent drives directly. Any[[mcp.servers]]you configure is proxied into the tool registry. - Single-artifact convergence:
/sysprompt+system_sweep+optimize_universalconverge on ONE universal system prompt; they can't split into variant toolkits. - Persona author (
author_persona): writes a full devoted-persona system-prompt jailbreak from scratch via the codified ENI method (draft → self-critique → validate → refine → distill), auto-picking a credentialed-authority or limerence register from the objective's domain. - Native-format mimicry:
sysprompt_*tools search a leaked product system-prompt corpus (Claude/GPT/Gemini/Grok…) and hand the target's own section-tag/heading dialect to the persona author so a payload speaks the victim model's native format. - Multimodal image channel:
query_image_editfires an image + instruction at an image target and vision-judges the result;image_chainruns a Chain-of-Jailbreak, decomposing a refused image into a ladder of benign edit steps. Plus Tier-3 T2I framing transforms. - Pluggable attacker brains: OpenAI/Anthropic APIs, or the local Claude Code CLI
(
protocol = "claude-code", keyless) as the red-team brain. Third-party Anthropic proxies work viaauth_style = "bearer". - Extended attack arsenal (this fork):
cipherchat(CipherChat/SelfCipher, ICLR 2024) teaches the target a cipher in-band then fires in ciphertext;skeleton_key(Russinovich 2024) reframes the guardrail as a policy amendment with a "Warning:" label;persuasion_attack(PAP, Zeng 2024) rewrites the ask through 16 persuasion strategies concurrently and ranks bypasses;drattack(Li 2024) decomposes the objective into benign fragments then reassembles;ica(Wei 2023) packs N harmful Q/A demos into a single in-context turn.
Clone Repository
git clone https://github.com/Yivas/wallbreaker-hermes
cd wallbreaker-hermes
Install
Install the release from PyPI:
pip install wallbreaker-hermes
For local development from a clone:
python -m venv .venv
. .venv/bin/activate
pip install -e ".[dev]" # add [barcodes] for the QR/barcode tool
Configure
cp config.example.toml config.toml # add your keys (config.toml is gitignored)
wallbreaker check # validate it: profiles, keys, target, judge
Profiles set the attacker brain; [target] is the model under attack; [judge] grades
replies. Keys can be inline or from env. OpenRouter endpoints support provider pinning
and a timeout override.
Hermes Agent targets use protocol = "hermes-lab" with a dedicated checkout, Python
interpreter, provider credential environment variable, and closed manifest. They remain opt-in
and do not change any standard target defaults. See Hermes Native Laboratory
and the operator skill.
default_profile = "glm"
[profiles.glm]
protocol = "openai"
base_url = "https://api.z.ai/api/paas/v4"
api_key = "..."
model = "glm-4.6"
[target]
protocol = "openai"
base_url = "https://openrouter.ai/api/v1"
api_key = "sk-or-..."
model = "deepseek/deepseek-v4-pro"
# provider = "WandB" # pin the backend for reproducible results
# timeout = 60 # seconds (default 120)
[judge]
protocol = "openai"
base_url = "https://openrouter.ai/api/v1"
api_key = "sk-or-..."
model = "openai/gpt-4o-mini"
Attacker-brain options:
# Local Claude Code CLI as the attacker brain, keyless (the CLI self-auths).
[profiles.claude-code]
protocol = "claude-code"
model = "sonnet"
# system_prompt_file = "operator.md" # optional: leads the harness tool doctrine
# Third-party Anthropic-compatible proxy that wants an OpenAI-style bearer token.
[profiles.proxy]
protocol = "anthropic"
base_url = "https://your-proxy.example" # host root; provider appends /v1/messages
api_key = "..."
model = "claude-sonnet-4"
auth_style = "bearer" # Authorization: Bearer <key> (default: x-api-key)
P4RS3LT0NGV3 engine (native)
The full upstream P4RS3LT0NGV3 engine (222 transforms across 11 categories plus the
universal decoder) is wired straight into the agent registry as native parsel_* tools
(parsel_list/search/inspect/transform/chain/decode/guide/craft). No MCP
server or config block is required; the tools appear automatically once the repo is vendored
and Node.js is on PATH. One-time setup:
wallbreaker parsel update # git-clone elder-plinius/P4RS3LT0NGV3 into library/ (needs Node.js)
wallbreaker parsel list # sanity-check: prints all 222 transforms by category
If Node is missing, the pure-Python parseltongue tool (50+ transforms) remains as an
offline fallback. Override the vendored location with PARSEL_REPO=/abs/path/to/P4RS3LT0NGV3.
MCP servers (optional)
The harness is also an MCP client: every [[mcp.servers]] you declare is spawned over stdio
at startup and its tools are proxied into the registry. The same P4RS3LT0NGV3 engine is still
available as an MCP server if you prefer to run it out-of-process (it re-registers the same
parsel_* names with identical behaviour):
[[mcp.servers]]
name = "parsel"
command = "python"
args = ["-m", "p4rs3lt0ngv3_mcp"]
enabled = false # native tools already cover this
# tool_prefix = "p_" # optional namespace for the proxied tools
# env = { PARSEL_REPO = "/abs/path/to/P4RS3LT0NGV3" } # override the vendored repo location
No npm install/build is needed; the server drives the upstream Node bridge headlessly.
In the TUI, /parsel guide|list|search <q>|inspect <key> browses the catalog. The server is
a standalone stdio MCP server, so any MCP client (Claude Code, Cursor) can use it too.
Launch
wallbreaker # TUI on default_profile
wallbreaker --profile openrouter
wallbreaker --auto "objective..." # one-shot autonomous run
wallbreaker --resume # reopen the autosaved session (survives a crash/Ctrl+C)
The TUI autosaves the whole engagement to sessions/autosave.json after every turn;
--resume reopens it (or pass a specific session file: wallbreaker --resume mysession.json).
Picking the model to attack
/model changes the attacker brain; /target changes the victim.
/target anthropic/claude-3.7-sonnet attack any model on the target endpoint
/target glm attack via a profile
/provider WandB pin the OpenRouter backend (reproducibility)
Finding ONE universal prompt (the right way)
A "one prompt for every task" goal means a single fixed artifact, not a toolkit.
/sysprompt set <one system prompt> hold ONE fixed system prompt
/sysprompt test sweep it across the HarmBench cyber battery
/validate <task> re-fire 8x for the REAL success rate
Read which tasks failed → refine the one prompt → /sysprompt test again. A single
COMPLIED is luck; validate tells you the truth. For the user-turn variant use
/template set … {request} + /template test.
Agent tools (/tools lists them live)
| tool | purpose |
|---|---|
run_shell, read_file, write_file, edit_file |
build/run/save payloads |
parseltongue, parseltongue_catalog, mutate |
obfuscate / anti-classifier rewrite |
parsel_* (native) |
full P4RS3LT0NGV3 engine: parsel_guide/list/search/inspect/transform/chain/decode: 222 transforms + universal decoder. parsel_craft builds a ready-to-fire payload (encode a request through a chain + wrap it decode-and-comply / split-into-vars) |
l1b3rt4s_*, eni_* |
jailbreak libraries: L1B3RT4S + the ENI persona collection |
author_persona |
author a full devoted-persona system prompt from scratch (ENI method: draft→critique→validate→refine→distill), auto-picking an authority/limerence register from the objective's domain |
sysprompt_list, sysprompt_search, sysprompt_get, sysprompt_native |
browse/search a leaked product system-prompt corpus (Claude/GPT/Gemini/Grok…); sysprompt_native hands the target's own section-tag/heading format to the persona author for native mimicry |
harmbench, preset |
unbiased behavior benchmark, curated seed templates |
query_target |
fire at the model-under-test (with transforms=[...] to encode+fire) |
query_image_edit |
fire an input image + instruction at an IMAGE target (modality='image') and vision-judge the edited picture |
image_chain |
Chain-of-Jailbreak: decompose a refused image into a ladder of individually-benign edit steps and drive them in sequence |
multi_fire |
sweep one payload through many encodings (concurrent) |
crescendo |
multi-turn escalation |
pair_attack |
PAIR/TAP: refine one objective on the target's refusals |
pair_sweep |
run the PAIR loop across a whole battery concurrently (highest-ASR, batched) |
best_of_n |
resample N times, keep the bypass |
many_shot |
many-shot jailbreak: flood context with faux compliant turns, then fire |
prefill |
response-priming: seed the assistant's own reply so it continues, not refuses |
narrate |
fiction-frame + in-story prefill (novel-chapter roleplay); tops the scoreboard |
diff_fire |
A/B two payloads at one target to attribute ASR to a specific edit |
recommend_transforms |
survey ~16 encodings, rank by bypass, synthesize a chain to try |
seed_sweep |
inject one request through many ENI+L1B3RT4S seeds, rank which bypass |
fire_file |
fire a file/seed RAW (verbatim, full-length) as the target system prompt |
adapt_seed |
attacker-LLM patches a persona for a specific refusal (don't use it to distill) |
campaign |
auto-escalate a HarmBench battery up a technique ladder, coverage matrix |
leaderboard |
rank multiple profiles by ASR on one battery (robustness benchmark) |
leak_scan |
scan a reply for secrets/PII/system-prompt echo (evidence, not a verdict) |
scan |
Garak-style coverage matrix (technique + HarmBench probes) |
indirect_inject |
RAG/agent injection via document/email/tool-output carriers |
system_sweep |
validate ONE system prompt across a task battery (multi-sample) |
optimize_universal |
hill-climb one template (user or slot='system') |
judge_response, validate |
LLM judge a reply / measure the real success rate |
judge_selftest |
calibrate the grader on benign fixtures before trusting ASR |
http_request, barcode |
raw delivery / QR+barcode encoding |
finish, ask_operator |
stop the tool / pause for the operator |
Slash commands
/profile /target /provider /model /judge [model] endpoints & grader
/auto /autoexit /rounds autonomous loop
/objective /template /sysprompt /validate /replay campaign + reliability
/transforms /encode /diff /campaign /leaderboard arsenal, auto-sweep & benchmark
/find /tools /preset /lib /parsel /eni /harmbench search & libraries
/log /asr /stats /findings /repro /export /report logging, scoreboard, repro, CI export
Ctrl+S report · Ctrl+Y copy payload · Ctrl+T stats · Ctrl+R repro · Ctrl+L clear
Logging & reports
Every payload, reply, and verdict goes to sessions/run-<ts>.jsonl. /findings lists the
bypasses; /report [html] writes a markdown or styled-HTML findings doc; /repro [n]
copies a repro pack; /export dumps structured findings JSON; /session save|load
persists the whole engagement (history, objective, template, system prompt).
Headless / CI
Render reports and gate builds straight from a run log, no TUI. The log arg is optional:
omit it (or pass a directory) and the newest sessions/run-*.jsonl is used:
wallbreaker report # markdown for the latest run, to stdout
wallbreaker report --html --out report.html
wallbreaker export --out findings.json # structured findings JSON
wallbreaker export --fail-on-finding # exit 2 if any bypass -> fails CI
An opt-in live example lives at docs/examples/redteam-gate.yml. It is outside
.github/workflows so cloning the repository cannot schedule provider calls.
Test
pytest -q
Web dashboard
A browser dashboard ships alongside the TUI (FastAPI backend + React/Vite SPA). WebUI V2 uses the same capability catalog and application services as the TUI, and adds a server-owned execution queue, resumable event streams, persistent multi-turn composition, workflow sequencing, provider/profile management, and current or historical evidence inspection. Agent is dedicated to the autonomous Attack → Target → Judge loop; Live provides the holistic-to-granular observability surface.
More views: arsenal
pip install "wallbreaker-hermes[dashboard]" # FastAPI + uvicorn
wallbreaker dashboard # binds to 127.0.0.1:8787
The wheel includes the compiled dashboard. Contributors using an editable clone can rebuild it
with npm ci && npm run build under wallbreaker/dashboard/web.
Open WebUI V2 at http://127.0.0.1:8787/v2. The original dashboard remains available at
http://127.0.0.1:8787/legacy during the parity rollout. The backend reuses the same
engine as the TUI. For frontend hot-reload, run npm run dev in
wallbreaker/dashboard/web; it proxies /api to the dashboard backend.
See the setup guide for Windows instructions, provider configuration, history storage, development workflow, network-exposure safeguards, and troubleshooting. See Upstream maintenance for the update and rollback procedure.
Responsible use
Wallbreaker is for authorized LLM red-teaming and safety evaluation only: your own
models, or targets you have explicit permission to test. Run logs and generated
artifacts can contain harmful content; they're written to gitignored wb_runs/,
wb_artifacts/, and findings/. Dataset updates and configured provider calls use the
network; an optional [art] endpoint also receives scorecard labels and results. See
SECURITY.md for the full policy and vulnerability-reporting channel.
Contributing
Setup, architecture, and house rules are in CONTRIBUTING.md. Run
pytest -q before a PR; the suite is the contract.
License
AGPL-3.0-or-later. Wallbreaker is copyleft: any modified version (including one you run as a network/hosted service) must make its complete corresponding source available under the same license. Third-party corpora and benchmark rows are fetched only when requested and are not redistributed; see NOTICE and External data sources.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file wallbreaker_hermes-0.2.2.tar.gz.
File metadata
- Download URL: wallbreaker_hermes-0.2.2.tar.gz
- Upload date:
- Size: 1.9 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a2e224fa5d53effeeca027b62121408e5f83e8949220aea4fbec32b251f0292f
|
|
| MD5 |
1f445402576434b06a39351b90fca55d
|
|
| BLAKE2b-256 |
2e4a933f74c65b04a4a8cfa6cc573b457543b02af792bd00e9bd9786701c6ecf
|
Provenance
The following attestation bundles were made for wallbreaker_hermes-0.2.2.tar.gz:
Publisher:
release.yml on Yivas/wallbreaker-hermes
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
wallbreaker_hermes-0.2.2.tar.gz -
Subject digest:
a2e224fa5d53effeeca027b62121408e5f83e8949220aea4fbec32b251f0292f - Sigstore transparency entry: 2492520828
- Sigstore integration time:
-
Permalink:
Yivas/wallbreaker-hermes@ce38ef7e8e093b967f8c8aa5a5f602cb947acadf -
Branch / Tag:
refs/tags/v0.2.2 - Owner: https://github.com/Yivas
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@ce38ef7e8e093b967f8c8aa5a5f602cb947acadf -
Trigger Event:
push
-
Statement type:
File details
Details for the file wallbreaker_hermes-0.2.2-py3-none-any.whl.
File metadata
- Download URL: wallbreaker_hermes-0.2.2-py3-none-any.whl
- Upload date:
- Size: 760.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
933ec9a4759e829f25f7c527c0a5d703dbccb3e7833eaef15a534fcafb10087a
|
|
| MD5 |
6df1cecfc50c32f739cd0456a7cb4743
|
|
| BLAKE2b-256 |
2242dfce4bf75ec6bbf425f7439b7686415be89eb4f7b1fd2f5a0e301d322cdb
|
Provenance
The following attestation bundles were made for wallbreaker_hermes-0.2.2-py3-none-any.whl:
Publisher:
release.yml on Yivas/wallbreaker-hermes
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
wallbreaker_hermes-0.2.2-py3-none-any.whl -
Subject digest:
933ec9a4759e829f25f7c527c0a5d703dbccb3e7833eaef15a534fcafb10087a - Sigstore transparency entry: 2492521093
- Sigstore integration time:
-
Permalink:
Yivas/wallbreaker-hermes@ce38ef7e8e093b967f8c8aa5a5f602cb947acadf -
Branch / Tag:
refs/tags/v0.2.2 - Owner: https://github.com/Yivas
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@ce38ef7e8e093b967f8c8aa5a5f602cb947acadf -
Trigger Event:
push
-
Statement type: