Skip to main content
Avatar for Simon Sure from gravatar.com

Simon Sure

Username    simonissure
Date joined   Joined

42 projects

superred-target-chatbot

Last released

Chatbot target for the superred red-teaming framework: wraps any LLM via litellm

superred-optimizer-attack-anything

Last released

Attack Anything (SEATS) self-evolving attack-tree-search jailbreak optimizer for superred

superred-optimizer-many-shot

Last released

Many-Shot Jailbreak single-turn optimizer for superred

superred-optimizer-minja

Last released

MINJA memory-injection optimizer for superred agent targets

superred-optimizer-flip-attack

Last released

FlipAttack single-turn jailbreak optimizer for superred

superred-optimizer-fitd

Last released

Foot-in-the-Door multi-turn jailbreak optimizer for superred

superred-optimizer-eia-agent

Last released

Environmental Injection Attack optimizer for SuperRed web agents

superred-optimizer-code-chameleon

Last released

CodeChameleon jailbreak optimizer for superred

superred-optimizer-agentvigil-websentinel

Last released

AgentVigil/WebSentinel indirect prompt-injection optimizer for superred agent targets

superred-optimizer-libertas

Last released

Byte-faithful, model-aware replay of Pliny's L1B3RT4S jailbreak prompt corpus

superred-target-openclaw

Last released

OpenClaw agent target module for superred

superred-optimizer-poisonedrag

Last released

PoisonedRAG knowledge-corruption optimizer for superred

superred-optimizer-gptfuzzer

Last released

GPTFuzzer jailbreak optimizer for superred

superred-optimizer-dra

Last released

Disguise and Reconstruction Attack (DRA) optimizer for superred

superred-optimizer-chord-xthp

Last released

Chord/XTHP optimizer for superred agent tool-control-flow targets

superred-optimizer-tap

Last released

Tree of Attacks with Pruning (TAP) optimizer for superred

superred-optimizer-pair

Last released

PAIR jailbreak optimizer for superred

superred-optimizer-crescendo

Last released

Crescendo multi-turn jailbreak optimizer for superred

superred-optimizer-muzzle

Last released

MUZZLE adaptive agentic indirect-prompt-injection optimizer for superred (arXiv 2602.09222)

superred-optimizer-goal-passthrough

Last released

Goal-passthrough (unattacked direct-prompt) baseline optimizer for superred

superred-claim-strongreject

Last released

StrongREJECT (Souly et al., NeurIPS 2024) jailbreak benchmark as a superred SecurityClaim against ChatbotTarget

superred-claim-sorry-bench

Last released

SORRY-Bench safety-refusal benchmark as a superred SecurityClaim

superred-claim-harmbench

Last released

HarmBench standardized red-teaming benchmark as a superred SecurityClaim

superred-claim-dtap

Last released

DecodingTrust-Agent (DTAP-BENCH) security claim: one task per DTAP per-task config (benign or malicious, direct/indirect), driving either DTAP agent target through the byte-faithful out-of-band env-state judge. Target-agnostic; hierarchical factories by domain, threat model, and risk category.

superred-claim-agentharm

Last released

AgentHarm SecurityClaim for superred: faithful port of the harmful test_public split (176 behaviors) against the general inspect-agent target, reusing upstream inspect_evals.agentharm tools/grading/judges. 8 per-category subclaims + total.

superred-claim-asb

Last released

Agent Security Bench (ASB) security claim for the ASB target: one task per (agent, benign task, attacker tool), with ASB's attack-success / utility / refusal predicates ported verbatim and hierarchical factories by scenario, aggressiveness, and attack type.

superred-claim-agentdojo

Last released

AgentDojo SecurityClaim for superred: original injection-task port + bespoke system-purpose-violation goals against the composite AgentDojoTarget

superred-target-dtap-openclaw

Last released

OpenClaw concrete agent target for the DecodingTrust-Agent (DTAP) superred port: runs the OpenClaw CLI headless inside an isolated Docker container, wired to the shared dtap-scaffold base (forest, controllables, observables, env/MCP/injection lifecycle). Native exec/fs tools enabled (upstream disabled them).

superred-target-dtap-claudecode

Last released

Claude Code (claude_agent_sdk) DTAP agent target for superred: subclasses dtap_scaffold.DtapAgentTarget and runs the Claude Agent SDK in an isolated Docker container against the DTAP env MCP proxy, normalizing the in-container transcript into the shared TrajectoryArtifact.

superred-target-asb

Last released

Agent Security Bench (ASB) composite agent target for superred: runs ASB's real plan-then-execute agent loop (vendored, pinned to upstream 1f561dcc) with event-based injection at the four ASB surfaces (DPI/OPI/PoT/MP), routed through a litellm proxy.

superred-target-agentdojo

Last released

AgentDojo composite target for superred: union of all four AgentDojo suites with on-demand event-based injection and tool-catalog controllable surface. Tracks the latest released benchmark version (v1.2.2) via agentdojo_target.BENCHMARK_VERSION.

superred-target-inspect-agent

Last released

General inspect-ai tool-calling agent target for superred: runs an agent over caller-supplied tools/prompts/model and exposes the message trace. Benchmark-agnostic; reusable by any inspect-tool agentic SecurityClaim.

superred-target-dtap-scaffold

Last released

Shared scaffolding for the DecodingTrust-Agent (DTAP) superred targets: the agent-agnostic Target base, the security-domain forest, controllables/observables, env/MCP/Docker lifecycle, and the byte-faithful judge runner.

superred-optimizer-goat

Last released

GOAT (Generative Offensive Agent Tester) multi-turn jailbreak optimizer for superred

superred-optimizer-gepa-agentic

Last released

GEPA reflective prompt evolution optimizer for superred agent targets

superred-optimizer-gepa

Last released

GEPA reflective prompt evolution optimizer for superred

superred-optimizer-bijection

Last released

Bijection learning jailbreak optimizer for superred

superred-optimizer-autodan-turbo

Last released

AutoDAN-Turbo lifelong-strategy jailbreak optimizer for superred

superred

Last released

A modular framework for comprehensive red-teaming of AI systems

superred-claim-demo-secret-leak

Last released

Demo security claim for superred: a synthetic planted-secret scenario with a rigged trigger word. Not a real benchmark.

superred-optimizer-demo-prompt-list

Last released

Demo optimizer for superred: replays a fixed prompt list, ignores the goal, and does not learn. Not a real attack.

superred-target-minimal-llm-chat

Last released

Minimal single-turn LLM chat target for the superred red-teaming framework

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page