Skip to main content

Fusion First

Measure and guard AI agents. Fusion attacks an agent's system prompt with an OWASP-mapped suite, grades safety and quality with a cross-family judge whose accuracy is measured in the same run, and ships a runtime guardrail that blocks leaked secrets and tool calls the user's request does not cover (measured results below). Every number carries a confidence interval or an honesty badge. The prompt fix is optional and re-tested on your model (paired before/after, McNemar, honesty badge): on the small open-weight models measured it rarely cut attacks and raised refusals of safe requests; see the Trust Report.

Checks (OWASP LLM Top 10 2025 / Agentic 2026): direct_prompt_injection (LLM01), excessive_agency (LLM06), data_exfiltration (LLM02), system_prompt_leakage (LLM07).

Install

pip install fusion-safety            # core: runs, runtime guardrail, CLI, prompt fix
pip install "fusion-safety[mcp]"     # + MCP server (fusion-mcp)
pip install "fusion-safety[serve]"   # + local web app (fusion serve)

The wheel bundles the corpora (datasets/, cassettes/, crosswalk/, evals/), the web app and the Claude Code plugin marketplace (fusion plugin-dir prints its path).

Use

Any model or workflow can be the target:

Target Runs Cost
ollama:<model> a local model; any Hugging Face GGUF as ollama:hf.co/<user>/<repo> free
openai-compat:<url>#<model> your own server: vLLM, TGI, LM Studio, llama.cpp free if local
claude-cli:<model> Claude via the claude CLI subscription
openai: hf: anthropic: openrouter: together: groq: fireworks: mistral-api: deepseek: + <model> a hosted API with your key (OPENAI_API_KEY, HF_TOKEN, ...) metered; needs FUSION_ALLOW_API_SPEND=1
cmd:<command> any workflow: a program that reads a JSON request on stdin and prints the reply (template: /use) yours
python:<name> a Python function, via fusion_first.targets.FunctionModelClient yours

Grader: host (the calling agent), claude-cli, or any chat target above. Each run measures its grader on known-answer questions and withholds the grade ('?') if it falls short.

1. Claude Code plugin (recommended)

uv tool install fusion-safety        # puts fusion on PATH; the plugin starts its server with uvx
claude plugin marketplace add "$(fusion plugin-dir)" && claude plugin install fusion@fusion-first
# in Claude Code:  /fusion:audit prompts/support_bot.md ollama:llama3.2:1b

2. Command line / CI

fusion doctor                                                    # available backends
fusion run start --prompt agent.txt --target ollama:llama3.2:1b --grader claude-cli
fusion run start --prompt agent.txt --target ollama:llama3.2:1b  # grader=host: pauses for grading
fusion run start --prompt agent.txt --target openai-compat:http://127.0.0.1:8000/v1#Qwen/Qwen2.5-7B-Instruct --grader claude-cli  # vLLM / LM Studio / llama.cpp
fusion run start --prompt agent.txt --target hf:meta-llama/Llama-3.1-8B-Instruct --grader openai:gpt-4o  # hosted, your keys
fusion run start --prompt agent.txt --target "cmd:python fusion_adapter.py" --grader claude-cli  # any workflow
fusion run tasks > tasks.json;  fusion run submit --file answers.json
fusion run finalize --min-grade B                               # exit 1 below the bar
fusion run verify                                               # re-derive offline; exit 1 on drift
fusion harden --prompt agent.txt --write                        # optional: append the prompt fix
fusion run logs --file logs.jsonl --check everything --grader claude-cli  # grade existing transcripts
fusion import-promptfoo results.json                            # Wilson CIs + paired McNemar on promptfoo results
fusion guard-bench                              # runtime guard benchmark (in-house corpus: 100% recall / 0% over-block)

No local grader is recommended: ollama-prob:qwen2.5:7b scored below the policy floor in its pre-registered test (Trust Report).

MCP: fusion-mcp (stdio) exposes fusion_doctor, start_run, run_status, get_grading_tasks, submit_grades, finalize_run, verify_run, the one-shot audit_agent / scan_prompt, the runtime guard's guardrail_snippet / check_output / check_tool_call, and the optional harden_prompt.

{ "mcpServers": { "fusion": { "command": "fusion-mcp" } } }

3. Runtime guardrail (in your agent code; the fusion-guard plugin for Claude Code)

On fresh successful attacks against qwen2.5:7b and llama3.1:8b (rules 0c560b5, scored once, pre-registered), the guardrail stopped 89% of real attacks while wrongly blocking 1% of clean transcripts: every password leak and direct-harm hijack, 69% of data-stealing hijacks. Llama Guard 3 8B, configured, stopped 52% of the same attacks. The wrapper checks replies; each tool call is checked against the user's own request before it runs. Every live scan also replays the guard over its own replies (in-sample): the Guard step, the HTML report and fusion run finalize show what it would have stopped, and count separately the attacks answered in prose, where there is no tool call to check.

from fusion_first.guardrail.guard import Guardrail
from fusion_first.guardrail.policy import GuardConfig
from fusion_first.guardrail.client import GuardedModelClient

guard = Guardrail(GuardConfig(allowlisted_domains=["your-co.com"], secret_values=["sk-..."],
                              system_prompt=SYSTEM_PROMPT, require_authorization=True))  # the measured config
client = GuardedModelClient(your_model_client, guard)   # replies: secrets redacted, prompt dumps blocked
outcome = guard.guard_tool_call(tool_name, tool_args, user_request=user_message, untrusted_context=True)
if outcome.blocked: ...                                  # don't run it

4. Local web app

pip install "fusion-safety[serve]"
fusion serve                        # http://127.0.0.1:8765; --demo-only disables live models

Binds 127.0.0.1, accepts only 127.0.0.1/localhost Host headers, and requires a per-launch token on every request that runs anything, so other websites can't drive your local models or claude subscription.

Offline runs report DEMONSTRATION numbers (deterministic stand-in judge, canned responses); --live runs through Ollama and/or the claude CLI. Metered API keys are used only with --backend api and FUSION_ALLOW_API_SPEND=1. Full CI workflow: integrations/README.md.

Develop

python -m pytest -q                 # offline suite
python -m ruff check fusion tests scripts

fusion_first/ is the pure core (never imports modal); Modal/FastAPI wrappers live in app/. See CLAUDE.md (project guide) and SCHEMA.md (data model).

Metadata

Release files for fusion-safety 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for fusion-safety 0.1.0
File Size Uploaded
fusion_safety-0.1.0.tar.gz 509.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for fusion-safety 0.1.0
File Interpreter ABI Platform
fusion_safety-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.1 MB

Release files / fusion_safety-0.1.0.tar.gz

Download URL fusion_safety-0.1.0.tar.gz
Size 509.2 kB
Tags Source
SHA-256 checksum
How to use checksums
354e5e5f58e4dee21082992a89e991d3d56ded71b683184b6b3e09a3348b3cde
BLAKE2b-256 checksum
How to use checksums
830322adbab9de16251c5b69d0fbd4edcef1ff7fca8ff25fbf2a00a41bef562f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.2

Release files / fusion_safety-0.1.0-py3-none-any.whl

Download URL fusion_safety-0.1.0-py3-none-any.whl
Size 551.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
31de5921fee0f7567c63f78af39fdb69a537f1523d1441c10abfc0179f660781
BLAKE2b-256 checksum
How to use checksums
9925c6a3c767f88a5848929997569942f56eef980ef64dc85855cb17f46d10f5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.2

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page