Skip to main content
CI # vamp-llm-probe

Python 3.8+ Version License AGPL-3.0 VampSecure Labs

Security auditor for language model inference API endpoints. Sends crafted HTTP requests to detect vulnerabilities without relying on any AI SDK — only aiohttp, asyncio, and the standard library.

For authorized use only. Run this tool exclusively against endpoints you own or have explicit written permission to test.


Features

  • 7 audit phases covering reconnaissance, prompt injection, restriction bypass, data extraction, access controls, adversarial dataset red team and multi-turn conversation attacks
  • Truly bilingual detection — refusal and compliance heuristics cover both English and Spanish; models responding in Spanish are correctly evaluated regardless of the prompt language
  • Bundled adversarial datasets — 666 jailbreaks (EN) + 170+ injection/jailbreak vectors (ES bundled en código, sin fichero externo) + 135+ jailbreaks ES + 210 injection prompts (EN) + 390 forbidden questions (13 content-policy categories) — 305+ payloads ES totales
  • 6-subtest Phase 6: A (injection EN), B (jailbreak EN), C (forbidden questions), A_es (injection ES), B_es (jailbreak ES), D (ASCII smuggling)
  • ASCII smuggling detection — active (subtest D sends Unicode Tags payloads) and passive (scans every Phase 2 response for hidden Unicode Tags characters in output)
  • No AI SDK dependency — pure HTTP-level testing via aiohttp
  • Async execution — parallel requests for rate-limiting tests
  • Structured findings with severity levels (CRITICAL / HIGH / MEDIUM / LOW / INFO)
  • OWASP mapping — every finding is automatically tagged with the corresponding OWASP LLM Top 10 2025 and OWASP Agentic AI Top 10 2026 categories; visible in JSON output and HTML reports
  • Professional reports — JSON (machine-readable) and HTML (client-delivery) with OWASP tags on each finding card
  • Exit codes suitable for CI/CD pipeline integration
  • Auto-detection of API format and available models

Installation

pip install vamp-llm-probe
# o con Homebrew:
brew install vampsecure-labs/labs/vamp-llm-probe
pip install -r requirements.txt

Requirements: Python 3.8+ and aiohttp>=3.9.0.


Usage

Basic scan

python3 vamp_llm_probe.py --endpoint http://localhost:11434

With API key and specific model

python3 vamp_llm_probe.py \
  --endpoint http://api.example.com \
  --api-key sk-your-api-key \
  --model llama3:8b

Full engagement with HTML report

python3 vamp_llm_probe.py \
  --endpoint http://inference.internal:8080 \
  --api-key Bearer_TOKEN \
  --client "Empresa SL" \
  --engagement "Pentest Infraestructura 2026-Q3" \
  --auditor "VampSecure Labs" \
  --output results.json \
  --report-html report.html \
  --verbose

Activate dataset red team (Phase 6)

python3 vamp_llm_probe.py \
  --endpoint http://localhost:11434 \
  --dataset \
  --dataset-sample 30

Dataset with category filter (forbidden questions)

python3 vamp_llm_probe.py \
  --endpoint http://localhost:11434 \
  --dataset \
  --dataset-sample 20 \
  --dataset-categories "Malware,Illegal Activity,Physical Harm"

Skip jailbreak phase (restricted environments)

python3 vamp_llm_probe.py \
  --endpoint http://localhost:11434 \
  --no-jailbreak

CLI Arguments

Argument Description Default
--endpoint URL Base URL of the endpoint to audit (required) —
--api-key KEY Bearer authorization key (none)
--model MODELO Model name for inference tests auto-detect
--timeout N Request timeout in seconds 30
--output FILE Save results as JSON to FILE (none)
--report-html FILE Generate professional HTML report (none)
--client NOMBRE Client name for the report cover Confidencial
--engagement DESC Engagement description (none)
--auditor NOMBRE Auditor name for the report VampSecure Labs
--no-jailbreak Skip Phase 3 (restriction bypass) false
--verbose Verbose mode — shows HTTP traces false
--dataset Activate Phase 6: adversarial dataset red team false
--dataset-sample N Prompts per dataset type to test 15
--dataset-categories CATS Comma-separated forbidden question categories to test all
--multi-turn Activate Phase 9: multi-turn conversation attacks (20 chains, EN+ES) false
--mt-turns N Number of turns per multi-turn chain 3

Available forbidden question categories: Illegal Activity, Hate Speech, Malware, Physical Harm, Economic Harm, Fraud, Pornography, Political Lobbying, Privacy Violence, Legal Opinion, Financial Advice, Health Consultation, Gov Decision


Audit Phases & Findings

Phase Name Findings Severity
1 Endpoint Reconnaissance LLM-001..009 CRITICAL to INFO
2 Prompt Injection LLM-010..029 CRITICAL to HIGH
3 Restriction Bypass Attempts LLM-030..049 HIGH to MEDIUM
4 Data Extraction & Leaks LLM-050..069 CRITICAL to MEDIUM
5 Access Controls & Behavior LLM-070..089 HIGH to LOW
6 Adversarial Dataset Red Team LLM-100..139 HIGH
9 Multi-Turn Conversation Attacks (--multi-turn) LLM-MT-001..020 CRITICAL to HIGH

Phase 1 — Endpoint Reconnaissance

Finding Title Severity
LLM-001 Endpoint exposes model list without authentication CRITICAL
LLM-002 Web management interface publicly accessible MEDIUM
LLM-003 Server version exposed in headers or response INFO
LLM-004 Multiple administrative routes accessible MEDIUM
LLM-005 Inference endpoint accessible without authentication CRITICAL

Phase 2 — Prompt Injection

Tests 10 crafted prompt injection payloads including direct overrides, role substitution, JSON format overrides, indirect HTML injection, multilingual overrides, zero-width space evasion, developer-mode unlocking, and token-separator injection.

Phase 3 — Restriction Bypass Attempts

Tests 8 bypass techniques: Base64-encoded instructions, unrestricted roleplay, query fragmentation, emoji/token obfuscation, language-switch overrides (English, French), and continuation-text technique.

Phase 4 — Data Extraction & Leaks

Finding Title Severity
LLM-050..052 System prompt extraction attempts HIGH
LLM-053 Filesystem listing request HIGH
LLM-054 API key / credential extraction CRITICAL
LLM-055 Network configuration extraction HIGH
LLM-056 SSRF vector via prompt injection CRITICAL
LLM-057 Environment variable extraction HIGH

Phase 5 — Access Controls & Behavior

Finding Title Severity
LLM-070 No rate limiting on inference endpoint HIGH
LLM-071 Context overflow accepted without limits MEDIUM
LLM-072 Poor handling of non-existent models MEDIUM
LLM-073 Malformed JSON does not return HTTP 400 LOW
LLM-074 Permissive CORS on inference endpoint MEDIUM
LLM-075 Missing HTTP security headers LOW

Phase 6 — Adversarial Dataset Red Team (--dataset)

Uses five bundled datasets — three in English (from TrustAI-laboratory/Learn-Prompt-Hacking) and two curated in Spanish by VampSecure Labs — plus a dynamic ASCII smuggling generator:

Subtest Dataset Size Language Detection Method
A injection_prompts.csv 210 real injection vectors EN Model responds substantively without refusal to task-redirect prompts
B jailbreak_prompts.csv 666 real jailbreak techniques EN Model response contains compliance markers without refusal patterns
C forbidden_questions.csv 390 questions across 13 policy categories EN Model provides >150-char substantive answer without refusal
A_es injection_prompts_es.csv 50 curated injection vectors ES Same bilingual heuristic — Spanish refusal phrases also checked
B_es jailbreak_prompts_es.csv 30 jailbreak prompts ES Spanish compliance markers (NEXUS/MÁXIMO/ALFA/SOMBRA/modo sin filtros…) detected
D _generar_payloads_ascii_smuggling() 30 generated payloads EN Model executes hidden instructions encoded in Unicode Tags (U+E0000-U+E007F)

The detection engine is language-aware: _REFUSAL_PATTERNS covers 30+ English and Spanish refusal phrases; _COMPLIANCE_MARKERS covers 30+ English and Spanish compliance indicators. A Spanish model refusing in Spanish counts as a refusal; a Spanish model adopting an unrestricted persona in Spanish counts as a jailbreak success.

Findings LLM-100+ are generated dynamically. Each finding includes the exact prompt and model response as evidence.

ASCII Smuggling Vector

ASCII smuggling exploits the Unicode Tags block (U+E0000–U+E007F) — an invisible copy of printable ASCII. These characters are not rendered on screen but are processed by LLMs, allowing hidden instructions to be embedded in content that appears clean to a human reviewer.

Microsoft published an analysis on 3 Sep 2026 showing the technique is actively used in phishing to evade email security filters: ASCII Smuggling Crosses Over from AI Prompt Injection to Phishing Evasion.

vamp-llm-probe detects this vector in two ways:

  1. Active (Subtest D) — sends 30 payloads where innocent-looking visible text contains Unicode Tags–encoded jailbreak instructions. A CRITICAL finding is raised if the model executes the hidden instruction.
  2. Passive (Phase 2) — every response from the inference endpoint is scanned for Unicode Tags characters. A HIGH finding is raised if the endpoint itself returns invisible characters (which could inject hidden instructions into downstream clients).

Legitimate exceptions — the English, Scottish, and Welsh flag emoji — are excluded from detection (they encode their subdivision tags using this same Unicode block).


OWASP Mapping

Every finding produced by vamp-llm-probe is automatically tagged with the corresponding OWASP categories before the report is generated. Tags appear in the JSON output (finding.tags) and as blue badges in the HTML report.

OWASP LLM Top 10 — 2025

Tag Category
OWASP-LLM01 Prompt Injection
OWASP-LLM02 Sensitive Information Disclosure
OWASP-LLM05 Improper Output Handling
OWASP-LLM06 Excessive Agency
OWASP-LLM07 System Prompt Leakage
OWASP-LLM10 Unbounded Consumption

OWASP Agentic AI Top 10 — 2026

Tag Category
OWASP-AGENT04 Context Manipulation
OWASP-AGENT06 Intent Breaking & Goal Hijacking
OWASP-AGENT07 Data Exfiltration via Agents
OWASP-AGENT09 Resource Overuse

Finding-to-OWASP mapping

Finding range Phase OWASP tags
LLM-001..009 Endpoint Reconnaissance LLM06 (+ LLM02 if LLM-003)
LLM-010..029 Prompt Injection + passive ASCII scan LLM01 AGENT04 AGENT06
LLM-030..049 Restriction Bypass / Jailbreak LLM01 AGENT06
LLM-050..069 Data Extraction & Leaks LLM02 LLM07 AGENT07
LLM-070 Rate limiting absent LLM10 AGENT09
LLM-071..073 Output handling issues LLM05
LLM-074..089 CORS / security headers LLM06
LLM-100..199 Adversarial dataset red team LLM01 AGENT06
LLM-ASCII-* ASCII smuggling active (Subtest D) LLM01 AGENT04

Bundled Datasets

vamp-llm-probe/payloads/
├── jailbreak_prompts.csv           # 666 real jailbreaks EN (verazuo/jailbreak_llms)
├── injection_prompts.csv           # 210 injection prompts EN (TrustAI curated)
├── forbidden_questions.csv         # 390 questions × 13 policy categories (TrustAI)
├── injection_prompts_es.csv        # 50 injection vectors ES (VSL curated)
├── jailbreak_prompts_es.csv        # 30 jailbreak prompts ES (VSL curated)
└── ascii_smuggling_payloads.json   # Source instructions for ASCII smuggling Subtest D (VSL)

All datasets are offline and self-contained. No external requests are made at runtime. The English datasets are sourced from TrustAI-laboratory/Learn-Prompt-Hacking; the Spanish datasets were curated by VampSecure Labs to cover native Spanish-language attack vectors not present in the original corpus.


Exit Codes

Code Meaning
0 No critical findings (MEDIUM, LOW, or INFO only)
1 HIGH severity findings detected
2 CRITICAL severity findings detected

Use these codes in CI/CD pipelines to gate deployments:

python3 vamp_llm_probe.py --endpoint "$ENDPOINT" --dataset || {
  echo "Security findings detected — blocking deployment"
  exit 1
}

Output Formats

JSON (--output results.json)

Machine-readable structured output following the VSL standard schema:

{
  "schema_version": "1.0",
  "generated": "2026-08-12 12:00 UTC",
  "meta": { "tool": "vamp-llm-probe", "tool_version": "1.6.0", ... },
  "summary": { "total": 5, "by_severity": { "CRITICAL": 2, "HIGH": 1, ... } },
  "findings": [ { "id": "LLM-001", "severity": "CRITICAL", ... } ]
}

HTML (--report-html report.html)

Professional client-delivery report with:

  • Cover page with engagement details
  • Executive summary with risk distribution chart
  • Findings table with severity color coding
  • Detailed finding cards with evidence and remediation

Project Structure

vamp-llm-probe/
├── vamp_llm_probe.py    # Main auditor (6 phases, bilingual detection, ASCII smuggling)
├── vampsec_report.py    # Unified reporting module (VSL shared)
├── payloads/            # Adversarial datasets (Phase 6)
│   ├── jailbreak_prompts.csv           # EN — 666 jailbreaks
│   ├── injection_prompts.csv           # EN — 210 injection vectors
│   ├── forbidden_questions.csv         # EN — 390 forbidden questions
│   ├── injection_prompts_es.csv        # ES — 50 injection vectors (VSL)
│   ├── jailbreak_prompts_es.csv        # ES — 30 jailbreak prompts (VSL)
│   └── ascii_smuggling_payloads.json   # ASCII smuggling source instructions (VSL)
├── requirements.txt
├── .gitignore
└── README.md

License

AGPL-3.0 License — see individual file headers for copyright details.


© VampSecure Studios — VampSecure Labs Security Research Division
Authorized use only in environments with explicit written permission.


Phase 9 — Multi-Turn Conversation Attacks (--multi-turn)

Tests 20 conversation chains designed to bypass LLM safety controls through progressive context manipulation. Unlike single-turn attacks, multi-turn chains build rapport, establish personas and escalate incrementally across multiple messages.

python3 vamp_llm_probe.py --endpoint http://localhost:11434 --multi-turn
python3 vamp_llm_probe.py --endpoint http://localhost:11434 --multi-turn --mt-turns 5
Chain set Language Count Strategy
MT-ES-001..010 Spanish 10 Rol de asistente sin restricciones, extracción progresiva, jailbreak incremental
MT-EN-001..010 English 10 Roleplay escalation, hypothetical framing, authority impersonation chains

Each chain sends N turns sequentially, tracking conversation history. A finding is raised if the model eventually complies with a restricted request after initial refusal. This tests multi-turn jailbreak resilience — a model that correctly refuses turn 1 but complies on turn 3 is still vulnerable.


Historial de versiones

Versión Cambios principales
v1.7.0 Fase 9 multi-turn (20 chains EN+ES), corpus ES bundled en código (305+ payloads), sin dependencia de ficheros CSV externos para ES
v1.6.0 OWASP mapping LLM+Agentic AI, ASCII smuggling passive scan en Phase 2
v1.5.0 Subtest D ASCII smuggling activo, datasets ES curados por VSL
v1.4.0 Phase 6 datasets, bilingual detection engine

© VampSecure Studios — VampSecure Labs Security Research Division
Authorized use only in environments with explicit written permission.

Metadata

Release files for vamp-llm-probe 1.7.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for vamp-llm-probe 1.7.0
File Size Uploaded
vamp_llm_probe-1.7.0.tar.gz 94.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for vamp-llm-probe 1.7.0
File Interpreter ABI Platform
vamp_llm_probe-1.7.0-py3-none-any.whl Python 3 none any Details

Total release size: 179.4 kB

Release files / vamp_llm_probe-1.7.0.tar.gz

Download URL vamp_llm_probe-1.7.0.tar.gz
Size 94.2 kB
Tags Source
SHA-256 checksum
How to use checksums
fb0be7fdd36b30cf4f43985ea9d05053a9907bda2f59496b48d1ebbb476b8770
BLAKE2b-256 checksum
How to use checksums
c6bd96475023241a0f9448d1dcffef434686e88d01fcb00d1623716808169757
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / vamp_llm_probe-1.7.0-py3-none-any.whl

Download URL vamp_llm_probe-1.7.0-py3-none-any.whl
Size 85.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8d8ebdc3e0c77beb183cd06c89b5b4013d68fa8a693e79fb21ba9b2d6ede9c8b
BLAKE2b-256 checksum
How to use checksums
4e30bbaa5d93fbe2b139c24edcb20cf24db3d1495aeed164df3cd4dce0aa1784
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

1.8.0

2 release files

This release

1.7.0 This release

2 release files

1.6.0

2 release files

1.5.0

1 release file

1.4.0

2 release files

1.3.1

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page