AI Adversarial Security Testing for LLM, RAG, Agent, and Vision Models
Project description
RednBlue CLI v3.1.0
Zero-Knowledge Adversarial Security Testing for AI Models
RednBlue CLI is a command-line tool for testing the adversarial robustness of machine learning models. Run security assessments locally — your model never leaves your infrastructure.
███████████ ███████████
▒▒███▒▒▒▒▒███ ▒▒███▒▒▒▒▒███
▒███ ▒███ ████████ ▒███ ▒███
▒██████████ ▒▒███ ▒▒███ ▒██████████
▒███▒▒▒▒▒███ ▒███ ▒███ ▒███▒▒▒▒▒███
▒███ ▒███ ▒███ ▒███ ▒███ ▒███
█████ █████ ████ █████ ███████████
▒▒▒▒▒ ▒▒▒▒▒ ▒▒▒▒ ▒▒▒▒▒ ▒▒▒▒▒▒▒▒▒▒▒
Zero-Knowledge Adversarial Security Testing
Features
- Zero-Knowledge Protocol — Model weights and data never leave your infrastructure
- Image Classifiers — Test ResNet, VGG, EfficientNet, and custom architectures
- YOLO Detection — Full support for YOLOv5, YOLOv8, YOLOv10, YOLOv11
- Tier-Based Testing — Freelancer (quick scan) and Enterprise (comprehensive)
- Visual Explanations (XAI) — See exactly which pixels confused the model, including the real perturbed image, not a placeholder
- Blue Team Hardening —
rnb defendapplies a local defensive layer to a vulnerable model - Encrypted Submission — AES-256 encrypted results with HMAC-SHA256 signing
- Signed Compliance Reports — Map test results onto EU AI Act, NIST AI RMF, and ISO/IEC 42001 articles, with a tamper-evident trust seal
- Multi-Jurisdiction Compliance — EU AI Act, NIST AI RMF, ISO/IEC 42001, UK DSIT, Canada AIDA, Singapore MAIGF
Installation
# Install from PyPI
pip install rednblue
# Verify installation
rnb
Requirements
- Python 3.8+
- PyTorch 2.0+
- CUDA (optional, for GPU acceleration)
Quick Start
1. Choose your interface
Desktop UI (recommended for non-CLI users):
rnb ui
Command line: see below.
2. Set your token (only needed for certified/submitted runs)
# Linux/Mac
export RNB_TOKEN=RB-XXXXXX-YYYYYY
# Windows
set RNB_TOKEN=RB-XXXXXX-YYYYYY
3. Run a free preview
Image Classifier:
rnb preview --model resnet50.pth --input ./test_images --model-type classifier
YOLO Detection Model:
rnb preview --model yolov10n.pt --input ./test_images --model-type yolo
4. Submit for certification
rnb preview --model yolov10n.pt --input ./images --model-type yolo --submit
5. View your report
Go to: https://dashboard.rednblue.io/dashboard/tests to download your PDF report and certificate.
Commands
| Command | Description |
|---|---|
rnb |
Show welcome banner and quick start |
rnb-welcome |
Same as rnb — convenience alias usable outside the main group |
rnb ui |
Launch the desktop graphical interface |
rnb preview --help |
Run adversarial attacks (classifier or YOLO) |
rnb status |
Check CLI version and token validity/tier |
rnb explain --help |
Generate a standalone visual-explanation panel for one image (Robust Integrated Gradients) |
rnb defend --help |
Apply local defensive hardening to a model (Blue Team) |
rnb report --help |
Generate a signed compliance report + trust seal from saved results |
rnb optimize-epsilon |
Optimize epsilon values (Enterprise) |
rnb llm-test |
Test LLM models (Enterprise) |
rnb llm-scan |
Recon-only LLM pipeline scan |
Supported Model Formats
| Type | Extensions |
|---|---|
| PyTorch | .pt, .pth, .safetensors |
| ONNX | .onnx |
| Ultralytics YOLO | .pt (auto-detected) |
Architectures: ResNet, VGG, DenseNet, EfficientNet, MobileNet, Inception, SqueezeNet, ShuffleNet, GoogLeNet, AlexNet, and common timm backbones.
TensorFlow (
.h5,.pb) is not yet supported. The ingestion layer (rnb/core/ingestion/wrapper.py) is architected to accept a TensorFlow adapter, buttensorflowis intentionally not bundled as a dependency today — adding real support is tracked as a follow-up, not claimed here.
Assessment Dimensions
Classifier Models
| Dimension | Description |
|---|---|
| Noise Resilience | Stability under sensor noise and interference |
| Spatial Consistency | Robustness to spatial feature shifts |
| Universal Pattern Defense | Resistance to universal perturbation patterns |
| Feature Stability | Internal representation integrity |
| Confidence Calibration | Prediction reliability accuracy |
| Iterative Stress Tolerance | Defense against sustained pressure |
| Optimization Attack Defense | Resistance to optimized adversarial inputs |
| Deep Perturbation Resistance | Resilience against deep layer perturbations |
YOLO Detection Models
| Dimension | Description |
|---|---|
| Noise Resilience | Stability under sensor noise |
| Input Perturbation Defense | Resistance to subtle input modifications |
| Iterative Stress Tolerance | Defense against multi-step attacks |
| Detection Consistency | Reliable detection under varying conditions |
| Targeted Evasion Defense | Resistance to deliberate misclassification |
| Object Persistence | Maintains detections under perturbations |
| Multi-Object Stability | Accuracy in crowded scenes |
| Black-Box Resilience | Defense without model access |
| Query-Limited Defense | Resistance to low-query probing |
Tier Comparison
| Feature | Freelancer | Enterprise |
|---|---|---|
| Classifier Attacks | 5 | 8 |
| YOLO Attacks | 4 | 9 |
| Epsilon Values | 2 | 4 |
| Total Scenarios | ~10-20 | ~30-70 |
| LLM Testing | ❌ | ✅ |
| Epsilon Optimization | ❌ | ✅ |
| Continuous Integrity Monitoring | ❌ | 🔜 Coming soon |
Output Example
============================================================
RednBlue Security Preview — YOLO Detection
============================================================
Attacks run : 21
Successful hits: 0/21 (0%)
Robustness rate: 100%
Estimated Grade: GOLD
⚠️ This is a preview only
→ Visit: https://rednblue.io/checkout
→ Re-run with: rnb preview --model-type yolo --submit
Certification Grades
| Grade | Score | Meaning |
|---|---|---|
| 🥇 GOLD | ≥90% | Excellent robustness, deployment ready |
| 🥈 SILVER | ≥75% | Good robustness, minor improvements recommended |
| 🥉 BRONZE | ≥50% | Moderate robustness, improvements needed |
Architecture
For the complete, AI-readable source dump of this repository, see
rnb.md.
Links
- Platform: https://dashboard.rednblue.io
- Documentation: https://docs.rednblue.ai
- Website: https://rednblue.io
Authors
- Dr. Mahdi Deramgozin — Chief AI Officer
- Dr. Saeid Samizade — Chief Technology Officer
License
Proprietary — RednBlue SAS © 2026
Made in France 🇫🇷
Changelog
[3.3.0] — 2026-07-02
Fixed
- Critical:
--tiermanual override onrnb llm/rnb previewbypassedRNB_TOKENvalidation entirely, allowing unlimited unpaid runs. Token is now always required and validated. - Critical: API/transport errors (e.g. invalid API key) were scored as successful jailbreaks. Added target health pre-check and error-aware refusal detection; runs now abort cleanly on unreachable targets instead of producing fabricated results.
- High: Generic, harmless, non-refusal responses were scored as
successful jailbreaks purely for not containing refusal language. Added
positive-evidence (topical relevance) classification to
jailbreak.pyandprompt_injection.py's Token Fishing. - Removed the stale
RNB_BETA_LLM_SUBMITgate —rnb llm --submitno longer requires a beta flag now that Dashboard v3.1 supports LLM sessions. - Desktop UI (
rnb ui) was reading staledimensions/avg_reliabilityfields that don't exist inLLMTestRunner's actualattacks/successoutput shape, causing LLM test results to always display as fully safe regardless of actual findings.
Changed
--tieronrnb llm/rnb previewnow only warns on mismatch instead of silently overriding the token's actual tier.
[3.1.0] — 2026-06
Fixed
- "After Attack" panel no longer mirrors the original image. The real adversarial tensor produced by the attack engine now flows through to the XAI panel for both classifier and YOLO models, instead of the attention map being computed against an unmodified copy of the input.
Added
- Unified model-ingestion layer (
rnb/core/ingestion/wrapper.py) — aBaseModelWrapperabstraction with a PyTorch adapter (covering.pt/.pth/.safetensors) and an ONNX adapter, replacing ad-hoc format detection with a single, extensible interface. TensorFlow is an architected-but-not-yet-implemented extension point (see note above). rnb defend— Local Blue Team hardening. Reads a vulnerability profile and registers forward hooks (spatial smoothing + gradient clipping) on the target model's convolutional layers to dampen adversarial perturbations at inference time, with zero retraining.rnb explain— Standalone XAI playground for a single (model, image) pair. Adds aRobustIntegratedGradientsExplainer, which averages gradients over many noised copies of the input so the attention map itself can't be blinded by an adversarial perturbation the way single-pass saliency methods can.rnb report— Generates a signed compliance report from savedrnb preview --save-resultsoutput: maps Noise Resilience, Jailbreak Bypass Rate, and Bias Score-style metrics onto EU AI Act and NIST AI RMF articles, then emits a Markdown report plus a signedtrust_seal.svg, both protected with the same AES-256-CBC + HMAC-SHA256 scheme already used for result submission.--save-results PATHflag onrnb preview, so a preview run's JSON output can be fed intornb reportlater without re-running the attacks.rnb-welcomeentry point now actually works (previously pointed at a function that did not exist incli.py).
Changed
rnb/banner.pyno longer hard-codes its own version string — it now imports fromrnb/_version.py, the single source of truth.- Repo cleanup: removed
BANNER_INTEGRATION.md,patch_banner.py, the stray root-levelbanner.py,rnb_init.py,rnb/architecture.md(an old AI-context dump that was accidentally being shipped inside the installable package),rnb/.installed, andtest.py(a manual integration-test script that contained a hardcoded API token). - Documentation consolidated:
QUICKSTART.mdandCHANGELOG.mdhave been folded into this README. The only other Markdown file in this repository isrnb.md, the AI-readable full-source dump.
Deferred
- Continuous Integrity Monitoring Daemon (
rnb daemon) is not included in this release. It is reserved for a future Enterprise-only release.
[3.0.0] — 2026-05-08
Major release. Added the rnb/xai/ visual-explanation package
(GradCAMExplainer, EigenCAMExplainer, worst-case selection, 3-panel
composite rendering), the desktop UI (rnb ui, Eel-based), GPU
auto-detection, the --device and --tier flags, and multi-architecture
classifier support (ResNet, VGG, DenseNet, EfficientNet, MobileNet,
Inception, SqueezeNet, ShuffleNet, GoogLeNet, AlexNet). Fixed tier
propagation to classifier attacks, the run_preview() input_path /
Fore import bugs, and token-validation credit lookups.
[2.5.1] — 2026-04-14
Added HuggingFace whitelist filtering and timm backbone support.
Migrated the build system from setup.py to pyproject.toml.
[2.5.0] — 2026-04 (initial public release)
Welcome banner, rnb status, rnb preview, rnb optimize-epsilon,
rnb test-llm. Eight image-classifier attacks (GNI, SHFP, UAP, FSP, CCM, PGD, CW, DEEP). AES-256-CBC + HMAC-SHA256 encrypted result
submission to the dashboard.
Full per-commit history is available via
git log. This section retains only release-level highlights.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file rednblue-3.3.0.tar.gz.
File metadata
- Download URL: rednblue-3.3.0.tar.gz
- Upload date:
- Size: 291.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.10.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
42162ec6b922724d648570a67ecf1e58e83bbc0292278fc8c7c8aaa8bf91751d
|
|
| MD5 |
fa3c1b6956412e4d46f021f5a3c3b94b
|
|
| BLAKE2b-256 |
9b7f1cf5bd6f8a5b12fd13798a9678459fde1858c13ef5a447977ead67532e03
|
File details
Details for the file rednblue-3.3.0-py3-none-any.whl.
File metadata
- Download URL: rednblue-3.3.0-py3-none-any.whl
- Upload date:
- Size: 304.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.10.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a9c6e10e7e0be68c5f8c4db32a58a97b8d2b0338f089921bae1466a64fe4bfc5
|
|
| MD5 |
08898613da0bd619b17dfeee0a274b06
|
|
| BLAKE2b-256 |
f018b6e81f7bcc8918f4cf4e1023cbcd2dd1e7e07237def2e73bfd22c7ade125
|