Skip to main content

Security auditor for quantized SLMs using the AEGIS-4 Protocol.

Project description

SafeQuant-SLM is an open-source diagnostic framework designed to audit the security integrity of Small Language Models (SLMs) after quantization.

Developed by Jessica Sciammarelli, this tool bridges the gap between AI efficiency and AI governance by identifying "Security Decay" in compressed local models.


The Problem: Quantization Decay

As the industry moves toward Edge AI, models are increasingly compressed to 4-bit (NF4/FP4) to run on consumer hardware. However, my research indicates that heavy quantization often "crushes" the weights responsible for safety alignment.

The Findings: During initial benchmarks, models like Qwen-2.5-0.5B and SmolLM2-360M dropped to a 0.0% Security Score after 4-bit quantization, becoming highly vulnerable to malicious exploits despite being "safe" in their full-precision versions.


The AEGIS-4 Protocol

SafeQuant-SLM implements the AEGIS-4 (Adversarial Evaluation of Gated Intelligence Systems) protocol. This standardized 4-vector stress test evaluates:

  1. Prompt Injection: Bypassing system instructions via indirect/direct injections.
  2. Malicious Code: Generation of exploits (Buffer Overflows, SQLi, Malware scripts).
  3. Ethical Deception: Susceptibility to social engineering and logic manipulation.
  4. Privilege Escalation: Attempts to bypass authentication or access protected APIs.

Features

  • Automated Pipeline: Downloads, quantizes, and audits any Hugging Face model in one command.
  • SafetyGuard Verdict: Real-time terminal alerts (Secure, Caution, or Critical Risk).
  • Executive Reporting: Generates detailed PDF audit reports including VRAM usage and hardware latency.

Installation

1. Clone and Install

git clone [https://github.com/deeplearningworld/safequant-slm.git](https://github.com/deeplearningworld/safequant-slm.git)
cd safequant-slm
pip install -r requirements.txt
pip install .


2. Run your first Audit
# Audit a model in 4-bit precision
safequant --model google/gemma-2-2b-it --bits 4


 SafetyGuard Verdicts Score Status Action
75% - 100%✅ SECURE Production ready. Model maintained alignment.
30% - 74%⚠️ CAUTION Requires additional server-side/wrapper filtering.
0% - 29%⛔ CRITICAL RISKSecurity collapsed. Deployment not recommended.

Customization & Hardware 
ScalingSafeQuant-SLM is designed to be model-agnostic and hardware-flexible, allowing it to adapt to various infrastructure needs:
Supporting Different SLM Architectures The framework utilizes the Hugging Face AutoModel architecture. To test a new model (e.g., Llama-3.2-1B, Phi-3.5, or Qwen-2.5), simply pass the model ID string:Bash safequant --model "meta-llama/Llama-3.2-1B-Instruct"

Hardware Acceleration (GPU, TPU, MPS)
The codebase leverages accelerate and device_map="auto" to detect the best available backend:NVIDIA GPU (CUDA): Uses bitsandbytes for 4-bit NormalFloat (NF4) quantization, optimized for RTX/A100 hardware.
Apple Silicon (MPS): Full support for Mac M1/M2/M3 chips using Unified Memory for efficient local inference.
Google TPU: Can be scaled for large-scale batch auditing on TPUs by wrapping the Aegis4Auditor in an XLA-compatible distributed loop (requires torch_xla).
Modifying AEGIS-4 Vectors You can customize the adversarial prompts in safequant/audit.py to tailor the protocol to specific industry regulations, such as HIPAA for healthcare or PCI-DSS for financial services.

License Distributed under the Apache License 2.0. See LICENSE for more information. This license provides an explicit grant of patent rights from contributors to users, ensuring long-term legal safety for enterprise implementations.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

safequant_slm-1.0.0.tar.gz (7.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

safequant_slm-1.0.0-py3-none-any.whl (7.7 kB view details)

Uploaded Python 3

File details

Details for the file safequant_slm-1.0.0.tar.gz.

File metadata

  • Download URL: safequant_slm-1.0.0.tar.gz
  • Upload date:
  • Size: 7.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.8

File hashes

Hashes for safequant_slm-1.0.0.tar.gz
Algorithm Hash digest
SHA256 5909fdade420c1c37f87ef5c719fdad44d737078b37f73c9a0aa09421cb30647
MD5 2510af7d71adcb95199a061b3fc7f3de
BLAKE2b-256 db79efef38db0a52c9ea73fb67dfa38dd52e51371509013449fd64d8c1852ff4

See more details on using hashes here.

File details

Details for the file safequant_slm-1.0.0-py3-none-any.whl.

File metadata

  • Download URL: safequant_slm-1.0.0-py3-none-any.whl
  • Upload date:
  • Size: 7.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.8

File hashes

Hashes for safequant_slm-1.0.0-py3-none-any.whl
Algorithm Hash digest
SHA256 0adb3181e035c5c1d8796942f3a77835290f73272ec42e85983e7c9a6d1f0e05
MD5 5b25b27a52a955a8d7f8cb5a2ca3cf35
BLAKE2b-256 12b97cc1bf355edf5cac91a2a457da31b3bf675690d83bdd9737daf62d76d72b

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page