Skip to main content
Anti-PLUTO Logo

Anti-PLUTO

Anti-Phishing Lexical Utilities and Threat Observation

A unified framework for email dataset preprocessing, PII masking, synthetic data generation, and semantic-stylometric phishing classification.

License: CC BY-NC 4.0 Python 3.10+


Overview

With the rapid proliferation of Large Language Models (LLMs) enabling threat actors to generate contextually fluent spear-phishing emails that evade traditional rule-based filters (like SpamAssassin and Rspamd), defensive systems must adapt. Anti-PLUTO addresses this emerging threat by utilizing a state-of-the-art dual-branch feature fusion architecture.

Key Components & Architecture

1. Dual-Branch Feature Fusion

  • High-Dimensional Stylometric Branch (100,000 dimensions): Extracts joint word- and character-level $n$-gram TF-IDF representations. It preserves functional stop-words, capitalization patterns, and punctuation marks to isolate the subtle, subconscious stylistic fingerprints of human writers vs. LLM token sampling distributions.
  • Deep Contextual Semantic Branch (768 dimensions): Utilizes a dense continuous vector manifold generated by fine-tuning DeBERTa-v3-small (Decoding-enhanced BERT with Disentangled Attention). It extracts contextual intent, emotional manipulation vectors, and social engineering semantics.

2. Unified 4-Cohort Classification

Anti-PLUTO is designed to effectively differentiate between four email cohorts:

  • Human Benign (HB): Legitimate, human-authored corporate correspondence.
  • Human Phishing (HP): Malicious, human-authored attacks (e.g., credential harvesting, advance-fee fraud).
  • LLM Benign (LB): Legitimate corporate communications written or polished by generative writing assistants.
  • LLM Phishing (LP): Malicious communications synthesized or polished by generative AI to execute social engineering.

3. Privacy-Preserving Lexical Processing Pipeline

Features an extensive 5-stage preprocessing pipeline designed to prevent target leakage:

  • RFC 5322 MIME Extraction: Parses complex multipart email structures.
  • HTML Stripping & Thread Slicing: Cleans markup and removes historic reply chains.
  • Unicode Normalization: Mitigates homoglyph attacks.
  • Two-Stage PII Masking: Employs deterministic regex tokenization and spaCy NER to mask URLs, IPs, email addresses, names, and organizations into standard tokens (e.g., [NAME], [URL]).

4. Production Serving

Includes an asynchronous FastAPI REST server supporting inference on pre-computed feature vectors at high throughput (up to 7,497 msgs/s), with ONNX model export support.


Installation

Anti-PLUTO is built with modularity in mind. You can install the core framework or include optional components like the classifier and REST API depending on your needs.

For the most streamlined installation experience, we have provided a setup script that installs the core framework, all optional dependencies, and the required language models in an editable state.

git clone https://github.com/mmaarij/antipluto.git
cd antipluto
chmod +x setup_env.sh
./setup_env.sh

Standard Installation

Once published to PyPI, you can install the core framework directly:

pip install antipluto

To install directly from the source repository:

git clone https://github.com/mmaarij/antipluto.git
cd antipluto
pip install -e .

Optional Dependencies

You can install specific components of the framework depending on your use case:

  • Classifier Mode: pip install -e ".[classifier]" (Installs ML dependencies)
  • API Mode: pip install -e ".[api]" (Installs FastAPI and Uvicorn for REST endpoints)
  • Development/Full: pip install -e ".[dev,classifier,api]"

Language Models (Required for Masking)

Because direct URL dependencies are restricted on PyPI, you must install the required spaCy language models manually after installing the framework:

# Standard model (Fast, recommended for general usage)
pip install https://github.com/explosion/spacy-models/releases/download/en_core_web_sm-3.8.0/en_core_web_sm-3.8.0-py3-none-any.whl

# Transformer model (Highest accuracy, requires 'dev' extra dependencies)
pip install https://github.com/explosion/spacy-models/releases/download/en_core_web_trf-3.8.0/en_core_web_trf-3.8.0-py3-none-any.whl

GPU Acceleration (Optional)

To enable GPU acceleration for the transformer-based masking and classification, you must install PyTorch with CUDA support manually. Ensure you have the CUDA Toolkit 13.x installed, then run:

pip install cupy-cuda13x
pip install torch --index-url https://download.pytorch.org/whl/cu132 --force-reinstall

Configuration

Before running the framework, you must configure your API keys (if you are generating datasets) and other system settings.

  1. Copy the example configuration file:
    cp antipluto/configs/default.example.yaml antipluto/configs/default.yaml
    
  2. Open antipluto/configs/default.yaml and fill in your respective keys (e.g., OPENROUTER_API_KEY, FOUNDRY_API_KEY).

Note: The default.yaml file is intentionally ignored by git to prevent sensitive data leakage.


Usage

Anti-PLUTO provides a robust CLI interface. You can invoke it after installation as antipluto or via python as python -m antipluto.

1. Preprocessing Data

Run the full ingestion pipeline to clean and process your raw corpora:

# Run with default configuration:
antipluto preprocess

# Run a specific dataset source (e.g., nazario):
antipluto preprocess --source nazario

# Override the export mode (full = 20 features, minimal = core NLP):
antipluto preprocess --mode full

2. PII Masking

Apply Named Entity Recognition (NER) and regex masking to anonymize sensitive information and prevent target leakage.

# Mask all preprocessed cohorts:
antipluto mask --input datasets/datasets_preprocessed/

# Mask a specific file:
antipluto mask --input datasets/datasets_preprocessed/human_written_phishing.jsonl

# High-accuracy transformer masking (requires GPU and en_core_web_trf):
antipluto mask --input datasets/datasets_preprocessed/ --spacy-model en_core_web_trf --batch-size 64

3. LLM Cohort Generation

Generate synthetic LLM cohorts using Azure OpenAI, Anthropic, or OpenRouter models (configurable via default.yaml).

# Generate benign and phishing cohorts:
antipluto generate --cohort benign
antipluto generate --cohort phishing

# Trace a synthetic email back to its original human seed email:
antipluto trace --record '{"source": "gpt-5-mini", "seed_idx": 42, "label": 1}'

4. Training the Classifiers

Train the dual-branch feature fusion classifier with a chosen fusion head (XGBoost, Random Forest, PyTorch MLP, 1D-CNN).

# Train the default XGBoost model on all cohorts:
antipluto train --classifier xgboost

# Compare all available models (XGBoost, RF, MLP, CNN) to find the best configuration:
antipluto compare --classifier all

5. Prediction & Exporting

Classify individual emails and export trained models to ONNX for portable deployment.

# Predict a single email's cohort and phishing probability:
antipluto predict --subject "Urgent Account Update" --body "Click here to secure your account."

# Predict from a raw email file using a specific fusion head:
antipluto predict --file suspicious_email.eml --classifier rf --mode fusion

# Export trained DeBERTa and classifier heads to ONNX format:
antipluto export --classifier all

6. Baseline Benchmarking

Evaluate Anti-PLUTO against industry-standard legacy engines (SpamAssassin & Rspamd) on unseen test emails.

antipluto baseline --engine all

7. Production API Serving

Launch the asynchronous FastAPI inference server to expose REST endpoints for real-time email security prediction.

# Start the API server on 127.0.0.1:8000
antipluto serve --classifier xgboost --mode fusion

# Enable auto-reload for local development:
antipluto serve --reload

8. Validation & Telemetry

Evaluate the statistics and schema integrity of your generated cohorts.

antipluto validate --input datasets/datasets_preprocessed/human_written_benign.jsonl
antipluto stats --input datasets/datasets_preprocessed/human_written_phishing.jsonl

Metadata

Release files for antipluto 1.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for antipluto 1.0.1
File Size Uploaded
antipluto-1.0.1.tar.gz 117.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for antipluto 1.0.1
File Interpreter ABI Platform
antipluto-1.0.1-py3-none-any.whl Python 3 none any Details

Total release size: 265.2 kB

Release files / antipluto-1.0.1.tar.gz

Download URL antipluto-1.0.1.tar.gz
Size 117.5 kB
Tags Source
SHA-256 checksum
How to use checksums
0d2b0367d59d44b4644e3ebeb2b7b0934492f02df6b9074f3e1b5595be959829
BLAKE2b-256 checksum
How to use checksums
36eb23864c80e2c0ae7c4245154f69443fb492dba547f5c2a3d0fcaac75d584f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.

Transparency log

Release files / antipluto-1.0.1-py3-none-any.whl

Download URL antipluto-1.0.1-py3-none-any.whl
Size 147.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5a35685876f041a4f33219dac8a38d49826f6d1d0bf183da6db4da4825e7cb7e
BLAKE2b-256 checksum
How to use checksums
edf91027e76dcecb6e94628a3b8a64b56d16b245e78f88ad686606488161e590
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.0.1 This release

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page