Skip to main content

laya-phishield logo

laya-phishield

Explainable phishing detection that runs locally.
Eight focused semantic signals and deterministic email checks become one risk score—with every reason attached.

Python 3.10+ PyPI: laya-phishield CI status FastAPI Apache 2.0 license Local inference Project status: early development

Quick start · How it works · Evaluation · Interactive explainer


Animated demo of laya-phishield analyzing an email and explaining its phishing verdict
Autoplays and loops · ▶ Open the 1080p video

The short version

Most phishing classifiers answer one broad question: “Is this phishing?” That can produce a useful score, but it gives an analyst little evidence to review.

laya-phishield decomposes that decision into eight narrow yes/no signals. A deterministic pre-pass catches header and URL artifacts, laya evaluates the semantic signals locally in one forward pass, and a logistic head combines everything into an auditable verdict.

0.9531
test AUC
0.913
precision @ 0.5
0.808
recall @ 0.5
$0
API cost / 1,000 emails

Why this approach

  • Explainable by construction. Each verdict includes signal probabilities, deterministic flags, and the features that contributed most to the score.
  • Local and inexpensive. Detection runs on a local laya decision model; email content does not need to be sent to a hosted LLM API.
  • Harder to fool with obfuscation. Header, domain, punycode, homoglyph, SPF, and DKIM checks complement semantic analysis.
  • Operationally flexible. Scan .eml, .mbox, or .jsonl files, call the HTTP API, or use the Streamlit demo.
  • Measured honestly. Near-duplicates are removed before a per-class temporal split, and the documented tail-recall limitation is kept visible.

How it works

flowchart LR
    A[Raw RFC 822 email] --> B[Deterministic pre-pass]
    B --> C[Structured email state]
    C --> D[Eight atomic laya signals]
    B --> E[Header and URL flags]
    D --> F[Logistic combination head]
    E --> F
    F --> G[Risk score + label + reasons]
  1. Parse and inspect. extract.py extracts the subject, sender and reply-to domains, first URL host, and a plain-text body excerpt. It also raises deterministic flags for reply-to mismatch, IP-literal or punycode links, brand homoglyphs, brands hidden in subdomains, display-name mismatch, and SPF/DKIM failures.
  2. Score eight atomic signals. signals.py asks the local decision model eight positively phrased questions in one forward pass. The full contract and hard negatives live in data/signals_schema.md.
  3. Combine and explain. A trained logistic head weighs the eight probabilities and deterministic flags. It returns a score, a label, and the top feature contributions rather than an opaque binary decision.

The eight signals

Signal What it asks
asks_credentials Does the email ask the recipient to confirm a password or account login?
urgency_pressure Does it demand immediate action or threaten consequences for delay?
payment_gift_request Does it request money, gift cards, a transfer, or cryptocurrency?
brand_impersonation Does it borrow a known brand while using an unrelated sender domain?
requests_pii Does it request identity, card, or other sensitive personal data?
too_good_to_be_true Does it promise an unearned prize, inheritance, or windfall?
suspicious_instructions Does it ask the recipient to keep secrets or bypass normal procedure?
external_link_risk Does it point to a host unrelated to the claimed sender or brand?

Quick start

Requires Python 3.10+ and uv. Released on PyPI, so either install directly (note: this pulls laya and with it torch/transformers, roughly a gigabyte of wheels) or clone for the eval tooling and demo:

uv tool install laya-phishield        # or: uv pip install laya-phishield
git clone https://github.com/Gjusev/laya-phishield.git
cd laya-phishield
uv venv
uv pip install -e .

# Scan one email, a mailbox, or JSONL records shaped as {"raw": "..."}
uv run laya-phishield scan suspicious.eml
uv run laya-phishield scan inbox.mbox --json
uv run laya-phishield scan mail-1.eml mail-2.eml --limit 100

The model checkpoint is downloaded on first use. Human-readable output names the strongest reasons; --json emits one complete verdict object per email.

PHISHING 0.990  suspicious.eml#1  [asks_credentials 2.56, urgency_pressure 1.86]
1 email scanned, 1 flagged

HTTP API

Start the app as a Uvicorn factory so the model remains lazy-loaded. The command below adds Uvicorn to the run environment without changing the project dependencies:

uv run --with uvicorn uvicorn laya_phishield.serve:create_app --factory --port 8000

curl -s http://localhost:8000/scan \
  -H 'content-type: application/json' \
  -d '{"raw": "<full RFC822 email>"}'

Interactive API documentation is available at http://localhost:8000/docs while the server is running.

Paste-an-email demo

uv pip install -e ".[demo]"
uv run streamlit run app.py

Visual walkthrough

The interactive explainer replays both examples with real measured values from the shipped model.

Phishing email Legitimate email
Phishing email walkthrough Legitimate email walkthrough

Evaluation

The benchmark compares the composite model with a keyword baseline and a single forced-choice laya question.

Metric Keyword Forced choice Composite
AUC 0.5947 0.9444 0.9531
Precision @ 0.5 1.000 0.902 0.913
Recall @ 0.5 0.051 0.590 0.808
FPR @ 95% TPR 1.000 0.124 0.210
API cost / 1,000 emails $0 $0 $0

Results use a temporal test split of 183 emails: 105 legitimate and 78 phishing. The source corpus combines Nazario phishing emails with a reservoir sample of legitimate Enron mail.

To reduce leakage, MinHash + LSH removes near-duplicates across both classes before splitting. Each class is then split chronologically: the oldest 70% for training and the newest 30% for testing; undated messages remain in training. Full coefficients, ablations, and run metadata are stored in evals/results.json.

The composite improves over forced choice by 0.9 AUC points and 21.8 recall points at the 0.5 operating threshold. The limitation matters too: at 95% recall, forced choice has the lower false-positive rate (0.124 vs 0.210). The composite is strongest at the practical operating point measured here, not at the extreme-recall tail.

Reproduce the benchmark

uv pip install -e ".[eval]"
uv run python evals/prepare_data.py --max-phish 350 --max-legit 350
uv run python evals/run_eval.py --skip-gpt

# Optional hosted-LLM baseline; incurs OpenAI API usage
OPENAI_API_KEY=... uv run python evals/run_eval.py

The featurization cache makes interrupted runs resumable.

Testing and development

uv pip install -e ".[dev]"
uv run pytest

The default suite is fast and uses a deterministic fake agent, so it never downloads a checkpoint. The opt-in smoke test uses the real model:

uv run pytest -m slow

Project map

src/laya_phishield/
├── extract.py       # email parsing and deterministic checks
├── signals.py       # eight semantic signal definitions
├── combine.py       # logistic scoring and feature attribution
├── pipeline.py      # raw email → explainable verdict
├── cli.py           # batch CLI
├── serve.py         # FastAPI application
└── data/head.json   # trained combination head

evals/               # dataset preparation, baselines, and benchmark
docs/                # interactive explainer and visual assets
tests/               # fast behavior tests and real-model smoke tests

Acknowledgements

Built on laya, the open-source System 1 decision engine. Evaluation data comes from the Nazario phishing corpus and the Enron email dataset.

License

Apache 2.0. See LICENSE.

Metadata

Release files for laya-phishield 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for laya-phishield 0.1.1
File Size Uploaded
laya_phishield-0.1.1.tar.gz 5.0 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for laya-phishield 0.1.1
File Interpreter ABI Platform
laya_phishield-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 5.0 MB

Release files / laya_phishield-0.1.1.tar.gz

Download URL laya_phishield-0.1.1.tar.gz
Size 5.0 MB
Tags Source
SHA-256 checksum
How to use checksums
ab8ae7180c2265ec2509046cdbaf8412f1cd6c3cf8d35fa4c52ca9b8655d7720
BLAKE2b-256 checksum
How to use checksums
ad49bee96fdbf985f6e58f68968bd382c651f383bd12ab8b6c3fd5ead8272d30
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / laya_phishield-0.1.1-py3-none-any.whl

Download URL laya_phishield-0.1.1-py3-none-any.whl
Size 25.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5c5e3a228f720d172da3331bbfb6081f67d9235efc4b578f3cad079ff193b605
BLAKE2b-256 checksum
How to use checksums
d653be4e5aaa1184bf526c8f06ad55d42f9e7cf066a18edb4f5110c597062a49
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page