laya-phishield
Explainable phishing detection that runs locally.
Eight focused semantic signals and deterministic email checks become one risk score—with every reason attached.
Quick start · How it works · Evaluation · Interactive explainer
Autoplays and loops · ▶ Open the 1080p video
The short version
Most phishing classifiers answer one broad question: “Is this phishing?” That can produce a useful score, but it gives an analyst little evidence to review.
laya-phishield decomposes that decision into eight narrow yes/no signals. A deterministic pre-pass catches header and URL artifacts, laya evaluates the semantic signals locally in one forward pass, and a logistic head combines everything into an auditable verdict.
| 0.9531 test AUC |
0.913 precision @ 0.5 |
0.808 recall @ 0.5 |
$0 API cost / 1,000 emails |
Why this approach
- Explainable by construction. Each verdict includes signal probabilities, deterministic flags, and the features that contributed most to the score.
- Local and inexpensive. Detection runs on a local laya decision model; email content does not need to be sent to a hosted LLM API.
- Harder to fool with obfuscation. Header, domain, punycode, homoglyph, SPF, and DKIM checks complement semantic analysis.
- Operationally flexible. Scan
.eml,.mbox, or.jsonlfiles, call the HTTP API, or use the Streamlit demo. - Measured honestly. Near-duplicates are removed before a per-class temporal split, and the documented tail-recall limitation is kept visible.
How it works
flowchart LR
A[Raw RFC 822 email] --> B[Deterministic pre-pass]
B --> C[Structured email state]
C --> D[Eight atomic laya signals]
B --> E[Header and URL flags]
D --> F[Logistic combination head]
E --> F
F --> G[Risk score + label + reasons]
- Parse and inspect.
extract.pyextracts the subject, sender and reply-to domains, first URL host, and a plain-text body excerpt. It also raises deterministic flags for reply-to mismatch, IP-literal or punycode links, brand homoglyphs, brands hidden in subdomains, display-name mismatch, and SPF/DKIM failures. - Score eight atomic signals.
signals.pyasks the local decision model eight positively phrased questions in one forward pass. The full contract and hard negatives live indata/signals_schema.md. - Combine and explain. A trained logistic head weighs the eight probabilities and deterministic flags. It returns a score, a label, and the top feature contributions rather than an opaque binary decision.
The eight signals
| Signal | What it asks |
|---|---|
asks_credentials |
Does the email ask the recipient to confirm a password or account login? |
urgency_pressure |
Does it demand immediate action or threaten consequences for delay? |
payment_gift_request |
Does it request money, gift cards, a transfer, or cryptocurrency? |
brand_impersonation |
Does it borrow a known brand while using an unrelated sender domain? |
requests_pii |
Does it request identity, card, or other sensitive personal data? |
too_good_to_be_true |
Does it promise an unearned prize, inheritance, or windfall? |
suspicious_instructions |
Does it ask the recipient to keep secrets or bypass normal procedure? |
external_link_risk |
Does it point to a host unrelated to the claimed sender or brand? |
Quick start
Requires Python 3.10+ and uv. Released on PyPI,
so either install directly (note: this pulls laya and with it
torch/transformers, roughly a gigabyte of wheels) or clone for the eval
tooling and demo:
uv tool install laya-phishield # or: uv pip install laya-phishield
git clone https://github.com/Gjusev/laya-phishield.git
cd laya-phishield
uv venv
uv pip install -e .
# Scan one email, a mailbox, or JSONL records shaped as {"raw": "..."}
uv run laya-phishield scan suspicious.eml
uv run laya-phishield scan inbox.mbox --json
uv run laya-phishield scan mail-1.eml mail-2.eml --limit 100
The model checkpoint is downloaded on first use. Human-readable output names the strongest reasons; --json emits one complete verdict object per email.
PHISHING 0.990 suspicious.eml#1 [asks_credentials 2.56, urgency_pressure 1.86]
1 email scanned, 1 flagged
HTTP API
Start the app as a Uvicorn factory so the model remains lazy-loaded. The command below adds Uvicorn to the run environment without changing the project dependencies:
uv run --with uvicorn uvicorn laya_phishield.serve:create_app --factory --port 8000
curl -s http://localhost:8000/scan \
-H 'content-type: application/json' \
-d '{"raw": "<full RFC822 email>"}'
Interactive API documentation is available at http://localhost:8000/docs while the server is running.
Paste-an-email demo
uv pip install -e ".[demo]"
uv run streamlit run app.py
Visual walkthrough
The interactive explainer replays both examples with real measured values from the shipped model.
| Phishing email | Legitimate email |
Evaluation
The benchmark compares the composite model with a keyword baseline and a single forced-choice laya question.
| Metric | Keyword | Forced choice | Composite |
|---|---|---|---|
| AUC | 0.5947 | 0.9444 | 0.9531 |
| Precision @ 0.5 | 1.000 | 0.902 | 0.913 |
| Recall @ 0.5 | 0.051 | 0.590 | 0.808 |
| FPR @ 95% TPR | 1.000 | 0.124 | 0.210 |
| API cost / 1,000 emails | $0 | $0 | $0 |
Results use a temporal test split of 183 emails: 105 legitimate and 78 phishing. The source corpus combines Nazario phishing emails with a reservoir sample of legitimate Enron mail.
To reduce leakage, MinHash + LSH removes near-duplicates across both classes before splitting. Each class is then split chronologically: the oldest 70% for training and the newest 30% for testing; undated messages remain in training. Full coefficients, ablations, and run metadata are stored in evals/results.json.
The composite improves over forced choice by 0.9 AUC points and 21.8 recall points at the 0.5 operating threshold. The limitation matters too: at 95% recall, forced choice has the lower false-positive rate (0.124 vs 0.210). The composite is strongest at the practical operating point measured here, not at the extreme-recall tail.
Reproduce the benchmark
uv pip install -e ".[eval]"
uv run python evals/prepare_data.py --max-phish 350 --max-legit 350
uv run python evals/run_eval.py --skip-gpt
# Optional hosted-LLM baseline; incurs OpenAI API usage
OPENAI_API_KEY=... uv run python evals/run_eval.py
The featurization cache makes interrupted runs resumable.
Testing and development
uv pip install -e ".[dev]"
uv run pytest
The default suite is fast and uses a deterministic fake agent, so it never downloads a checkpoint. The opt-in smoke test uses the real model:
uv run pytest -m slow
Project map
src/laya_phishield/
├── extract.py # email parsing and deterministic checks
├── signals.py # eight semantic signal definitions
├── combine.py # logistic scoring and feature attribution
├── pipeline.py # raw email → explainable verdict
├── cli.py # batch CLI
├── serve.py # FastAPI application
└── data/head.json # trained combination head
evals/ # dataset preparation, baselines, and benchmark
docs/ # interactive explainer and visual assets
tests/ # fast behavior tests and real-model smoke tests
Acknowledgements
Built on laya, the open-source System 1 decision engine. Evaluation data comes from the Nazario phishing corpus and the Enron email dataset.
License
Apache 2.0. See LICENSE.
Metadata
Release files for laya-phishield 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| laya_phishield-0.1.1.tar.gz | 5.0 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| laya_phishield-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 5.0 MB
Release files / laya_phishield-0.1.1.tar.gz
| Download URL | laya_phishield-0.1.1.tar.gz |
|---|---|
| Size | 5.0 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ab8ae7180c2265ec2509046cdbaf8412f1cd6c3cf8d35fa4c52ca9b8655d7720
|
|
BLAKE2b-256 checksum How to use checksums |
ad49bee96fdbf985f6e58f68968bd382c651f383bd12ab8b6c3fd5ead8272d30
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency logRelease files / laya_phishield-0.1.1-py3-none-any.whl
| Download URL | laya_phishield-0.1.1-py3-none-any.whl |
|---|---|
| Size | 25.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5c5e3a228f720d172da3331bbfb6081f67d9235efc4b578f3cad079ff193b605
|
|
BLAKE2b-256 checksum How to use checksums |
d653be4e5aaa1184bf526c8f06ad55d42f9e7cf066a18edb4f5110c597062a49
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency log