Read-only breach-intelligence MCP: reports THAT an organization was breached and which data types leaked, across full history from public feeds (HaveIBeenPwned, RansomLook, ransomwatch archive, SEC 8-K 1.05). Never the leaked data.
Project description
data-breach-detector
A read-only breach-intelligence MCP server. It answers "has this organization ever been breached, what's the recent breach news, what does two decades of breach history look like, how severe is this threat text" from public disclosure feeds — and reports intelligence, not contents: the existence, timing, scale, category and exposed data-types of a breach, never the leaked records themselves.
Built for defenders and for agents that work on their behalf.
Why this instead of the alternatives
Most breach tooling sits in one of three camps, and each has a structural gap:
- Consumer checkers (HaveIBeenPwned's site) answer one question — "is my email in a breach" — one account at a time, one source at a time.
- Leak-data brokers (DeHashed, IntelX, LeakCheck and the like) sell access to the leaked records themselves. Wiring one into an AI agent hands the agent stolen credentials.
- Enterprise intel platforms (SpyCloud, Recorded Future, Flashpoint) do the join properly — behind five-figure contracts and closed APIs.
This server takes a fourth position:
- Four primary sources, one queryable surface. The verified breach directory (HIBP), a live ransomware leak-site tracker (RansomLook), a ~16k-victim leak-site archive back to 2020 (ransomwatch), and SEC 8-K Item 1.05 filings — companies' own legally mandated "material cybersecurity incident" disclosures. Regulator-grade and criminal-infrastructure-grade evidence in the same index. No key, no contract.
- History is first-class.
breach_history,breach_timelineandbreach_statstreat 2007→today as the product, not a cache: every breach of 2013, an organization's full incident chronology, repeat-victim flagging, per-year and per-actor aggregates. - The ethical boundary is in the code, not the terms of service. No
fetch/crawl/proxy primitives, no
.onionaccess, and a redaction pass strips emails, hashes, IPs, crypto addresses and credential-shaped tokens from every string served. That makes it the breach feed you can safely hand to an autonomous agent. - Honesty is instrumented.
feed_sourcesreports each feed's newest item, a staleness flag and the last fetch error — a dead upstream is a served fact, not a silent hole. (The ransomwatch project itself froze in June 2025; this server says so instead of pretending.) - MCP-native, free, MIT, self-hostable. One
pip install, stdio or streamable-HTTP.
What it does not do
- No arbitrary URL fetch, no crawl, no proxy — no general scraping primitives.
- No
.onionmarketplace access, no transactions. - Never returns the raw text of a dump, paste or leak. A redaction layer strips emails, hashes, IPs, crypto addresses and credential-shaped tokens from every string returned.
Sources (public, no key)
- HaveIBeenPwned
/api/v3/breaches— the verified breach directory back to 2007: domain, breach date, pwn count, exposed data categories. - RansomLook (
ransomlook.io) — live ransomware leak-site tracker. - ransomwatch (
joshhighet/ransomwatch) — frozen archive of ~16k leak-site posts, Jan 2020 → Jun 2025, retained as history. - SEC EDGAR — 8-K filings carrying Item 1.05 Material Cybersecurity Incidents (mandatory first-party disclosure since Dec 2023).
Tools
| tool | what it returns |
|---|---|
breach_news(since_days, sector, source, limit) |
recent disclosures — entity, date, scale, exposed data types, severity |
check_exposure(query) |
does a domain/company appear anywhere in breach data — yes/no + metadata |
breach_history(query, year_from, year_to, sector, data_type, min_accounts, order, limit) |
search the full archive back to 2007 |
breach_timeline(entity) |
one organization's incident-by-incident chronology + repeat-victim assessment |
breach_stats(group_by, sector) |
aggregates per year / source / data type / threat level / ransomware actor |
assess_threat(text) |
classify a piece of security text — level, categories, action (no network) |
feed_sources() |
feeds, per-source freshness, staleness flags, last fetch errors |
Run
pip install data-breach-detector
data-breach-detector # stdio (for MCP clients)
data-breach-detector --http # streamable-HTTP on 127.0.0.1:8790/mcp
Or point an MCP client at the config:
{ "mcpServers": { "data_breach_detector": {
"command": "data-breach-detector"
} } }
Hosted remote: https://breach.seiche.info/mcp
License
MIT. The breach data belongs to its sources (HaveIBeenPwned, RansomLook, ransomwatch, SEC EDGAR); this tool only aggregates their public disclosure metadata, with attribution.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file data_breach_detector-0.2.0.tar.gz.
File metadata
- Download URL: data_breach_detector-0.2.0.tar.gz
- Upload date:
- Size: 17.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a892e6f4ae57d1691432a7222584154c0676ed1ac4053b1df63b2bcf9fb643e0
|
|
| MD5 |
31b386e7f911ae391afd5883c621f2af
|
|
| BLAKE2b-256 |
805532e997fd3018a8a97f1cf58fb48dbce0c0d882abb902ed397f74e0c49f9c
|
Provenance
The following attestation bundles were made for data_breach_detector-0.2.0.tar.gz:
Publisher:
publish-pypi.yml on beepboop2025/data-breach-detector
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
data_breach_detector-0.2.0.tar.gz -
Subject digest:
a892e6f4ae57d1691432a7222584154c0676ed1ac4053b1df63b2bcf9fb643e0 - Sigstore transparency entry: 2313845282
- Sigstore integration time:
-
Permalink:
beepboop2025/data-breach-detector@e5a4be231a08d372eb238ae27b28b217faac6158 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/beepboop2025
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish-pypi.yml@e5a4be231a08d372eb238ae27b28b217faac6158 -
Trigger Event:
workflow_dispatch
-
Statement type:
File details
Details for the file data_breach_detector-0.2.0-py3-none-any.whl.
File metadata
- Download URL: data_breach_detector-0.2.0-py3-none-any.whl
- Upload date:
- Size: 16.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
96e0338c56f41745d7376debce8250016dbebf8c76f39736011929ebe1d7143a
|
|
| MD5 |
64a8300286560d8de16e7183ed0aee4f
|
|
| BLAKE2b-256 |
b0e0af55c6caea48a221cf0ebcb84c166fe9e4eecfca6c9ed78e734b6dc733f2
|
Provenance
The following attestation bundles were made for data_breach_detector-0.2.0-py3-none-any.whl:
Publisher:
publish-pypi.yml on beepboop2025/data-breach-detector
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
data_breach_detector-0.2.0-py3-none-any.whl -
Subject digest:
96e0338c56f41745d7376debce8250016dbebf8c76f39736011929ebe1d7143a - Sigstore transparency entry: 2313845333
- Sigstore integration time:
-
Permalink:
beepboop2025/data-breach-detector@e5a4be231a08d372eb238ae27b28b217faac6158 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/beepboop2025
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish-pypi.yml@e5a4be231a08d372eb238ae27b28b217faac6158 -
Trigger Event:
workflow_dispatch
-
Statement type: