Predict categories based on domain names and their content
Project description
piedomains: Classify website content using ML Models or LLMs
🚀 What's New in v0.6.0
- Streamlined JSON API: Simple, consistent JSON responses for easy integration with any workflow
- Enhanced LLM Support: Built-in support for OpenAI, Anthropic, and Google AI models with custom category definitions
- Advanced Archive Analysis: Analyze historical website versions from archive.org with intelligent rate limiting
- Separated Data Collection: Collect website content once, run multiple classification approaches (ML + LLM + ensemble)
- 39 Content Categories: Comprehensive classification including news, shopping, social media, education, finance, and more
Installation
pip install piedomains
Requires Python 3.11+
Basic Usage
from piedomains import DomainClassifier, DataCollector
classifier = DomainClassifier()
run = classifier.classify(["cnn.com", "amazon.com", "wikipedia.org"])
for result in run["results"]:
print(f"{result['domain']}: {result['category']} ({result['confidence']:.3f})")
# Output:
# cnn.com: news (0.876)
# amazon.com: shopping (0.923)
# wikipedia.org: education (0.891)
Knowing What Failed
Every call returns both the per-domain rows and a run report, so a long URL list
never fails silently. Each row carries a machine-readable status, the stage it
reached, and a stable error_code:
run = classifier.classify(open("domains.txt").read().split())
print(run["report"])
# {'run_id': '8fe4cc80eeb5', 'total': 500, 'classified': 461, 'failed': 39,
# 'by_reason': {'dns_error': 12, 'timeout': 9, 'cannot_classify': 11, 'thin_content': 7},
# 'by_stage': {'fetch': 32, 'infer': 7},
# 'by_source': {'live': 448, 'archive': 13},
# 'missing': ['foo.com', 'bar.org', ...],
# 'started_at': ..., 'finished_at': ..., 'elapsed_ms': 184203}
# Retry only what is worth retrying
retry = [r["domain"] for r in run["results"] if r.get("retryable")]
error_code is a closed, stable set — safe to group on: invalid_domain,
dns_error, connection_error, timeout, http_error, robots_blocked,
content_type_rejected, content_too_large, no_archive_snapshot,
archive_rate_limited, empty_text, missing_input_path, missing_screenshot,
model_load_error, model_error, llm_error, bot_blocked, thin_content,
cannot_classify, unknown.
cannot_classify is the umbrella terminal state: branch on it when you do not want
to enumerate every cause.
Bot Walls
Roughly one domain in seven serves an anti-bot interstitial rather than a page.
Changing the user-agent does not help — DataDome and Cloudflare fingerprint headless
Chromium itself — so piedomains detects the interstitial and refetches the page from
archive.org, which already has it. No evasion, and no challenge page classified as
though it were the site.
run = classifier.classify(["etsy.com", "reuters.com", "indeed.com"])
for r in run["results"]:
print(r["domain"], r["category"], r["source"], r["snapshot_timestamp"])
# etsy.com shopping archive 20260727020309
# reuters.com news archive 20260718201522
# indeed.com jobsearch archive 20260724170521
A capture older than archive_max_age_days (default 365) is refused rather than
passed off as the live page; those domains report cannot_classify. Set
archive_fallback=False to turn this off and have blocked domains report
bot_blocked instead.
From the command line:
classify_domains --file domains.txt --output json --report run-report.json
# stdout: {"results": [...], "report": {...}}
# stderr: 461/500 classified, 39 failed (run 8fe4cc80eeb5)
# dns_error: 12
# timeout: 9
# exit status is non-zero if any domain failed
Structured logs
Human-readable text is the default. For pipelines, opt into JSON lines — every
record carries the run_id, plus domain/stage/error_code where relevant, so
logs join against the report:
PIEDOMAINS_LOG_FORMAT=json classify_domains --file domains.txt
# {"ts":"...","level":"ERROR","run_id":"8fe4cc80eeb5","domain":"foo.com",
# "stage":"fetch","error_code":"timeout","msg":"navigation timed out"}
Classification Methods
# Combined text + image analysis (most accurate)
run = classifier.classify(["github.com"])
# Text-only classification (faster)
run = classifier.classify_by_text(["news.google.com"])
# Image-only classification
run = classifier.classify_by_images(["instagram.com"])
# Batch processing with separated workflow
collector = DataCollector()
collection = collector.collect_batch(domains, batch_size=50)
results = classifier.classify_from_collection(collection, method="text")
Historical Analysis
# Analyze archived versions from archive.org
old_run = classifier.classify(["facebook.com"], archive_date="20100101")
# Batch processing with archive.org (respects rate limits)
domains = ["google.com", "wikipedia.org", "cnn.com"]
collector = DataCollector(archive_date="20050101")
collection = collector.collect_batch(domains, batch_size=10) # Archive.org uses conservative defaults
historical_results = classifier.classify_from_collection(collection, method="text")
How archive analysis works
Snapshot discovery and retrieval go through the
wayback library (CDX + Memento):
- Only status-200 captures are used. An archived redirect or 404 is never classified
as if it were content; the domain reports
no_archive_snapshotinstead. - The capture actually used is reported as
snapshot_timestamp— not the date you asked for. Requesting20100101forcnn.comyields20100101041727. - Text is fetched raw via Wayback's
id_playback mode: no injected Wayback JavaScript, no rewritten URLs, and no browser required — so the text path is fast. - Screenshots render via
if_, which hides the Wayback toolbar but keeps archived CSS and images, so the page looks as it did. (id_would render an unstyled skeleton.) - Rate limiting, retries and exponential backoff are handled by the
waybacksession.
The cache key includes the archive date, so a live fetch and snapshots from different years coexist rather than overwriting one another:
cache/html/cnn.com.html # live
cache/html/cnn.com@20050101.html # 2005 snapshot
cache/html/cnn.com@20150101.html # 2015 snapshot
Configure via piedomains.config: archive_max_parallel, archive_window_days,
archive_search_rate, archive_memento_rate, archive_retries, archive_backoff.
LLM Classification
# Configure LLM provider
classifier.configure_llm(
provider="openai",
model="gpt-4o",
api_key="sk-...",
categories=["news", "shopping", "social", "tech"]
)
# LLM-powered classification
result = classifier.classify_by_llm(["example.com"])
# With custom instructions
result = classifier.classify_by_llm(
["site.com"],
custom_instructions="Classify by educational value"
)
Set API keys via environment variables:
export OPENAI_API_KEY="sk-..."
export ANTHROPIC_API_KEY="sk-ant-..."
export GOOGLE_API_KEY="..."
Categories
39 categories: news, finance, shopping, education, government, adult content, gambling, social networks, search engines, and others based on Shallalist taxonomy.
Security & Docker
v0.5.0 includes production-ready Docker containerization for secure domain analysis:
# Build secure sandbox container
docker build -t piedomains-sandbox .
# Run with security constraints (2GB RAM, 2 CPU, read-only filesystem)
docker run --rm --memory=2g --cpus=2 --read-only \
--tmpfs /tmp --tmpfs /var/tmp \
piedomains-sandbox python -c "
from piedomains import DomainClassifier
classifier = DomainClassifier()
run = classifier.classify(['example.com'])
print(run['results'][0]['category'])
"
Batch Processing in Container:
# Use the included secure classification script
cd examples/sandbox
echo -e "wikipedia.org\ngithub.com\ncnn.com" > domains.txt
python3 secure_classify.py --file domains.txt
For testing, use known-safe domains: ["wikipedia.org", "github.com", "cnn.com"]
Documentation
Development
git clone https://github.com/themains/piedomains
cd piedomains
uv sync --all-groups
uv run pytest tests/ -v
License
MIT License
Citation
@software{piedomains,
title={piedomains: AI-powered domain content classification},
author={Chintalapati, Rajashekar and Sood, Gaurav},
year={2024},
url={https://github.com/themains/piedomains}
}
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file piedomains-0.7.0.tar.gz.
File metadata
- Download URL: piedomains-0.7.0.tar.gz
- Upload date:
- Size: 3.7 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f77004d983b003392056d781befad75df39146b8c42682d245e4b75b695bc5bf
|
|
| MD5 |
0c2bdbbb83ab7a3a3c03a0e64e03e28f
|
|
| BLAKE2b-256 |
b9ad9d6b068b5e0fe53dfd7e93b54b4e36186138b8286b8d67ee34aafc1b715c
|
Provenance
The following attestation bundles were made for piedomains-0.7.0.tar.gz:
Publisher:
python-publish.yml on themains/piedomains
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
piedomains-0.7.0.tar.gz -
Subject digest:
f77004d983b003392056d781befad75df39146b8c42682d245e4b75b695bc5bf - Sigstore transparency entry: 2256948407
- Sigstore integration time:
-
Permalink:
themains/piedomains@d6ce245010dc8f2abaccfa001b37c94bad1cd674 -
Branch / Tag:
refs/tags/v0.7.0 - Owner: https://github.com/themains
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
python-publish.yml@d6ce245010dc8f2abaccfa001b37c94bad1cd674 -
Trigger Event:
push
-
Statement type:
File details
Details for the file piedomains-0.7.0-py3-none-any.whl.
File metadata
- Download URL: piedomains-0.7.0-py3-none-any.whl
- Upload date:
- Size: 167.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7243ea1530d3d223eb16595603af5655b025812ea50b7d7e79cef1d05c824b0c
|
|
| MD5 |
c679161bd3e5f4ba461fcf35f52d3ac8
|
|
| BLAKE2b-256 |
12596855f92ea8f42f2d1b0e3fccd65d524ab065fbe11a4f312b05a43569e83b
|
Provenance
The following attestation bundles were made for piedomains-0.7.0-py3-none-any.whl:
Publisher:
python-publish.yml on themains/piedomains
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
piedomains-0.7.0-py3-none-any.whl -
Subject digest:
7243ea1530d3d223eb16595603af5655b025812ea50b7d7e79cef1d05c824b0c - Sigstore transparency entry: 2256948416
- Sigstore integration time:
-
Permalink:
themains/piedomains@d6ce245010dc8f2abaccfa001b37c94bad1cd674 -
Branch / Tag:
refs/tags/v0.7.0 - Owner: https://github.com/themains
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
python-publish.yml@d6ce245010dc8f2abaccfa001b37c94bad1cd674 -
Trigger Event:
push
-
Statement type: