Skip to main content

piedomains

Guess what a website is about from its homepage.

CI PyPI Version Downloads Documentation Text model Image model

Give it a list of domains. It fetches each homepage, reads the text, and returns one of 44 categories with a calibrated probability.

from piedomains import DomainClassifier

classifier = DomainClassifier()
run = classifier.classify(["nytimes.com", "wikipedia.org"])

for r in run["results"]:
    if r["status"] == "ok":
        print(f"{r['domain']:16s} {r['category']:10s} {r['confidence']:.3f}")
    else:
        print(f"{r['domain']:16s} failed: {r['error_code']}")

# nytimes.com      news       0.992
# wikipedia.org    library    0.554

A domain that could not be fetched comes back with category set to None and a reason in error_code, so check status before reading a label. Homepages change, so the exact probabilities move a little between runs.

Install

pip install piedomains

Python 3.11 or newer. The wheel is about 245 KB; model weights download from Hugging Face on first use and cache locally.

Live pages render in headless Chromium, so install the browser once:

playwright install chromium

Archived pages are fetched over plain HTTP and need no browser.

What you get back

Every call returns per-domain rows and a run report. One row:

run = classifier.classify(["nytimes.com"])
row = run["results"][0]
{
  "domain": "nytimes.com",
  "category": "news",
  "confidence": 0.992,
  "categories": [{"category": "news", "probability": 0.992}],
  "raw_predictions": {"news": 0.992, "politics": 0.002, "...": "all 44"},
  "model_used": "text/shallalist_ml",
  "source": "live",
  "snapshot_timestamp": null,
  "status": "ok",
  "stage": "infer",
  "error_code": null,
  "retryable": false
}

category is the argmax and confidence is its probability.

The label set is not mutually exclusive. shopping says what a site does, automobile says what it is about, and a car dealership is honestly both, so categories reports every label above a probability floor. Most domains get one label and ambiguous ones get several:

nytimes.com      news 0.992
wikipedia.org    library 0.554, searchengines 0.184
nasa.gov         automobile 0.305, science 0.136, news 0.124, military 0.113

Reporting the runners-up raises the chance of covering the right answer. Whether any particular runner-up is itself correct is not something the evaluation data can answer, because its gold labels are single-label. Treat the extra labels as candidates, not findings.

How well does it work

Depends on the domain, and the honest answer is that this is not settled.

The nasa.gov row above is the useful example. The top label is automobile at 0.305, which is wrong, and the closest thing to a right answer sits third. Low confidence is doing its job there, and a caller reading confidence would have known not to trust it.

There are two numbers, and they disagree. On documents held out of its own training corpus, the shipped text checkpoint reports 0.818 accuracy and 0.758 macro-F1. On 155 popular domains carrying independent human labels from Curlie, the previous checkpoint agreed 0.543 of the time.

Part of that gap is real. Held-out documents share the training distribution and its labeling conventions, so the first number is an in-distribution ceiling rather than what an arbitrary crawl will give you. But the gap is not all generalization loss, because the two taxonomies have never been reconciled. A site Curlie files under Reference and this model calls library is scored as a disagreement under a hand-written mapping nobody has audited, and the same goes for every boundary the two schemes draw differently. Until the taxonomies are reconciled and the labeled set is itself audited, both numbers are weak evidence about accuracy on your data.

What is worth acting on is that quality is very uneven across the 44 classes. The text checkpoint reports F1 of 0.99 on parked and 0.97 on science and religion, against 0.15 on urlshortener, 0.33 on library, 0.37 on socialnet, and 0.45 on shopping. If the categories you care about are in the second group, measure before trusting.

Read any of this back from the pinned checkpoint rather than taking it here:

import json
from huggingface_hub import hf_hub_download
from piedomains.text import DEFAULT_TEXT_MODEL, DEFAULT_TEXT_REVISION

path = hf_hub_download(
    DEFAULT_TEXT_MODEL, "test_metrics.json", revision=DEFAULT_TEXT_REVISION
)
metrics = json.load(open(path))
print(metrics["accuracy"], metrics["per_class"]["urlshortener"]["f1"])

Confidence is a temperature-scaled probability rather than a raw softmax, so it is meaningful enough to threshold on. The training corpus is overwhelmingly English, and non-English pages score measurably worse.

Knowing what failed

A long domain list never fails silently. Each row carries a status, the stage it reached, and a stable error_code, and the report aggregates them:

run = classifier.classify(open("domains.txt").read().split())
print(run["report"])
{
  "run_id": "69c9c2e30071",
  "started_at": "2026-08-20T17:57:03.348685+00:00",
  "finished_at": "2026-08-20T17:57:07.854997+00:00",
  "elapsed_ms": 4506,
  "total": 7,
  "classified": 6,
  "failed": 1,
  "by_reason": {"dns_error": 1},
  "by_stage": {"fetch": 1},
  "by_source": {"live": 6},
  "missing": ["this-domain-does-not-exist-9z8x7.com"]
}

Retry only what is worth retrying:

retry = [r["domain"] for r in run["results"] if r.get("retryable")]

error_code is a closed set of 21 values, safe to group on: invalid_domain, dns_error, private_address, connection_error, timeout, http_error, robots_blocked, content_type_rejected, content_too_large, no_archive_snapshot, archive_rate_limited, empty_text, missing_input_path, missing_screenshot, model_load_error, model_error, llm_error, bot_blocked, thin_content, cannot_classify, unknown. Branch on cannot_classify when you do not want to enumerate every cause.

Command line

classify_domains --file domains.txt --report run-report.json
6/7 classified, 1 failed (run 961cfaf7a9f7)
  dns_error: 1
  no result for: this-domain-does-not-exist-9z8x7.com
wikipedia.org                          ok      library      0.554
github.com                             ok      downloads    0.387
nytimes.com                            ok      news         0.992
this-domain-does-not-exist-9z8x7.com   failed  None         n/a     dns_error
etsy.com                               ok      shopping     0.321
espn.com                               ok      news         0.368
nasa.gov                               ok      automobile   0.305

The counts go to stderr and the rows to stdout, so they redirect separately. Exit status is 1 if any domain failed. --output json emits the full run object instead. --method takes text (the default), images, or combined. --archive-date YYYYMMDD classifies an archived snapshot instead of the live page.

For pipelines, opt into JSON logs. Every record carries the run_id, so logs join against the report, and the closing record repeats the failure counts:

PIEDOMAINS_LOG_FORMAT=json classify_domains --file domains.txt
{"ts": "2026-08-20T10:58:42-0700", "level": "WARNING", "logger": "piedomains",
 "msg": "Failed to fetch data for example.invalid: refused address: dns_error",
 "run_id": "4e9146942d05"}
{"ts": "2026-08-20T10:58:54-0700", "level": "INFO", "logger": "piedomains",
 "msg": "Run 4e9146942d05 finished: 1/2 classified, 1 failed",
 "run_id": "4e9146942d05", "by_reason": {"dns_error": 1}}

Any keyword the package logs through extra= is promoted to a top-level key, so records carry more than msg where the call site supplies it.

Bot walls

Roughly one domain in seven serves an anti-bot interstitial rather than a page. Changing the user-agent does not help, because DataDome and Cloudflare fingerprint headless Chromium itself. So piedomains detects the interstitial and refetches the page from archive.org, which already has it. No evasion, and no challenge page classified as though it were the site.

run = classifier.classify(["etsy.com", "reuters.com", "indeed.com"])
for r in run["results"]:
    print(r["domain"], r["category"], r["source"], r["snapshot_timestamp"])
# etsy.com     shopping   archive  20260820065309
# reuters.com  news       archive  20260816153131
# indeed.com   jobsearch  live     None

Which domains hit a wall changes week to week, so source is worth reading rather than assuming. A capture older than archive_max_age_days (default 365) is refused rather than passed off as the live page, and those domains report cannot_classify. Set archive_fallback=False to turn this off and have blocked domains report bot_blocked.

Crawling politely

The fetcher reads robots.txt through protego, Scrapy's parser, and obeys it before making any other request. It throttles per host and bounds concurrency. Its user-agent names the package and carries a contact URL.

Robots failures are directional. A 5xx or an unparseable robots body fails closed, while an unreachable host fails open, so you get the real dns_error rather than a claim that the host refused you.

Historical analysis

old_run = classifier.classify(["facebook.com"], archive_date="20100101")

from piedomains import DataCollector

collector = DataCollector(archive_date="20050101")
collection = collector.collect_batch(["google.com", "cnn.com"], batch_size=10)
results = classifier.classify_from_collection(collection, method="text")

Snapshot discovery and retrieval go through the wayback library. Only status-200 captures are used, so an archived redirect or 404 is never classified as if it were content; the domain reports no_archive_snapshot instead. The capture actually used comes back as snapshot_timestamp, which is not necessarily the date you asked for: requesting 20100101 for cnn.com yields 20100101041727.

Text is fetched raw through Wayback's id_ playback mode, so there is no injected Wayback JavaScript, no rewritten URLs, and no browser. Screenshots render through if_, which hides the Wayback toolbar but keeps archived CSS and images so the page looks as it did.

The cache key includes the archive date, so a live fetch and snapshots from different years coexist rather than overwriting one another:

cache/html/cnn.com.html            # live
cache/html/cnn.com@20050101.html   # 2005 snapshot
cache/html/cnn.com@20150101.html   # 2015 snapshot

Rate limits, retries, and backoff are configurable through piedomains.config: archive_max_parallel, archive_window_days, archive_search_rate, archive_memento_rate, archive_retries, archive_backoff.

Other ways to classify

Text is the default and the most accurate option available here.

run = classifier.classify_by_text(["news.google.com"])

Screenshots are opt-in and weaker. The image checkpoint reports 0.501 accuracy on its own held-out split against the text checkpoint's 0.818, and lower still on pages captured from the web as it looks today. Four ways of combining the two were measured and all four came out worse than text alone, so the ensemble was built and not shipped. Reach for screenshots when there is no text to read, which is the case they exist for.

run = classifier.classify(["github.com"], use_screenshots=True)
run = classifier.classify_by_images(["github.com"])

An LLM can classify into your own label set instead of the built-in 44, which is the escape hatch when the taxonomy here does not fit your question:

classifier.configure_llm(
    provider="openai",
    model="gpt-4o",
    categories=["news", "shopping", "social", "tech"],
)
run = classifier.classify_by_llm(["example.com"])
run = classifier.classify_by_llm(
    ["site.com"], custom_instructions="Classify by educational value"
)

Set OPENAI_API_KEY, ANTHROPIC_API_KEY, or GOOGLE_API_KEY in the environment.

To fetch once and classify several ways, separate collection from inference:

collector = DataCollector()
collection = collector.collect_batch(domains, batch_size=50)
results = classifier.classify_from_collection(collection, method="text")

Categories

44 labels: adult, alcohol, automobile, cooking, dating, downloads, drugs, education, finance, fortunetelling, forum, gamble, games, gardening, government, homestyle, hospitals, humor, imagehosting, isp, jobsearch, library, military, movies, music, news, parked, pets, politics, radiotv, realestate, religion, restaurants, science, searchengines, shopping, socialnet, sports, travel, unavailable, urlshortener, weapons, webmail, wellness.

parked and unavailable describe domains with no site behind them.

The set derives from Shallalist, with one rule deciding every case: is the category visible in the page text? Classes describing how a site is built and paid for rather than what it says (adv, tracker, spyware, redirector) are gone, because a homepage selling handmade goods reads identically whether or not its operator runs trackers. See piedomains.training.taxonomy for the reasoning on each one.

The taxonomy is the live problem rather than a settled foundation. It does not line up with Curlie or with the other public schemes, so any number comparing them rests on a mapping that has not been audited, and some boundaries it draws are hard to answer from a homepage alone. If you have a labeled set, or an opinion about where the seams should be, that is the most useful thing you could contribute.

Running in a container

docker build -t piedomains-sandbox .

docker run --rm --memory=2g --cpus=2 --read-only \
  --tmpfs /tmp --tmpfs /var/tmp \
  piedomains-sandbox python -c "
from piedomains import DomainClassifier
run = DomainClassifier().classify(['example.com'])
print(run['results'][0]['category'])
"

examples/sandbox/secure_classify.py runs a batch under the same constraints.

Retraining and evaluation

The scripts that built and scored the checkpoints ship with the package:

classify_domains --training-scripts   # prints where they are installed

Links

API reference | Examples | Changelog | Sandbox guide

Development

git clone https://github.com/themains/piedomains
cd piedomains
uv sync --all-groups
uv run pytest tests/ -v

License

MIT

Citation

@software{piedomains,
  title={piedomains: classify website content from homepage text},
  author={Chintalapati, Rajashekar and Sood, Gaurav},
  year={2026},
  url={https://github.com/themains/piedomains}
}

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

piedomains-0.14.0.tar.gz (217.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

piedomains-0.14.0-py3-none-any.whl (252.1 kB view details)

Uploaded Python 3

File details

Details for the file piedomains-0.14.0.tar.gz.

File metadata

  • Download URL: piedomains-0.14.0.tar.gz
  • Upload date:
  • Size: 217.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for piedomains-0.14.0.tar.gz
Algorithm Hash digest
SHA256 b2b8dfad0bbfbf25eadfc21d3d817976a97d7f6b87a1f2b4915423f397eee2a7
MD5 58b14914c55ff503b306515bfaf41b57
BLAKE2b-256 b9a5cb4bf100bc7d6d8ef1d756c17cb15635243e7d37c5f22c792d9437b95356

See more details on using hashes here.

Provenance

The following attestation bundles were made for piedomains-0.14.0.tar.gz:

Publisher: release.yml on themains/piedomains

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file piedomains-0.14.0-py3-none-any.whl.

File metadata

  • Download URL: piedomains-0.14.0-py3-none-any.whl
  • Upload date:
  • Size: 252.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for piedomains-0.14.0-py3-none-any.whl
Algorithm Hash digest
SHA256 10c2fbf10e4eebf800eceba85abb35e250d6c79f1d33cd210242f35864767707
MD5 be28fd32f5f6d8587c482109acf24275
BLAKE2b-256 44ff38d4b28891b4c066042e816cf944148ebfae479c905fd1cb1ea5e330a9e7

See more details on using hashes here.

Provenance

The following attestation bundles were made for piedomains-0.14.0-py3-none-any.whl:

Publisher: release.yml on themains/piedomains

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.14.0 This release

2 files

0.5.0

2 files

0.3.10

2 files

0.0.19

2 files

0.0.11

2 files

0.0.4

2 files

0.0.1

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page