Skip to main content

tldextract PyPI version Build Status

tldextract accurately separates a URL's subdomain, domain, and public suffix, using the Public Suffix List (PSL).

Why? Naive URL parsing like splitting on dots fails for domains like forums.bbc.co.uk (gives "co" instead of "bbc"). tldextract handles the edge cases, so you don't have to.

Quick Start

>>> import tldextract

>>> tldextract.extract('http://forums.news.cnn.com/')
ExtractResult(subdomain='forums.news', domain='cnn', suffix='com', is_private=False)

>>> tldextract.extract('http://forums.bbc.co.uk/')
ExtractResult(subdomain='forums', domain='bbc', suffix='co.uk', is_private=False)

>>> # Access the parts you need
>>> ext = tldextract.extract('http://forums.bbc.co.uk')
>>> ext.domain
'bbc'
>>> ext.top_domain_under_public_suffix
'bbc.co.uk'
>>> ext.fqdn
'forums.bbc.co.uk'

Install

pip install tldextract

How-to Guides

How to disable HTTP suffix list fetching for production

no_fetch_extract = tldextract.TLDExtract(suffix_list_urls=())
no_fetch_extract("http://www.google.com")

Or set the default for new extractors with an empty environment variable. Set it before importing tldextract to affect the module-level extract function:

export TLDEXTRACT_PUBLIC_SUFFIX_LIST_URLS=""

The cache key includes the configured URLs. An empty value does not reuse a cache entry fetched with the standard URLs; if no entry exists for the empty URL list, tldextract uses its bundled snapshot.

How to set a custom cache location

Via environment variable:

export TLDEXTRACT_CACHE="/path/to/cache"

Or in code:

custom_cache_extract = tldextract.TLDExtract(cache_dir="/path/to/cache/")

How to update TLD definitions

Command line:

tldextract --update

Or delete the cache folder:

rm -rf $HOME/.cache/python-tldextract

How to treat private domains as suffixes

extract = tldextract.TLDExtract(include_psl_private_domains=True)
extract("waiterrant.blogspot.com")
# ExtractResult(subdomain='', domain='waiterrant', suffix='blogspot.com', is_private=True)

How to use a local suffix list

extract = tldextract.TLDExtract(
    suffix_list_urls=["file:///path/to/your/list.dat"],
    cache_dir="/path/to/cache/",
    fallback_to_snapshot=False,
)

An existing local path in TLDEXTRACT_PUBLIC_SUFFIX_LIST_URLS is converted to a file:// URL, like the --suffix_list_url command line option.

How to use a remote suffix list

extract = tldextract.TLDExtract(
    suffix_list_urls=["https://myserver.com/suffix-list.dat"]
)

The default can also be set through a newline-delimited environment variable:

export TLDEXTRACT_PUBLIC_SUFFIX_LIST_URLS="https://myserver.com/suffix-list.dat"

New TLDExtract() instances read this value when constructed. An explicit suffix_list_urls argument takes precedence.

How to set the suffix list fetch timeout

Set a default timeout in seconds before starting Python:

export TLDEXTRACT_DEFAULT_FETCH_TIMEOUT="1.2"

Or set it for one extractor in code:

extract = tldextract.TLDExtract(cache_fetch_timeout=1.2)

A single value applies separately to the connection and response-read phases of each remote Public Suffix List request.

How to add extra suffixes

extract = tldextract.TLDExtract(extra_suffixes=["foo", "bar.baz"])

How to validate URLs before extraction

from urllib.parse import urlsplit

extract = tldextract.TLDExtract()
split_url = urlsplit("https://example.com/path")
result = extract.extract_urllib(split_url)

Command Line

$ tldextract http://forums.bbc.co.uk
forums bbc co.uk

$ tldextract --update  # Update cached suffix list
$ tldextract --help    # See all options

Understanding Domain Parsing

Public Suffix List

tldextract uses the Public Suffix List, a community-maintained list of domain suffixes. The PSL contains both:

  • Public suffixes: Where anyone can register a domain (.com, .co.uk, .org.kg)
  • Private suffixes: Operated by companies for customer subdomains (blogspot.com, github.io)

Web browsers use this same list for security decisions like cookie scoping.

Suffix vs. TLD

While .com is a top-level domain (TLD), many suffixes like .co.uk are technically second-level. The PSL uses "public suffix" to cover both.

Default behavior with private domains

By default, tldextract treats private suffixes as regular domains:

>>> tldextract.extract('waiterrant.blogspot.com')
ExtractResult(subdomain='waiterrant', domain='blogspot', suffix='com', is_private=False)

To treat them as suffixes instead, see How to treat private domains as suffixes.

Default behavior with unlisted suffixes

Unlike the formal PSL algorithm, tldextract does not apply the implicit * rule when no suffix matches. For an unlisted final label, suffix remains empty so callers can distinguish configured public suffixes from unknown, invalid, or internal hostnames.

Hostname representation

tldextract identifies public suffix boundaries but does not validate or canonicalize hostnames. Returned strings can retain the input's casing and its Unicode or ASCII-compatible encoding (ACE, commonly called Punycode). Equivalent IDNA hostnames can therefore produce textually different values even when tldextract identifies the same suffix boundary.

Before using suffix, fqdn, or joined-domain properties in equality checks, allowlists, blocklists, or other security decisions, validate and normalize both hostnames to the same IDNA representation. IDNA defines label equivalence in terms of A-labels. See RFC 5890, section 2.3.2.4.

Caching behavior

By default, tldextract fetches the latest Public Suffix List on first use and caches it indefinitely in $HOME/.cache/python-tldextract.

URL validation

tldextract accepts any string and is very lenient. It prioritizes ease of use over strict validation, extracting domains from any string, even partial URLs or non-URLs.

FAQ

Can you add/remove suffix ____?

tldextract doesn't maintain the suffix list. Submit changes to the Public Suffix List.

Meanwhile, use the extra_suffixes parameter, or fork the PSL and pass it to this library with the suffix_list_urls parameter.

My suffix is in the PSL but not extracted correctly

Check if it's in the "PRIVATE" section. See How to treat private domains as suffixes.

Why does it parse invalid URLs?

See URL validation and How to validate URLs before extraction.

Contribute

Setting up

  1. Install uv.
  2. git clone this repository.
  3. Change into the new directory.
  4. uv sync

Running tests

uv run pytest          # Fast local suite
uv run tox -e py310    # Supported-version suite
uv run tox --parallel  # Full matrix
uv run ruff format .   # Format code

History

This package started from a StackOverflow answer about regex-based domain extraction. The regex approach fails for many domains, so this library switched to the Public Suffix List for accuracy.

Metadata

Release files for tldextract 5.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tldextract 5.4.0
File Size Uploaded
tldextract-5.4.0.tar.gz 197.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tldextract 5.4.0
File Interpreter ABI Platform
tldextract-5.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 304.8 kB

Release files / tldextract-5.4.0.tar.gz

Download URL tldextract-5.4.0.tar.gz
Size 197.3 kB
Tags Source
SHA-256 checksum
How to use checksums
6c9223212c15c25c0da2bf7313893c14f175cb36b64a0c42da67a468e0c61ee3
BLAKE2b-256 checksum
How to use checksums
fd5d45ece871390ccc985f821353543165bcf3784fa97d8484fd0ca5f2726612
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.16

Release files / tldextract-5.4.0-py3-none-any.whl

Download URL tldextract-5.4.0-py3-none-any.whl
Size 107.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7f02aed30bd3b6ad5717192eb859a39b20aafc7caf3917d9cf6cb00a58efb34f
BLAKE2b-256 checksum
How to use checksums
b8e0d5760e222a7e3f3aec7ef59f6dff952af27d09ec168bcca50751d9698151
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.16

Release history Release notifications | RSS feed

This release

5.4.0 This release

2 release files

5.3.2

2 release files

5.3.1

2 release files

5.3.0

2 release files

5.2.0

2 release files

5.1.3

2 release files

5.1.2

2 release files

5.1.1

2 release files

5.1.0

2 release files

5.0.1

2 release files

5.0.0

2 release files

4.0.0

2 release files

3.5.0

2 release files

3.4.4

2 release files

3.4.3

2 release files

3.4.2

2 release files

3.4.1

2 release files

3.4.0

2 release files

3.3.1

2 release files

3.3.0

2 release files

3.2.1

2 release files

3.2.0

2 release files

3.1.2

2 release files

3.1.1

2 release files

3.1.0

2 release files

3.0.2

2 release files

3.0.1

2 release files

3.0.0

2 release files

2.2.3

2 release files

2.2.2

2 release files

2.2.1

2 release files

2.2.0

2 release files

2.1.0

2 release files

2.0.3

2 release files

2.0.2

2 release files

2.0.1

1 release file

2.0.0

1 release file

1.7.5

1 release file

1.7.4

1 release file

1.7.3

1 release file

1.7.2

1 release file

1.7.1

1 release file

1.7

1 release file

1.6

1 release file

1.5.1

1 release file

1.5

1 release file

1.4

1 release file

1.3.1

1 release file

1.3

1 release file

1.2.2

1 release file

1.2.1

1 release file

1.2

1 release file

1.1.3

1 release file

1.1.2

1 release file

1.1.1

1 release file

1.1

1 release file

1.0

1 release file

0.4

1 release file

0.3.2

1 release file

0.3.1

1 release file

0.3

1 release file

0.2

1 release file

0.1.1

1 release file

0.1

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page