Accurately separates a URL’s subdomain, domain, and public suffix, using the Public Suffix List (PSL).
Python
2,026
574 commits
updated Sep 29, 2026
tldextract accurately separates a URL's subdomain, domain, and public suffix,
using the Public Suffix List (PSL).
Why? Naive URL parsing like splitting on dots fails for domains like
forums.bbc.co.uk (gives "co" instead of "bbc"). tldextract handles the edge
cases, so you don't have to.
>>> import tldextract
>>> tldextract.extract('http://forums.news.cnn.com/')
ExtractResult(subdomain='forums.news', domain='cnn', suffix='com', is_private=False)
>>> tldextract.extract('http://forums.bbc.co.uk/')
ExtractResult(subdomain='forums', domain='bbc', suffix='co.uk', is_private=False)
>>> # Access the parts you need
>>> ext = tldextract.extract('http://forums.bbc.co.uk')
>>> ext.domain
'bbc'
>>> ext.top_domain_under_public_suffix
'bbc.co.uk'
>>> ext.fqdn
'forums.bbc.co.uk'
pip install tldextract
no_fetch_extract = tldextract.TLDExtract(suffix_list_urls=())
no_fetch_extract("http://www.google.com")
Or set the default for new extractors with an empty environment variable. Set
it before importing tldextract to affect the module-level extract function:
export TLDEXTRACT_PUBLIC_SUFFIX_LIST_URLS=""
The cache key includes the configured URLs. An empty value does not reuse a
cache entry fetched with the standard URLs; if no entry exists for the empty
URL list, tldextract uses its bundled snapshot.
Via environment variable:
export TLDEXTRACT_CACHE="/path/to/cache"
Or in code:
custom_cache_extract = tldextract.TLDExtract(cache_dir="/path/to/cache/")
Command line:
tldextract --update
Or delete the cache folder:
rm -rf $HOME/.cache/python-tldextract
extract = tldextract.TLDExtract(include_psl_private_domains=True)
extract("waiterrant.blogspot.com")
# ExtractResult(subdomain='', domain='waiterrant', suffix='blogspot.com', is_private=True)
extract = tldextract.TLDExtract(
suffix_list_urls=["file:///path/to/your/list.dat"],
cache_dir="/path/to/cache/",
fallback_to_snapshot=False,
)
An existing local path in TLDEXTRACT_PUBLIC_SUFFIX_LIST_URLS is converted to
a file:// URL, like the --suffix_list_url command line option.
extract = tldextract.TLDExtract(
suffix_list_urls=["https://myserver.com/suffix-list.dat"]
)
The default can also be set through a newline-delimited environment variable:
export TLDEXTRACT_PUBLIC_SUFFIX_LIST_URLS="https://myserver.com/suffix-list.dat"
New TLDExtract() instances read this value when constructed. An explicit
suffix_list_urls argument takes precedence.
Set a default timeout in seconds before starting Python:
export TLDEXTRACT_DEFAULT_FETCH_TIMEOUT="1.2"
Or set it for one extractor in code:
extract = tldextract.TLDExtract(cache_fetch_timeout=1.2)
A single value applies separately to the connection and response-read phases of each remote Public Suffix List request.
extract = tldextract.TLDExtract(extra_suffixes=["foo", "bar.baz"])
from urllib.parse import urlsplit
extract = tldextract.TLDExtract()
split_url = urlsplit("https://example.com/path")
result = extract.extract_urllib(split_url)
$ tldextract http://forums.bbc.co.uk
forums bbc co.uk
$ tldextract --update # Update cached suffix list
$ tldextract --help # See all options
tldextract uses the Public Suffix List, a
community-maintained list of domain suffixes. The PSL contains both:
.com, .co.uk,
.org.kg)blogspot.com, github.io)Web browsers use this same list for security decisions like cookie scoping.
While .com is a top-level domain (TLD), many suffixes like .co.uk are
technically second-level. The PSL uses "public suffix" to cover both.
By default, tldextract treats private suffixes as regular domains:
>>> tldextract.extract('waiterrant.blogspot.com')
ExtractResult(subdomain='waiterrant', domain='blogspot', suffix='com', is_private=False)
To treat them as suffixes instead, see How to treat private domains as suffixes.
Unlike the formal PSL algorithm,
tldextract does not apply the implicit * rule when no suffix matches. For an
unlisted final label, suffix remains empty so callers can distinguish configured
public suffixes from unknown, invalid, or internal hostnames.
tldextract identifies public suffix boundaries but does not validate or
canonicalize hostnames. Returned strings can retain the input's casing and its
Unicode or ASCII-compatible encoding (ACE, commonly called Punycode).
Equivalent IDNA hostnames can therefore produce textually different values even
when tldextract identifies the same suffix boundary.
Before using suffix, fqdn, or joined-domain properties in equality checks,
allowlists, blocklists, or other security decisions, validate and normalize
both hostnames to the same IDNA representation. IDNA defines label equivalence
in terms of A-labels. See
RFC 5890, section 2.3.2.4.
By default, tldextract fetches the latest Public Suffix List on first use and
caches it indefinitely in $HOME/.cache/python-tldextract.
tldextract accepts any string and is very lenient. It prioritizes ease of use
over strict validation, extracting domains from any string, even partial URLs or
non-URLs.
tldextract doesn't maintain the suffix list. Submit changes to
the Public Suffix List.
Meanwhile, use the extra_suffixes parameter, or fork the PSL and pass it to
this library with the suffix_list_urls parameter.
Check if it's in the "PRIVATE" section. See How to treat private domains as suffixes.
See URL validation and How to validate URLs before extraction.
git clone this repository.uv syncuv run pytest # Fast local suite
uv run tox -e py310 # Supported-version suite
uv run tox --parallel # Full matrix
uv run ruff format . # Format code
This package started from a StackOverflow answer about regex-based domain extraction. The regex approach fails for many domains, so this library switched to the Public Suffix List for accuracy.
2,129 followers · starred May 2024
1,348 followers · starred Feb 2011
70 followers · starred Mar 2018
6 followers · starred Dec 2014
Python
100.0%
Accurately separates a URL’s subdomain, domain, and public suffix, using the Public Suffix List (PSL).
Python
2,026
574 commits
updated Sep 29, 2026
tldextract accurately separates a URL's subdomain, domain, and public suffix,
using the Public Suffix List (PSL).
Why? Naive URL parsing like splitting on dots fails for domains like
forums.bbc.co.uk (gives "co" instead of "bbc"). tldextract handles the edge
cases, so you don't have to.
>>> import tldextract
>>> tldextract.extract('http://forums.news.cnn.com/')
ExtractResult(subdomain='forums.news', domain='cnn', suffix='com', is_private=False)
>>> tldextract.extract('http://forums.bbc.co.uk/')
ExtractResult(subdomain='forums', domain='bbc', suffix='co.uk', is_private=False)
>>> # Access the parts you need
>>> ext = tldextract.extract('http://forums.bbc.co.uk')
>>> ext.domain
'bbc'
>>> ext.top_domain_under_public_suffix
'bbc.co.uk'
>>> ext.fqdn
'forums.bbc.co.uk'
pip install tldextract
no_fetch_extract = tldextract.TLDExtract(suffix_list_urls=())
no_fetch_extract("http://www.google.com")
Or set the default for new extractors with an empty environment variable. Set
it before importing tldextract to affect the module-level extract function:
export TLDEXTRACT_PUBLIC_SUFFIX_LIST_URLS=""
The cache key includes the configured URLs. An empty value does not reuse a
cache entry fetched with the standard URLs; if no entry exists for the empty
URL list, tldextract uses its bundled snapshot.
Via environment variable:
export TLDEXTRACT_CACHE="/path/to/cache"
Or in code:
custom_cache_extract = tldextract.TLDExtract(cache_dir="/path/to/cache/")
Command line:
tldextract --update
Or delete the cache folder:
rm -rf $HOME/.cache/python-tldextract
extract = tldextract.TLDExtract(include_psl_private_domains=True)
extract("waiterrant.blogspot.com")
# ExtractResult(subdomain='', domain='waiterrant', suffix='blogspot.com', is_private=True)
extract = tldextract.TLDExtract(
suffix_list_urls=["file:///path/to/your/list.dat"],
cache_dir="/path/to/cache/",
fallback_to_snapshot=False,
)
An existing local path in TLDEXTRACT_PUBLIC_SUFFIX_LIST_URLS is converted to
a file:// URL, like the --suffix_list_url command line option.
extract = tldextract.TLDExtract(
suffix_list_urls=["https://myserver.com/suffix-list.dat"]
)
The default can also be set through a newline-delimited environment variable:
export TLDEXTRACT_PUBLIC_SUFFIX_LIST_URLS="https://myserver.com/suffix-list.dat"
New TLDExtract() instances read this value when constructed. An explicit
suffix_list_urls argument takes precedence.
Set a default timeout in seconds before starting Python:
export TLDEXTRACT_DEFAULT_FETCH_TIMEOUT="1.2"
Or set it for one extractor in code:
extract = tldextract.TLDExtract(cache_fetch_timeout=1.2)
A single value applies separately to the connection and response-read phases of each remote Public Suffix List request.
extract = tldextract.TLDExtract(extra_suffixes=["foo", "bar.baz"])
from urllib.parse import urlsplit
extract = tldextract.TLDExtract()
split_url = urlsplit("https://example.com/path")
result = extract.extract_urllib(split_url)
$ tldextract http://forums.bbc.co.uk
forums bbc co.uk
$ tldextract --update # Update cached suffix list
$ tldextract --help # See all options
tldextract uses the Public Suffix List, a
community-maintained list of domain suffixes. The PSL contains both:
.com, .co.uk,
.org.kg)blogspot.com, github.io)Web browsers use this same list for security decisions like cookie scoping.
While .com is a top-level domain (TLD), many suffixes like .co.uk are
technically second-level. The PSL uses "public suffix" to cover both.
By default, tldextract treats private suffixes as regular domains:
>>> tldextract.extract('waiterrant.blogspot.com')
ExtractResult(subdomain='waiterrant', domain='blogspot', suffix='com', is_private=False)
To treat them as suffixes instead, see How to treat private domains as suffixes.
Unlike the formal PSL algorithm,
tldextract does not apply the implicit * rule when no suffix matches. For an
unlisted final label, suffix remains empty so callers can distinguish configured
public suffixes from unknown, invalid, or internal hostnames.
tldextract identifies public suffix boundaries but does not validate or
canonicalize hostnames. Returned strings can retain the input's casing and its
Unicode or ASCII-compatible encoding (ACE, commonly called Punycode).
Equivalent IDNA hostnames can therefore produce textually different values even
when tldextract identifies the same suffix boundary.
Before using suffix, fqdn, or joined-domain properties in equality checks,
allowlists, blocklists, or other security decisions, validate and normalize
both hostnames to the same IDNA representation. IDNA defines label equivalence
in terms of A-labels. See
RFC 5890, section 2.3.2.4.
By default, tldextract fetches the latest Public Suffix List on first use and
caches it indefinitely in $HOME/.cache/python-tldextract.
tldextract accepts any string and is very lenient. It prioritizes ease of use
over strict validation, extracting domains from any string, even partial URLs or
non-URLs.
tldextract doesn't maintain the suffix list. Submit changes to
the Public Suffix List.
Meanwhile, use the extra_suffixes parameter, or fork the PSL and pass it to
this library with the suffix_list_urls parameter.
Check if it's in the "PRIVATE" section. See How to treat private domains as suffixes.
See URL validation and How to validate URLs before extraction.
git clone this repository.uv syncuv run pytest # Fast local suite
uv run tox -e py310 # Supported-version suite
uv run tox --parallel # Full matrix
uv run ruff format . # Format code
This package started from a StackOverflow answer about regex-based domain extraction. The regex approach fails for many domains, so this library switched to the Public Suffix List for accuracy.
2,129 followers · starred May 2024
1,348 followers · starred Feb 2011
70 followers · starred Mar 2018
6 followers · starred Dec 2014
Python
100.0%