Skip to main content

urlps

Lightweight, secure URL parsing and building library with RFC 3986 compliance. Features comprehensive security protections including SSRF prevention, DNS rebinding detection, path traversal protection, and homograph attack detection.

Installation

pip install urlps

Development setup:

python -m venv .venv
. .venv/Scripts/activate  # Windows: .venv\Scripts\activate
pip install -e ".[dev]"

Search cleanup: repository searches can use .rgignore to skip local/IDE/build artifacts.

Quick Start

from urlps import parse_url, build

# Secure by default - blocks SSRF, private IPs, localhost
url = parse_url("https://api.example.com/data?token=abc#section")
print(url.host)  # api.example.com
print(url.query_params)  # [("token", "abc")]

# Build URLs
url_str = build("https", "example.com", port=8443, path="/api", query="x=1")
# https://example.com:8443/api?x=1

# Immutable with functional updates
url = parse_url("https://example.com/path")
new_url = url.with_host("other.com").with_port(8080)
print(new_url)  # https://other.com:8080/path

# Policy-based validation (policy="strict" is the default; shown explicitly
# here -- see "Security" below for what it blocks and when to relax it)
strict_url = parse_url("https://example.com", policy="strict")

Security

parse_url() defaults to policy="strict". Rejection is reserved for input that is genuinely malformed or genuinely dangerous -- cosmetic differences are normalized, not rejected, so HTTP://EXAMPLE.COM./, https://example.com:443/, /a/./b and %7E all parse fine and resolve to one canonical form. That is what makes url.host safe to compare against an allowlist directly.

Blocked under both strict and balanced:

  • SSRF -- private IPs (10.x, 172.16-31.x, 192.168.x), loopback, link-local (169.254.x), .local/.internal, cloud metadata endpoints (169.254.169.254, metadata.google.internal), kubernetes service names, and obfuscated spellings (decimal 2130706433, octal, hex, IPv4-mapped IPv6, NAT64)
  • Path traversal -- ../, null bytes, and encoded variants
  • Open redirect -- leading //, backslashes, raw or percent-encoded
  • Double-encoded characters -- %25xx filter bypass
  • Parser confusion -- URLs that different parsers disagree about
  • Homograph attacks -- mixed scripts and whole-script confusables, evaluated per label on the Punycode-decoded host
  • Invisible characters -- bidi controls, zero-width, malformed Punycode

Opt-in (off by default; enable via SecurityPolicy):

  • block_dangerous_ports -- SSRF already covers the internal-service case, and blocking port 22 on a public host prevents nothing
  • reject_credentials -- user:pass@host is legal RFC 3986; a non-blocking advisory finding is emitted regardless

policy="balanced" differs from strict only in those two opt-in checks, so in practice the two are very close. Use balanced when you intend to inspect and canonicalize URLs yourself rather than reject them outright:

from urlps import parse_url

balanced_url = parse_url("HTTP://EXAMPLE.com", policy="balanced")

Use parse_url_local() for development and internal URLs. It turns the heuristic checks off and permits loopback/RFC1918 hosts, but narrows SSRF enforcement rather than disabling it -- cloud metadata endpoints, the link-local range, .internal and kubernetes service names stay blocked:

from urlps import SecurityPolicy, parse_url_local

dev_url = parse_url_local("http://localhost:3000/api")
internal = parse_url_local("http://192.168.1.100/metrics")

# If policy is passed, parse_url_local uses it exactly.
trusted_policy = SecurityPolicy.local(check_dns=True)
internal_checked = parse_url_local("http://intranet.local/service", policy=trusted_policy)

parse_url_unsafe() is the former name for the same function. It still works but is deprecated -- it emits a DeprecationWarning and will be removed in a future major release; use parse_url_local() instead.

Need to adjust the tradeoff? Use policy presets:

  • policy="strict" (default): maximum protections, DNS connect checks fail-closed by default
  • policy="balanced": fewer false positives, DNS connect checks fail-open by default
  • policy="internal": trusted traffic -- heuristics off, SSRF still enforced
  • policy="local": development -- heuristics off, loopback/private hosts allowed, metadata endpoints still blocked

To genuinely disable SSRF enforcement you must say so explicitly: SecurityPolicy.internal(enforce_ssrf=False).

DNS connect behavior can be customized per policy:

from urlps import SecurityPolicy, parse_url

policy = SecurityPolicy.strict(check_dns=True, dns_fail_open_on_connect_error=True)
url = parse_url("https://api.example.com", policy=policy)

Recommended for multi-tenant or concurrent applications: inject a dedicated DNS limiter.

from urlps import DNSRateLimiter, DNSRateLimiterConfig, parse_url

limiter = DNSRateLimiter(
    DNSRateLimiterConfig(max_lookups_per_second=20, max_lookups_per_host=50)
)

url = parse_url(
    "https://api.example.com",
    policy="strict",
    check_dns=True,
    dns_rate_limiter=limiter,
)

Core Features

Immutable URL Objects

from urlps import parse_url

url = parse_url("https://user:pass@example.com:8080/path?token=abc", policy="balanced")
print(url.netloc)         # user:pass@example.com:8080
print(url.effective_port) # 8080

# with_* methods return new URL objects
url2 = url.with_netloc("admin@example.com")
url3 = url.with_host("other.com").with_port(443).with_path("/api")
url4 = url.with_query_param("new", "value")
url5 = url.without_query_param("token")

# Genuinely immutable, not immutable by convention -- a URL that passed
# validation cannot be re-pointed afterwards.
try:
    url._host = "evil.com"
except AttributeError as exc:
    print(exc)  # URL is immutable; use with_*() or copy() to derive a new URL ...

assert url.host == "example.com"

Query strings round-trip exactly

Parsing never rewrites the query. This matters if you verify signatures over a raw query string, or proxy URLs onward:

from urlps import parse_url

url = parse_url("https://api.example.com/search?sig=aGVsbG8%3D&q=a+%26+b")
print(url.query)         # sig=aGVsbG8%3D&q=a+%26+b   (byte-for-byte)
print(str(url) )         # ...unchanged...
print(url.query_params)  # [('sig', 'aGVsbG8='), ('q', 'a & b')]

q=a+%26+b is one parameter whose value contains &. Re-encoding is only performed when you explicitly change the query (with_query_param(), canonicalize()), and never turns one parameter into two.

Reference resolution (RFC 3986 §5)

join() is the security-preserving equivalent of urllib.parse.urljoin — the resolved target is validated, so resolution can't be used to slip past the checks parse_url() applies:

from urlps import join

join("https://example.com/a/b", "../c")     # https://example.com/c
join("https://example.com/a/b", "?q=1")     # https://example.com/a/b?q=1
join("https://example.com/a/b", "#frag")    # https://example.com/a/b#frag

# '..' can never escape the authority
join("https://example.com/a/b", "../../../../etc/passwd")
# https://example.com/etc/passwd

# A protocol-relative reference legitimately replaces the host, which is
# exactly why the *result* is re-validated rather than trusted:
join("https://example.com/a/", "//localhost/admin")   # raises InvalidURLError

Security Checks

from urlps import parse_url, InvalidURLError

# SSRF protection (enabled by default)
try:
    parse_url("http://localhost/admin")  # Blocked
except InvalidURLError as e:
    print(f"Rejected: {e}")

# DNS rebinding detection (optional - rate-limited to prevent DoS)
url_dns = parse_url("https://api.example.com/", check_dns=True)

# URL canonicalization (policy="balanced": the raw non-canonical/credentialed
# forms below are exactly what strict's require_canonical/reject_credentials
# would block -- use balanced when you want to parse first and canonicalize
# after, rather than reject upfront)
url_raw = parse_url("HTTP://EXAMPLE.COM:80/path?z=1&a=2", policy="balanced")
canonical = url_raw.canonicalize()
print(canonical.scheme)  # "http"
print(canonical.host)    # "example.com"
print(canonical.port)    # None (default port removed)
print(canonical.query)   # "a=2&z=1" (sorted)

# Password masking
url = parse_url("https://admin:secret123@api.example.com/", policy="balanced")
print(url.as_string(mask_password=True))  # https://admin:***@api.example.com/

Using check_dns/check_phishing from asyncio code

check_dns=True and check_phishing=True both do blocking network I/O (DNS resolution plus a verification connect; a synchronous HTTP download, respectively). Calling either directly inside an async def request handler blocks the event loop for the duration of that call -- exactly the kind of mistake that's easy to make in a FastAPI/aiohttp/Starlette handler doing SSRF-guarded URL validation, which is precisely where this library is most useful. Run them in an executor instead:

import asyncio
import functools

from urlps import parse_url


async def parse_url_async(url, **kwargs):
    loop = asyncio.get_running_loop()
    return await loop.run_in_executor(None, functools.partial(parse_url, url, **kwargs))


async def main():
    url = await parse_url_async("https://api.example.com/data", check_dns=True)
    print(url.host)  # api.example.com


asyncio.run(main())

The same pattern applies to check_phishing=True and to build_secure(). There is no bundled async wrapper -- run_in_executor (or an equivalent from your framework, e.g. starlette.concurrency.run_in_threadpool) is sufficient and avoids committing this library to a specific async runtime.

Audit Logging

Audit callbacks are supplied per call via AuditConfig, so different callers can log differently without sharing global state:

import logging
from urlps import AuditConfig, parse_url

def audit_url_parsing(logged_url, parsed_url, exception):
    if exception:
        logging.warning(f"Failed to parse URL: {exception}")
    else:
        logging.info(f"Parsed URL to host: {parsed_url.host}")

url = parse_url(
    "https://api.example.com/data",
    audit=AuditConfig(callback=audit_url_parsing),
)

Structured event callback:

from urlps import AuditConfig, parse_url

def on_event(event):
    # event includes: timestamp, level, operation, raw_url, host,
    # error_type, error_code, correlation_id
    print(event)

url = parse_url(
    "https://api.example.com/data",
    correlation_id="request-42",
    audit=AuditConfig(event_callback=on_event),
)

URLs are redacted before being passed to callbacks (credentials and sensitive query values are masked). Pass AuditConfig(..., redact_urls=False) to opt out. A callback that raises is recorded as a failure and never breaks the parse.

The same audit= parameter is accepted by parse_url_local(), join() and build_secure().

Component Length Limits

Conservative limits to prevent DoS attacks:

Component Max Length
URL (total) 32 KB
Scheme 16 chars
Host 253 chars
Path 4 KB
Query 8 KB
Fragment 1 KB
Userinfo 128 chars

Environment Variables

Override length limits via environment variables:

# PowerShell
$env:URLPS_MAX_URL_LENGTH = "65536"
python -c "import urlps.constants as c; print(c.MAX_URL_LENGTH)"

# Bash
export URLPS_MAX_URL_LENGTH=65536
python -c 'import urlps.constants as c; print(c.MAX_URL_LENGTH)'

Supported variables:

  • URLPS_MAX_URL_LENGTH
  • URLPS_MAX_SCHEME_LENGTH
  • URLPS_MAX_HOST_LENGTH
  • URLPS_MAX_PATH_LENGTH
  • URLPS_MAX_QUERY_LENGTH
  • URLPS_MAX_FRAGMENT_LENGTH
  • URLPS_MAX_USERINFO_LENGTH
  • URLPS_MAX_IPV6_STRING_LENGTH
  • URLPS_PHISHING_DATABASE_URL -- overrides the feed check_phishing=True downloads hostnames from (default: a third-party list at phish.co.za). Must be an http:// or https:// URL; set this to self-host the list or point at a mirror you trust instead. Must be set before import urlps, same as the cache-size variables below.

Internal @lru_cache sizes are also overridable this way -- see Cache Sizing below.

API Reference

Main Functions

Function Description
parse_url(url, *, allow_custom_scheme=False, check_dns=False, check_phishing=False, dns_rate_limiter=None, policy=None, correlation_id=None, audit=None) Parse URL with policy-aware security checks (recommended)
parse_url_local(url, *, allow_custom_scheme=False, debug=False, check_dns=False, dns_rate_limiter=None, policy=None, correlation_id=None, audit=None) Parse URL for trusted/internal input with optional policy overrides
parse_url_unsafe(...) Deprecated alias for parse_url_local(); emits DeprecationWarning
join(base, reference, *, policy=None, strict_resolution=True, ...) Resolve a reference against a base URI (RFC 3986 §5), then validate
build(*scheme_and_host, port=None, path="/", query=None, fragment=None, userinfo=None) Build URL string from components
build_secure(*scheme_and_host, policy=None, check_dns=False, check_phishing=False, dns_rate_limiter=None, correlation_id=None, audit=None, ...) Build and then validate a URL under a selected security policy
compose_url(components) Build URL from components dict

build()/compose_url() validate that host is a syntactically valid hostname, IPv4 literal, or bracketed IPv6 literal (raising HostValidationError otherwise), so their output always round-trips through parse_url(). This is structural validation only, not security policy -- use build_secure() when the host isn't already trusted.

Note: get_dns_rate_limiter() and reset_dns_rate_limiter() remain available for compatibility, but explicit dns_rate_limiter= injection is preferred.

URL Methods

Method Description
url.as_string(mask_password=False) Convert to string, optionally masking password
url.canonicalize() Return canonicalized copy
url.is_semantically_equal(other) Compare URLs by meaning after canonicalization
url.same_origin(other) Check if URLs have same origin
url == "https://...", <, <=, >, >= Compare a URL directly against another URL or a plain string, against as_string() (no canonicalization)
url.origin Return origin string (e.g., https://example.com)
url.copy(**overrides) Create copy with optional component overrides
url.with_*() Functional updates: with_scheme, with_host, with_port, with_path, with_fragment, with_userinfo, with_netloc, with_query_param, without_query_param
url.get_query_param(key, default=None) First value for key, or default if absent (None for a present, value-less key like ?flag)
url.get_query_param_all(key) All values for key, in order; [] if absent

Cache Management

from urlps import get_cache_info, clear_all_caches

# Get cache statistics
stats = get_cache_info()
print(stats['parser']['normalize_path']['hits'])

# Clear all caches (useful for long-running apps)
previous = clear_all_caches()

Cache Sizing

Every internal @lru_cache (security/host checks, validation predicates, path normalization, percent-encoding, policy resolution) is sized from a small set of environment variables, all read once, at import time. Python's functools.lru_cache bakes maxsize in when the decorated function is defined, so these must be set before import urlps runs -- setting them afterwards, or after the first parse_url() call, has no effect:

import os
os.environ["URLPS_CACHE_SIZE_SECURITY"] = "8192"
import urlps  # cache sizes are now locked in for this process
Variable Default Covers
URLPS_CACHE_SIZE_SECURITY 512 is_ssrf_risk, is_private_ip, has_parser_confusion, has_mixed_scripts, find_authority_marker -- keyed on host or full URL
URLPS_CACHE_SIZE_VALIDATION 512 is_valid_host, is_valid_scheme, is_url_safe_string, is_valid_fragment, etc.
URLPS_CACHE_SIZE_PARSER 1024 normalize_path
URLPS_CACHE_SIZE_BUILDER_QUERY_ENCODE 8192 percent-encoding of query keys/values
URLPS_CACHE_SIZE_BUILDER_PATH_ENCODE 1024 percent-encoding of path segments
URLPS_CACHE_SIZE_POLICY 16 resolved named policies (strict/balanced/internal x overrides) -- rarely worth changing, the working set is inherently tiny

Which value fits your workload? A cache only helps when the same input (same host, same URL shape) is seen again within the cache's window -- otherwise every lookup is a miss and the cache is pure overhead with no benefit. Use get_cache_info() after a representative burst of real traffic to check hit rates before guessing:

  • Short-lived script/CLI (parses a handful of URLs and exits): defaults are fine either way -- there's rarely enough repetition for cache size to matter, and the downside of an "oversized" cache here is negligible.
  • Long-running service with a bounded set of upstream hosts (an internal proxy, a service validating callback URLs from a fixed partner list): keep the defaults, or size URLPS_CACHE_SIZE_SECURITY/_VALIDATION to comfortably exceed your distinct-host count. A cache that fits your whole working set converges to a near-100% hit rate and stays there.
  • High-diversity, public-facing workload (a crawler, a webhook receiver from many tenants, a link-checker over arbitrary user-submitted URLs): the default 512 is easy to blow through in a single request burst, at which point the cache is being evicted before it's ever reused -- raise URLPS_CACHE_SIZE_SECURITY/_VALIDATION substantially (several thousand), or accept that there may be little to gain from caching this workload at all if hosts are effectively unique per request.
  • Memory-constrained environment: lower the values, especially the builder encode caches (8192/1024 by default) if you build many large, distinct query strings -- each cache entry holds a copy of the encoded string, and eviction under memory pressure is not automatic the way it is for cache-key diversity.

Command Line

pip install urlps also installs a urlps script for validating URLs from shell scripts, CI pipelines, or pre-commit hooks, without writing a throwaway Python file:

urlps check https://example.com/path 'HTTP://EXAMPLE.COM:80/'
# https://example.com/path
# http://example.com/

urlps check http://localhost/admin
# http://localhost/admin: Host poses SSRF risk and is disallowed. ...  (to stderr)
# exit code 1

urlps check --policy local http://localhost:3000/api   # like parse_url_local()
urlps check --check-dns https://api.example.com/        # also verify DNS resolution

# One URL per line on stdin when no URL arguments are given
# ('#'-prefixed and blank lines are skipped):
cat urls.txt | urlps check

Exits 0 and prints the canonical form of every URL if all pass; exits 1 and prints the rejection reason (to stderr) for each URL that fails, alongside the canonical form of the ones that passed. --quiet suppresses success output so only failures are printed. --policy, --check-dns and --check-phishing mirror parse_url()'s own options.

Comparison with urllib.parse

Feature urllib.parse urlps
Basic URL parsing ✓ ✓
RFC 3986 strict compliance Partial ✓
SSRF protection ✗ ✓
DNS rebinding detection ✗ ✓ (with rate limiting)
Path traversal detection ✗ ✓
Homograph detection ✗ ✓
URL parser confusion protection ✗ ✓
Query parameter injection detection ✗ ✓
Dangerous port validation ✗ ✓
Canonical form validation ✗ ✓
Immutable URL objects ✗ ✓
URL canonicalization ✗ ✓
Password masking ✗ ✓
Audit logging ✗ ✓
Component length limits ✗ ✓

Use urllib.parse when: You need zero dependencies and basic parsing is sufficient.

Use urlps when: Security matters, you need RFC 3986 strict compliance, or you want immutable URL objects with ergonomic manipulation methods.

Exceptions

from urlps import InvalidURLError, URLParseError, parse_url

user_input = "https://example.com"

try:
    url = parse_url(user_input)
except URLParseError:
    print("Malformed URL")
except InvalidURLError:
    print("Rejected by security policy")

Exception hierarchy:

  • InvalidURLError — Base exception for all URL errors
  • URLParseError — Parsing errors
  • URLBuildError — Building errors
  • HostValidationError / PortValidationError — Component validation errors
  • QueryParsingError, FragmentEncodingError, UserInfoParsingError, UnsupportedSchemeError — Specific errors

Running Tests

pytest
pytest -v -k "test_parse"     # Run specific tests
pytest -m ipv6                # Run IPv6 tests
pytest -m idna                # Run IDNA tests

Changelog

See CHANGELOG.md for a summary of every release, and changelogs/ for detailed per-release notes.

License

MIT

Metadata

Release files for urlps 1.1.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for urlps 1.1.4
File Size Uploaded
urlps-1.1.4.tar.gz 109.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for urlps 1.1.4
File Interpreter ABI Platform
urlps-1.1.4-py3-none-any.whl Python 3 none any Details

Total release size: 216.5 kB

Release files / urlps-1.1.4.tar.gz

Download URL urlps-1.1.4.tar.gz
Size 109.5 kB
Tags Source
SHA-256 checksum
How to use checksums
3ed7bda07c16aeff7b8aa67f8d8250592e851be159fd67d5d1b85afa8b5787a9
BLAKE2b-256 checksum
How to use checksums
6217b92c9da986058f265916c83035a6c671d4fba91a87e8974ad796f5d975d5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 14, 2026.

Transparency log

Release files / urlps-1.1.4-py3-none-any.whl

Download URL urlps-1.1.4-py3-none-any.whl
Size 107.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ab9c43ddb0118e23c11c4a39ea5c604ebe0e814c526e4dbd8d066a120a786638
BLAKE2b-256 checksum
How to use checksums
bc534f1075047014aa64d34d5959c8ac060e74fedda4af581031ca01de168377
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 14, 2026.

Transparency log

Release history Release notifications | RSS feed

1.2.1

2 release files

1.2.0

2 release files

This release

1.1.4 This release

2 release files

1.0.0

2 release files

0.7.0

2 release files

0.6.1

2 release files

0.5.1

2 release files

0.3.5

2 release files

0.2.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page