Protego
Protego is a pure-Python robots.txt parser. It implements the parsing and URL matching rules of RFC 9309, and additionally supports the Crawl-delay, Request-rate, Visit-time and Host extensions.
Fetching robots.txt is up to you, and so are the parts of RFC 9309 that govern it, such as the handling of HTTP status codes and redirects, caching, and imposing a parsing limit.
Install
To install Protego, simply use pip:
pip install protego
Usage
>>> from protego import Protego
>>> robotstxt = """
... User-agent: *
... Disallow: /
... Allow: /about
... Allow: /account
... Disallow: /account/contact$
... Disallow: /account/*/profile
... Crawl-delay: 4
... Request-rate: 10/1m # 10 requests every 1 minute
...
... Sitemap: http://example.com/sitemap-index.xml
... Host: http://example.co.in
... """
>>> rp = Protego.parse(robotstxt)
>>> rp.can_fetch("http://example.com/profiles", "mybot")
False
>>> rp.can_fetch("http://example.com/about", "mybot")
True
>>> rp.can_fetch("http://example.com/account", "mybot")
True
>>> rp.can_fetch("http://example.com/account/myuser/profile", "mybot")
False
>>> rp.can_fetch("http://example.com/account/contact", "mybot")
False
>>> rp.crawl_delay("mybot")
4.0
>>> rp.request_rate("mybot")
RequestRate(requests=10, seconds=60, start_time=None, end_time=None)
>>> list(rp.sitemaps)
['http://example.com/sitemap-index.xml']
>>> rp.preferred_host
'http://example.co.in'
Using Protego with Requests:
>>> from protego import Protego
>>> import requests
>>> r = requests.get("https://google.com/robots.txt")
>>> rp = Protego.parse(r.text)
>>> rp.can_fetch("https://google.com/search", "mybot")
False
>>> rp.can_fetch("https://google.com/search/about", "mybot")
True
>>> list(rp.sitemaps)
['https://www.google.com/sitemap.xml']
Comparison
The following table compares Protego to the most popular robots.txt parsers implemented in Python. Performance is the speed difference against Protego, so a positive value means faster than Protego. It is measured over the robots.txt of 100 of the most visited websites: the time taken to check the URLs their homepages link to, and the time taken to parse the files themselves.
Protego |
RobotFileParser |
robotspy |
Robotexclusionrulesparser |
|
|---|---|---|---|---|
Version tested |
Python 3.14.7 |
0.13.0 |
1.7.1 |
|
Reference specification |
||||
✓ |
✓ |
✓ |
✓ |
|
✓ |
✓ |
✓ |
||
Crawl-delay |
✓ |
✓ |
||
Request-rate |
✓ |
|||
Visit-time |
✓ |
|||
Sitemaps |
✓ |
✓ |
✓ |
✓ |
Host |
✓ |
|||
Matching performance |
-62% |
-74% |
-96% |
|
Parsing performance |
-71% |
+39% |
+56% |
API Reference
Class protego.Protego:
Properties
sitemaps {list_iterator} A list of sitemaps specified in robots.txt.
preferred_host {string} Preferred host specified in robots.txt.
Methods
parse(robotstxt_body) Parse robots.txt and return a new instance of protego.Protego.
can_fetch(url, user_agent) Return True if the user agent can fetch the URL, otherwise return False.
user_agent may be a product token, such as "mybot", or a whole User-Agent header value, such as "Mozilla/5.0 (compatible; mybot/1.0)"; a group applies when its product token appears in user_agent at a token boundary.
crawl_delay(user_agent) Return the crawl delay specified for the user agent as a float. If nothing is specified, return None.
request_rate(user_agent) Return the request rate specified for the user agent as a named tuple RequestRate(requests, seconds, start_time, end_time). If nothing is specified, return None.
visit_time(user_agent) Return the visit time specified for the user agent as a named tuple VisitTime(start_time, end_time). If nothing is specified, return None.
Metadata
Release files for Protego 0.7.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| protego-0.7.0.tar.gz | 3.1 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| protego-0.7.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 3.2 MB
Release files / protego-0.7.0.tar.gz
| Download URL | protego-0.7.0.tar.gz |
|---|---|
| Size | 3.1 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
2c032d9736a1f4f0c4318f3558353ae34da5cd038f1a5e064ded7548df315e5a
|
|
BLAKE2b-256 checksum How to use checksums |
7ad95026b9e75db1172f02441a84eaf42efb199b4cea14dda7651a620d1acd40
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.
Transparency logRelease files / protego-0.7.0-py3-none-any.whl
| Download URL | protego-0.7.0-py3-none-any.whl |
|---|---|
| Size | 13.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
944638aaee608f6c83fcad442695f3f9381a60be8d5007248f46b00f201fa5fc
|
|
BLAKE2b-256 checksum How to use checksums |
c020e4382a85e4cf4746d11746481f83360b990165725dd36ba73d55e029e6a9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.
Transparency log