mudfish
Python bindings for Mudfish Crawler, a Rust-native web crawling and intelligence engine.
This package wraps the Rust crawl engine directly (via PyO3) — there is no subprocess or network hop to a separate server. It currently exposes one function, synchronous/blocking crawling; the Rust project itself is at an early stage (Phase 1: HTTP-only crawling — see the main repo's ROADMAP_HONEST.md for full status).
Install
pip install mudfish
Usage
import mudfish
result = mudfish.crawl("https://example.com", depth=2, concurrency=30)
print(result["stats"])
for page in result["pages"]:
print(page["status_code"], page["url"], page["metadata"]["title"])
crawl() blocks until the crawl finishes (or hits max_urls / max_duration_secs) and returns a plain dict:
{
"crawl_id": str,
"pages": [
{
"url": str, "final_url": str, "status_code": int,
"content_type": str | None, "depth": int,
"metadata": {"title": str | None, "meta_description": str | None, "canonical": str | None},
"links": [{"url": str, "anchor_text": str | None, "rel": str | None}, ...],
"body_bytes": int, "fetched_via": "Http",
},
...
],
"errors": [{"url": str, "message": str}, ...],
"stats": {
"urls_discovered": int, "urls_fetched": int, "urls_skipped": int,
"errors": int, "bytes_downloaded": int, "duration_ms": int,
},
}
Options
All keyword-only, matching the CLI's flags:
| Parameter | Default | Meaning |
|---|---|---|
depth |
3 |
Max link-following depth |
concurrency |
50 |
Max requests in flight, whole crawl |
per_host_concurrency |
4 |
Max requests in flight per host |
same_domain |
True |
Restrict crawl to the seed's resolved domain |
max_urls |
10000 |
Hard cap on URLs fetched |
max_duration_secs |
None |
Optional wall-clock budget |
timeout_secs |
30 |
Per-request timeout |
request_delay_ms |
0 |
Min delay between requests to the same host |
max_response_bytes |
20 MB |
Per-response size cap |
respect_robots |
True |
Honor robots.txt |
allow_private_networks |
False |
Permit loopback/private targets (internal use only) |
Limitations
- Blocking, not async.
crawl()spins up its own Tokio runtime internally and blocks until done. The GIL is released while it runs (other Python threads keep going), but there is noasynciointegration in this version. - Single-platform wheels initially. Built and published from the maintainer's machine, not a cross-platform CI matrix yet — check PyPI for which platforms have wheels; others will need a Rust toolchain to build from source.
License
Apache-2.0.
Release files for mudfish 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| mudfish-0.1.1.tar.gz | 54.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| mudfish-0.1.1-cp39-abi3-macosx_11_0_arm64.whl | CPython 3.9 | abi3 | macOS 11.0+ ARM64 | Details |
Total release size: 3.4 MB
Release files / mudfish-0.1.1.tar.gz
| Download URL | mudfish-0.1.1.tar.gz |
|---|---|
| Size | 54.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4f0cf8d2ee47b6dd77e1c1a46aa5f0b270c2eeb0749ab5b2f85d0344071fbe4c
|
|
BLAKE2b-256 checksum How to use checksums |
564c95cf8cfce5733306d1da9fa94ef1a1dc06441dabeb6fb80c58d8643b7bbc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.6
|
Release files / mudfish-0.1.1-cp39-abi3-macosx_11_0_arm64.whl
| Download URL | mudfish-0.1.1-cp39-abi3-macosx_11_0_arm64.whl |
|---|---|
| Size | 3.4 MB |
| Tags | CPython 3.9 abi3 macOS 11.0+ ARM64 |
|
SHA-256 checksum How to use checksums |
fc6197a7d19b18943ba11c62dab13df322a0b31a344f7dff0516df07891aa4f2
|
|
BLAKE2b-256 checksum How to use checksums |
b922e5a4f49c68b79b8378a736da0d908c01032ad4472e9ebd996b4ad6840130
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.6
|