silkworm-rs
Silkworm is an async-first Python web scraping framework built on wreq and scraper-rs. It combines a small, typed Spider/Request/Response API with middleware, output pipelines, and production crawl controls.
Documentation · Getting started · Examples · API reference
Highlights
- Async crawling with concurrency, priorities, request deduplication, timeouts, and deadlock-free queue backpressure.
- Browser-impersonating HTTP through wreq, with redirects, proxies, cookies, retries, throttling, and robots.txt support.
- Push-style typed callbacks using
await self.emit(...)andawait response.follow(...). - Async CSS/XPath selection and optional declarative
Item,Text, andAttrextraction. - File, database, cloud, queue, and message-stream pipelines, with batch processing available on every built-in pipeline.
- Production controls for failure policies, stop limits, pause/resume, caching, graceful shutdown, metrics, and structured crawl statistics.
- Optional CDP, Servo, and OnionLink clients for rendered or onion-service pages.
Install
Silkworm supports Python 3.13–3.15.
pip install silkworm-rs
With uv:
uv add silkworm-rs
Integrations are installed as optional extras. For example:
pip install "silkworm-rs[uvloop,polars]"
See Getting Started for the complete extras and Python-version compatibility table.
Quick start
from silkworm import HTMLResponse, Response, Spider, run_spider
from silkworm.pipelines import JsonLinesPipeline
class QuotesSpider(Spider):
name = "quotes"
start_urls = ("https://quotes.toscrape.com/",)
async def parse(self, response: Response) -> None:
if not isinstance(response, HTMLResponse):
return
for quote in await response.select(".quote"):
text = await quote.select_first(".text")
author = await quote.select_first(".author")
if text is not None and author is not None:
await self.emit({"text": text.text, "author": author.text})
next_link = await response.select_first("li.next > a")
if next_link is not None and (href := next_link.attr("href")):
await response.follow(href, callback=self.parse)
run_spider(
QuotesSpider,
item_pipelines=[JsonLinesPipeline("data/quotes.jl")],
)
Callbacks return None; they report items and requests with emit and follow,
which apply backpressure while the callback is running.
Command line
Run a spider module or test its parser without writing a runner script:
silkworm crawl examples/quotes_spider.py -o data/quotes.jl -s max_items=100
silkworm parse https://quotes.toscrape.com/ --spider examples/quotes_spider.py
See the CLI reference for configuration, output formats, and exit codes.
Documentation
| Topic | Guide |
|---|---|
| Installation, extras, and first spider | Getting Started |
| Spider, request, response, selectors, and callbacks | Core Concepts |
| Declarative item extraction | Declarative Extraction |
| Engine, HTTP, CDP, Servo, and OnionLink | Engine and HTTP Client |
| Request and response middleware | Middlewares |
| Output destinations and batch processing | Pipelines |
| asyncio, rsloop, uvloop, winloop, and Trio | Runners |
| Resumable and observable crawls | Production Crawling |
| Version changes | Migration Guide |
| Constraints and workarounds | Limitations |
Development
uv venv --python python3.13
uv sync --group dev
just fmt && just lint && just typecheck && just test
See the development workflow and open an issue or pull request on GitHub.
Acknowledgements
Silkworm builds on wreq, scraper-rs, fast-h2m, rxml, and optional integrations including OnionLink and Servo. Thank you to their maintainers and contributors.
License
MIT. See LICENSE.
Metadata
Release files for silkworm-rs 0.14.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| silkworm_rs-0.14.0.tar.gz | 480.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| silkworm_rs-0.14.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 671.8 kB
Release files / silkworm_rs-0.14.0.tar.gz
| Download URL | silkworm_rs-0.14.0.tar.gz |
|---|---|
| Size | 480.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1d7c40ff9294197449fe356ff9468e2fd80363f4d5195bacb2720cc29566543e
|
|
BLAKE2b-256 checksum How to use checksums |
9d2614df41fb32a01b57498284b11e24ab9237b857cb1d021047935bff6ad1ad
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency logRelease files / silkworm_rs-0.14.0-py3-none-any.whl
| Download URL | silkworm_rs-0.14.0-py3-none-any.whl |
|---|---|
| Size | 190.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
3fbb1d083161c78e14a0f67765ea0a8f1fce6ccf1272aecebd5f837207608f94
|
|
BLAKE2b-256 checksum How to use checksums |
cd03f9c6eef59c8601582d0cb7ca4ccf3d3464f6396f3002bf3f268553124c3a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency log