silkworm-rs
Async-first web scraping framework built on wreq (HTTP with browser impersonation) and scraper-rs (fast HTML parsing). Silkworm gives you a minimal Spider/Request/Response model, middlewares, and pipelines so you can script quick scrapes or build larger crawlers without boilerplate.
📖 Documentation: https://bitingsnakes.github.io/silkworm/
Features
- Async engine with configurable concurrency, priority-aware queueing, bounded, deadlock-free backpressure (defaults to
concurrency * 10; hard for seeding, soft when several callbacks produce requests at once), and per-request timeouts. - wreq-powered HTTP client: browser impersonation, redirect following with loop detection, query merging, and proxy support via
request.meta["proxy"]. - Optional OnionLink client integration for scraping Tor v3
.onionsites without routing through wreq. - Optional Servo rendering via
ServoFetchClientfor JavaScript-rendered pages without changing the default HTTP client. - Typed async spiders with a push-style callback API:
await self.emit(item)streams items to pipelines andawait self.follow(...)/await response.follow(href)schedule requests, both with backpressure;HTMLResponseships selector helpers. - Optional declarative extraction with compiled
Item,Text, andAttrfield plans while keeping the callback API available. - HTML-to-Markdown conversion via
fast-h2m, including richfull, leanminimal, and streaming modes. - Middlewares: User-Agent rotation/default, proxy rotation, cookie jars with save/load, retries of error statuses and transient network failures with exponential backoff, per-host AutoThrottle, robots.txt enforcement (
DisallowandCrawl-delay), flexible delays,SkipNonHTMLMiddlewareto drop non-HTML callbacks, andCloudflareCrawlMiddlewarefor Browser Rendering crawl jobs. - Pipelines: JSON Lines, SQLite, XML (nested data preserved), and CSV (flattens dicts and lists) out of the box, plus schema validation with
ValidationPipeline. - Production controls: a
CrawlResultwith a failure policy (error rate, minimum items, drop rate) so broken spiders fail loudly; stop limits (max_items,max_requests,max_depth,max_duration,max_errors);allowed_domains; per-domain concurrency; graceful SIGINT/SIGTERM shutdown; response size limits; pause/resume withjob_dir; an on-disk HTTP cache; Prometheus metrics; and layered settings (environment,custom_settings, CLI). - A
silkwormcommand line (crawl,parse,fetch) andsilkworm.testinghelpers for testing callbacks offline. - Structured logging via the standard library (
SILKWORM_LOG_LEVEL=DEBUG), plus periodic/final crawl statistics with per-status, per-domain, and per-error breakdowns.
Installation
From PyPI with pip:
pip install silkworm-rs
From PyPI with uv (recommended for faster installs):
uv pip install silkworm-rs
# or if using uv's project management:
uv add silkworm-rs
From source:
uv venv # install uv from https://docs.astral.sh/uv/getting-started/ if needed
source .venv/bin/activate # Windows: .venv\Scripts\activate
uv pip install -e .
Targets Python 3.13+; dependencies are pinned in pyproject.toml.
Quick start
Define a spider by subclassing Spider and implementing an async parse that reports items with await self.emit(...) and schedules follow-up pages with await response.follow(...). This example writes quotes to data/quotes.jl and enables basic user agent, retry, and non-HTML filtering middlewares.
from silkworm import HTMLResponse, Response, Spider, run_spider
from silkworm.middlewares import (
RetryMiddleware,
SkipNonHTMLMiddleware,
UserAgentMiddleware,
)
from silkworm.pipelines import JsonLinesPipeline
class QuotesSpider(Spider):
name = "quotes"
start_urls = ("https://quotes.toscrape.com/",)
async def parse(self, response: Response) -> None:
if not isinstance(response, HTMLResponse):
return
html = response
for quote in await html.select(".quote"):
text_el = await quote.select_first(".text")
author_el = await quote.select_first(".author")
if text_el is None or author_el is None:
continue
tags = await quote.select(".tag")
await self.emit(
{
"text": text_el.text,
"author": author_el.text,
"tags": [t.text for t in tags],
}
)
if next_link := await html.select_first("li.next > a"):
if href := next_link.attr("href"):
await html.follow(href, callback=self.parse)
if __name__ == "__main__":
run_spider(
QuotesSpider,
request_middlewares=[UserAgentMiddleware()],
response_middlewares=[
SkipNonHTMLMiddleware(),
RetryMiddleware(max_times=3, sleep_http_codes=[429, 503]),
],
item_pipelines=[JsonLinesPipeline("data/quotes.jl")],
concurrency=16,
request_timeout=10,
log_stats_interval=30,
)
Upgrading from 0.10
Silkworm 0.11 replaced yield/return-based callbacks with emit/follow.
Callbacks, errbacks and start_requests() are now async functions returning
None: turn yield item into await self.emit(item), yield request into
await self.follow(request), and yield response.follow(href) into
await response.follow(href). Legacy generator callbacks fail with a
SpiderError that explains the fix. See the
migration table.
Upgrading to 0.12
- Runners and
Engine.run()return aCrawlResultinstead ofNone. - The default deduplication key is the request fingerprint (method, canonical URL
with
params, body) instead of the raw URL, so?a=1&b=2and?b=2&a=1are one page while POSTs with different bodies are not. items_scrapedcounts items that passed every pipeline; pipelines can raiseDropItemto discard items (counted asitems_dropped).- Responses carry the exact downloaded bytes (earlier versions re-encoded bodies
as UTF-8 text, corrupting binary files and non-UTF-8 pages), and bodies over
max_response_size_bytes(default 50 MB) fail withResponseTooLargeError. - Timeouts and connection failures raise
HttpTimeoutError/HttpConnectionError(bothHttpErrorsubclasses), whichRetryMiddlewarenow retries. - Requests record their link depth in
meta["depth"], and the new counters' names (such asretries) are reserved inSpider.stats_payload. - The sync runners stop gracefully on SIGINT/SIGTERM (
handle_signals=Falseopts out). - Requests time out after 60 seconds by default (
request_timeout, also theHttpClient(timeout=...)default); passNoneto restore the previous unlimited behaviour.
Production crawling
Declare what a successful crawl means, stop safely, stay polite, and resume after interruptions:
from silkworm import run_spider
from silkworm.middlewares import AutoThrottleMiddleware, RetryMiddleware, RobotsTxtMiddleware
from silkworm.pipelines import JsonLinesPipeline, ValidationPipeline
throttle = AutoThrottleMiddleware(start_delay=0.5, max_delay=30)
result = run_spider(
QuotesSpider,
request_middlewares=[RobotsTxtMiddleware(), throttle],
response_middlewares=[throttle, RetryMiddleware(max_times=3)],
item_pipelines=[ValidationPipeline(Quote), JsonLinesPipeline("data/quotes.jl")],
request_timeout=30,
concurrency_per_domain=4,
max_depth=5,
max_error_rate=0.05, # raise CrawlFailedError above 5% failed requests
min_items=50, # ...or when fewer than 50 valid items were scraped
job_dir="state/quotes", # Ctrl+C, then run again to resume
metrics_port=9410, # Prometheus metrics at http://127.0.0.1:9410/metrics
)
print(result.close_reason, result.items_scraped, result.error_rate)
Or from the command line, with settings from -s, SILKWORM_* environment
variables, or Spider.custom_settings:
silkworm crawl examples/quotes_spider.py -o data/quotes.jl -s max_items=100 --job-dir state/quotes
silkworm parse https://quotes.toscrape.com/ --spider examples/quotes_spider.py
silkworm crawl exits with status 1 when the failure policy is violated, so cron
jobs and CI notice broken spiders. See the
Production Crawling guide
and the CLI reference.
Declarative extraction
Use silkworm.declarative when the extraction is regular and keep ordinary
callbacks for pagination or site-specific behavior:
from silkworm import HTMLResponse, Response, Spider
from silkworm.declarative import Attr, Item, Text
def parse_price(value: str) -> float:
return float(value.removeprefix("$").strip())
class Product(Item):
__selector__ = ".product"
title: str = Text("h2", strip=True)
price: float = Text(".price", transform=parse_price)
url: str = Attr("a", "href", absolute=True)
image: str | None = Attr("img", "src", absolute=True)
tags: list[str] = Text(".tag")
class ProductsSpider(Spider):
start_urls = ("https://shop.example.com/products/",)
async def parse(self, response: Response) -> None:
if not isinstance(response, HTMLResponse):
return
async for product in Product.extract(response):
# Existing pipelines consume JSON-compatible values.
await self.emit(product.to_dict())
The annotation controls selector cardinality:
| Annotation | Selector operation | Missing result |
|---|---|---|
T |
select_first() |
raises MissingFieldError |
T | None |
select_first() |
None |
list[T] |
select() |
[] |
Fields support these options:
default=...supplies the final scalar value when an element or attribute is missing.transform=...applies a synchronous conversion to each extracted string. Declarative extraction does not implicitly coerce annotations.Text(..., strip=True)strips surrounding whitespace before transformation.Attr(..., absolute=True)resolves the attribute throughresponse.url_join().
Plans are compiled from annotations once per Item class and then cached. A
required-field or transform failure includes the item, field, selector, response
URL, and root index. after_extract() is available for the irregular part of an
otherwise declarative extraction:
class Article(Item):
title: str = Text("h1")
async def after_extract(self, response: HTMLResponse) -> None:
self.title = self.title.strip()
Item.extract() intentionally does not replace Spider, callbacks, requests,
middlewares, or pipelines. See examples/declarative_quotes_spider.py for a
complete spider with pagination.
run_spider/crawl knobs:
concurrency: number of concurrent HTTP requests; default 16; must be positive.max_pending_requests: queue bound to avoid unbounded memory use (defaults toconcurrency * 10); if provided, must be positive.start_requests()always respects it, and so does a single producing callback; when several callbacks produce requests at once, they enqueue past it rather than parking workers or deadlocking.request_timeout: per-request timeout (seconds); defaults to 60,Nonedisables it.keep_alive: reuse HTTP connections when supported by the underlying client (sendsConnection: keep-alive).http_client: use a custom client instance such asOnionLinkClient(...)orServoFetchClient(...)instead of the default wreq-backed client.dedup_key: optionalCallable[[Request], str]used for request deduplication; defaults tolambda req: req.url.html_max_size_bytes: limit HTML parsed intoAsyncDocumentto avoid huge payloads.log_stats_interval: seconds between periodic stats logs; final stats are always emitted.request_middlewares/response_middlewares/item_pipelines: plug-ins run on every request/response/item.- use
run_spider_rsloop(...)instead ofrun_spider(...)to run under rsloop (requirespip install silkworm-rs[rsloop]). - use
run_spider_uvloop(...)instead ofrun_spider(...)to run under uvloop (requirespip install silkworm-rs[uvloop]). - use
run_spider_winloop(...)instead ofrun_spider(...)to run under winloop on Windows (requirespip install silkworm-rs[winloop]).
Built-in middlewares and pipelines
from silkworm.middlewares import (
CloudflareCrawlMiddleware,
CookiesMiddleware,
DelayMiddleware,
ProxyMiddleware,
RequestResponseStreamMiddleware,
RetryMiddleware,
RobotsTxtDelayMiddleware,
SkipNonHTMLMiddleware,
UserAgentMiddleware,
)
from silkworm.pipelines import (
CallbackPipeline, # invoke a custom callback function on each item
CSVPipeline,
JsonLinesPipeline,
MsgPackPipeline, # requires: pip install silkworm-rs[msgpack]
RssPipeline,
SQLitePipeline,
XMLPipeline,
TaskiqPipeline, # requires: pip install silkworm-rs[taskiq]
ZenohPipeline, # requires: pip install silkworm-rs[zenoh]
PolarsPipeline, # requires: pip install silkworm-rs[polars]
ExcelPipeline, # requires: pip install silkworm-rs[excel]
YAMLPipeline, # requires: pip install silkworm-rs[yaml]
AvroPipeline, # requires: pip install silkworm-rs[avro]
ElasticsearchPipeline, # requires: pip install silkworm-rs[elasticsearch]
MongoDBPipeline, # requires: pip install silkworm-rs[mongodb]
MySQLPipeline, # requires: pip install silkworm-rs[mysql]
PostgreSQLPipeline, # requires: pip install silkworm-rs[postgresql]
S3JsonLinesPipeline, # requires: pip install silkworm-rs[s3]
VortexPipeline, # requires: pip install silkworm-rs[vortex]
WebhookPipeline, # sends items to webhook endpoints using wreq
GoogleSheetsPipeline, # requires: pip install silkworm-rs[gsheets]
SnowflakePipeline, # requires: pip install silkworm-rs[snowflake]
FTPPipeline, # requires: pip install silkworm-rs[ftp]
SFTPPipeline, # requires: pip install silkworm-rs[sftp]
CassandraPipeline, # requires: pip install silkworm-rs[cassandra]
CouchDBPipeline, # requires: pip install silkworm-rs[couchdb]
DynamoDBPipeline, # requires: pip install silkworm-rs[dynamodb]
DuckDBPipeline, # requires: pip install silkworm-rs[duckdb]
)
run_spider(
QuotesSpider,
request_middlewares=[
UserAgentMiddleware(), # rotate/custom user agent
DelayMiddleware(min_delay=0.3, max_delay=1.2), # polite throttling
# Read Crawl-delay/Request-rate from robots.txt and serialize same-origin requests
# RobotsTxtDelayMiddleware("https://quotes.toscrape.com", user_agent="silkworm"),
# ProxyMiddleware with round-robin selection (default)
# ProxyMiddleware(proxies=["http://user:pass@proxy1:8080", "http://proxy2:8080"]),
# ProxyMiddleware with random selection
# ProxyMiddleware(proxies=["http://proxy1:8080", "http://proxy2:8080"], random_selection=True),
# ProxyMiddleware from file with random selection
# ProxyMiddleware(proxy_file="proxies.txt", random_selection=True),
],
response_middlewares=[
RetryMiddleware(max_times=3, sleep_http_codes=[403, 429]), # backoff + retry
SkipNonHTMLMiddleware(), # drop callbacks for images/APIs/etc
],
item_pipelines=[
JsonLinesPipeline("data/quotes.jl"),
SQLitePipeline("data/quotes.db", table="quotes"),
XMLPipeline("data/quotes.xml", root_element="quotes", item_element="quote"),
CSVPipeline("data/quotes.csv", fieldnames=["author", "text", "tags"]),
MsgPackPipeline("data/quotes.msgpack"),
],
)
DelayMiddlewarestrategies:delay=1.0(fixed),min_delay/max_delay(random), ordelay_func(custom).RobotsTxtDelayMiddleware("https://example.com", user_agent="silkworm")downloadshttps://example.com/robots.txt, appliesCrawl-delayorRequest-rate, and serializes matching-origin requests so concurrency cannot bypass the configured spacing. Usefallback_delay=...to keep a conservative delay when robots.txt cannot be fetched.ProxyMiddlewaresupports three modes:- Round-robin (default):
ProxyMiddleware(proxies=["http://proxy1:8080", "http://proxy2:8080"])cycles through proxies in order. - Random selection:
ProxyMiddleware(proxies=["http://proxy1:8080", "http://proxy2:8080"], random_selection=True)randomly selects a proxy for each request. - From file:
ProxyMiddleware(proxy_file="proxies.txt")loads proxies from a file (one proxy per line, blank lines ignored). Combine withrandom_selection=Truefor random selection from the file.
- Round-robin (default):
CookiesMiddlewarestoresSet-Cookieresponse headers, applies matchingCookierequest headers, supports named jars viarequest.meta["cookiejar"], per-request cookies viarequest.meta["cookies"], opt-out viarequest.meta["dont_merge_cookies"], and Netscape/Mozilla cookie filesave(...)/load(...). Use the same instance inrequest_middlewaresandresponse_middlewares.RetryMiddlewarebacks off withasyncio.sleep; any status insleep_http_codesis retried even if not inretry_http_codes.SkipNonHTMLMiddlewarechecksContent-Typeand optionally sniffs the body (sniff_bytes) to avoid running HTML callbacks on binary/API responses.CloudflareCrawlMiddlewareis opt-in per request viarequest.meta["cloudflare_crawl"]; it submits a Cloudflare Browser Rendering crawl job, polls until completion, and hands your callback a synthetic JSONResponsewith the final API payload.RequestResponseStreamMiddlewarestreams paired request/response telemetry events to a collector endpoint; use the same instance in both middleware lists.JsonLinesPipelinewrites items to a local JSON Lines file and, whenopendalis installed, appends asynchronously via the filesystem backend (use_opendal=Falseto stick to a regular file handle).CSVPipelineflattens nested dicts (e.g.,{"user": {"name": "Alice"}}->user_name) and joins lists with commas;XMLPipelinepreserves nesting.RssPipelinewrites buffered RSS 2.0 feeds from items with configurable title/link/description fields.MsgPackPipelinewrites items in binary MessagePack format using ormsgpack for fast and compact serialization (requirespip install silkworm-rs[msgpack]).TaskiqPipelinesends items to a Taskiq queue for distributed processing (requirespip install silkworm-rs[taskiq]).ZenohPipelinepublishes JSON items to a static or dynamically resolved Zenoh key expression (requirespip install silkworm-rs[zenoh]).PolarsPipelinewrites items to a Parquet file using Polars for efficient columnar storage (requirespip install silkworm-rs[polars]).ExcelPipelinewrites items to an Excel .xlsx file (requirespip install silkworm-rs[excel]).YAMLPipelinewrites items to a YAML file (requirespip install silkworm-rs[yaml]).AvroPipelinewrites items to an Avro file with optional schema (requirespip install silkworm-rs[avro]).ElasticsearchPipelinesends items to an Elasticsearch index (requirespip install silkworm-rs[elasticsearch]).MongoDBPipelinesends items to a MongoDB collection (requirespip install silkworm-rs[mongodb]).MySQLPipelinesends items to a MySQL database table as JSON (requirespip install silkworm-rs[mysql]).PostgreSQLPipelinesends items to a PostgreSQL database table as JSONB (requirespip install silkworm-rs[postgresql]).S3JsonLinesPipelinewrites items to AWS S3 in JSON Lines format using async OpenDAL (requirespip install silkworm-rs[s3]).VortexPipelinewrites items to a Vortex file for high-performance columnar storage with 100x faster random access and 10-20x faster scans compared to Parquet (requirespip install silkworm-rs[vortex]).WebhookPipelinesends items to webhook endpoints via HTTP POST/PUT using wreq (same HTTP client as the spider) with support for batching and custom headers.GoogleSheetsPipelineappends items to Google Sheets with automatic flattening of nested data structures (requirespip install silkworm-rs[gsheets]and service account credentials).SnowflakePipelinesends items to Snowflake data warehouse tables as JSON (requirespip install silkworm-rs[snowflake]).FTPPipelinewrites items to an FTP server in JSON Lines format (requirespip install silkworm-rs[ftp]).SFTPPipelinewrites items to an SFTP server in JSON Lines format with support for password or key-based authentication (requirespip install silkworm-rs[sftp]).CassandraPipelinesends items to Apache Cassandra database tables (requirespip install silkworm-rs[cassandra]).CouchDBPipelinesends items to CouchDB databases as documents (requirespip install silkworm-rs[couchdb]).DynamoDBPipelinesends items to AWS DynamoDB tables with automatic table creation (requirespip install silkworm-rs[dynamodb]).DuckDBPipelinesends items to a DuckDB database table as JSON (requirespip install silkworm-rs[duckdb]).CallbackPipelineinvokes a custom callback function (sync or async) on each item, enabling inline processing logic without creating a full pipeline class. See example below.
Using CallbackPipeline for custom processing
Process items with custom callback functions without creating a full pipeline class:
from silkworm.pipelines import CallbackPipeline
# Sync callback
def print_item(item, spider):
print(f"[{spider.name}] {item}")
return item
# Async callback
async def validate_item(item, spider):
# Could do async operations like database checks
if len(item.get("text", "")) < 10:
print(f"Warning: Short text in item")
return item
# Modifying callback
def enrich_item(item, spider):
item["spider_name"] = spider.name
item["processed"] = True
return item
run_spider(
QuotesSpider,
item_pipelines=[
CallbackPipeline(callback=print_item),
CallbackPipeline(callback=validate_item),
CallbackPipeline(callback=enrich_item),
],
)
Callbacks receive (item, spider) and should return the processed item (or None to return the original item unchanged).
Streaming items to a queue with TaskiqPipeline
Stream scraped items to a Taskiq queue for distributed processing:
from taskiq import InMemoryBroker
from silkworm.pipelines import TaskiqPipeline
broker = InMemoryBroker()
@broker.task
async def process_item(item):
# Your item processing logic here
print(f"Processing: {item}")
# Save to database, send to another service, etc.
pipeline = TaskiqPipeline(broker, task=process_item)
run_spider(MySpider, item_pipelines=[pipeline])
This enables distributed processing, retries, rate limiting, and other Taskiq features. See examples/taskiq_quotes_spider.py for a complete example.
Publishing items with ZenohPipeline
Publish every item immediately to a Zenoh key expression:
from silkworm.pipelines import ZenohPipeline
pipeline = ZenohPipeline("scraping/quotes")
run_spider(MySpider, item_pipelines=[pipeline])
Use a sync or async resolver when items need separate keys:
async def item_key(item, spider):
return f"scraping/{spider.name}/{item['category']}"
pipeline = ZenohPipeline(item_key)
The pipeline opens and closes its own session by default. Pass a configured
zenoh.Config instance to customize that session, or session=existing_session to
reuse a caller-owned session.
Injected sessions remain open when the pipeline closes. Publisher options include encoding,
congestion_control, priority, express, reliability, and allowed_destination.
Dynamic publishers are cached by key until pipeline shutdown.
Handling non-HTML responses
Keep crawls cheap when URLs mix HTML and binaries/APIs:
response_middlewares=[SkipNonHTMLMiddleware(sniff_bytes=1024)]
# Tighten HTML parsing size (bytes) to avoid loading huge bodies into scraper-rs
run_spider(MySpider, html_max_size_bytes=1_000_000)
Performance optimization with rsloop
For improved async performance, enable rsloop as a drop-in replacement for asyncio's event loop:
pip install silkworm-rs[rsloop]
# or with uv:
uv pip install silkworm-rs[rsloop]
Then call run_spider_rsloop (same signature as run_spider):
from silkworm import run_spider_rsloop
run_spider_rsloop(
QuotesSpider,
concurrency=32,
)
Performance optimization with uvloop
For improved async performance, enable uvloop (a fast, drop-in replacement for asyncio's event loop):
pip install silkworm-rs[uvloop]
# or with uv:
uv pip install silkworm-rs[uvloop]
Then call run_spider_uvloop (same signature as run_spider):
from silkworm import run_spider_uvloop
run_spider_uvloop(
QuotesSpider,
concurrency=32,
)
uvloop can provide 2-4x performance improvement for I/O-bound workloads.
Performance optimization with winloop (Windows)
For Windows users who want improved async performance, enable winloop (a Windows-compatible alternative to uvloop):
pip install silkworm-rs[winloop]
# or with uv:
uv pip install silkworm-rs[winloop]
Then call run_spider_winloop (same signature as run_spider):
from silkworm import run_spider_winloop
run_spider_winloop(
QuotesSpider,
concurrency=32,
)
winloop provides significant performance improvements on Windows, similar to what uvloop offers on Unix-like systems.
Running spiders with trio
If you prefer trio over asyncio, you can use run_spider_trio instead of run_spider:
pip install silkworm-rs[trio]
# or with uv:
uv pip install silkworm-rs[trio]
Then use run_spider_trio:
from silkworm import run_spider_trio
run_spider_trio(
QuotesSpider,
concurrency=16,
request_timeout=10,
)
This runs your spider using trio as the async backend via trio-asyncio compatibility layer. The Trio runner currently requires Python 3.13 because trio-asyncio 0.16 is not compatible with Python 3.14 or newer.
JavaScript rendering with Servo
For pages that need JavaScript execution but do not require driving an external browser process, install the optional Servo renderer and pass ServoFetchClient as the spider HTTP client.
Install a wheel from this page: https://github.com/RustedBytes/servofetch-py/releases
from silkworm import HTMLResponse, Response, ServoFetchClient, Spider, run_spider
class RenderedSpider(Spider):
name = "rendered"
start_urls = ("https://example.com/",)
async def parse(self, response: Response) -> None:
if isinstance(response, HTMLResponse):
title = await response.select_first("title")
await self.emit({"title": title.text if title else ""})
run_spider(RenderedSpider, http_client=ServoFetchClient(settle_ms=500))
Per-request render options live in Request.meta: servo_javascript, servo_settle_ms, servo_user_agent, servo_screenshot, and servo_full_page. Request.timeout overrides the client timeout for that request.
ServoFetchClient embeds Servo through servofetch; the existing CDP client connects to an external Lightpanda/Chrome-compatible browser over WebSocket. Use the default wreq client when pages do not need client-side rendering.
JavaScript rendering with Lightpanda (CDP)
For pages that require JavaScript execution, you can use Lightpanda (or any CDP-compatible browser) instead of the standard HTTP client. This uses the Chrome DevTools Protocol (CDP) to control a browser.
Installation
pip install silkworm-rs[cdp]
# or with uv:
uv pip install silkworm-rs[cdp]
Starting Lightpanda
lightpanda --remote-debugging-port=9222
Or use Chrome/Chromium:
chromium --remote-debugging-port=9222 --headless
The same ws://127.0.0.1:9222 endpoint works for both. Chrome only accepts its per-session ws://.../devtools/browser/<id> URL, so when a bare host:port endpoint is refused, CDPClient looks that URL up at http://host:port/json/version and connects to it (keeping your host and port). You can also pass the full URL yourself.
Using CDP in your spider
There are two ways to use CDP: the convenience API or custom spider integration.
Convenience API (simple one-off fetches)
import asyncio
from silkworm import fetch_html_cdp
async def main():
# Fetch HTML with JavaScript rendering
text, doc = await fetch_html_cdp(
"https://example.com",
ws_endpoint="ws://127.0.0.1:9222",
timeout=30.0
)
# Extract data from rendered page
title = await doc.select_first("title")
print(title.text if title else "No title")
asyncio.run(main())
Full Spider Integration
from silkworm import HTMLResponse, Request, Response, Spider
from silkworm.cdp import CDPClient
class LightpandaSpider(Spider):
name = "lightpanda"
start_urls = ("https://example.com/",)
def __init__(self, **kwargs):
super().__init__(**kwargs)
self._cdp_client = None
async def start_requests(self) -> None:
# Connect to CDP endpoint
self._cdp_client = CDPClient(
ws_endpoint="ws://127.0.0.1:9222",
timeout=30.0
)
await self._cdp_client.connect()
for url in self.start_urls:
await self.follow(url, callback=self.parse)
async def parse(self, response: Response) -> None:
if not isinstance(response, HTMLResponse):
return
# Extract links from JavaScript-rendered page
for link in await response.select("a"):
href = link.attr("href")
if href:
await self.emit({"url": href})
async def close(self):
if self._cdp_client:
await self._cdp_client.close()
See examples/lightpanda_simple.py and examples/lightpanda_spider.py for complete working examples.
Note: CDP support is experimental. For production use, consider using dedicated browser automation tools or the standard HTTP client when JavaScript rendering is not required.
Onion services with OnionLink
For Tor v3 .onion sites, install the optional OnionLink extra and pass OnionLinkClient as the spider HTTP client:
pip install "silkworm-rs[onionlink]"
from silkworm import HTMLResponse, OnionLinkClient, Response, Spider, run_spider
class OnionSpider(Spider):
name = "onion"
start_urls = ("http://exampleexampleexampleexampleexampleexampleexampleexampleexampleexample.onion/",)
async def parse(self, response: Response) -> None:
if isinstance(response, HTMLResponse):
title = await response.select_first("title")
await self.emit({"title": title.text if title else ""})
run_spider(
OnionSpider,
http_client=OnionLinkClient(concurrency=4, timeout=30),
)
OnionLinkClient supports Silkworm Request headers, params, body/data, JSON payloads, redirects, HTML detection, and request.meta["redirect_times"]. Override OnionLink's response byte cap per request with request.meta["onionlink_response_limit"].
Logging and crawl statistics
- Structured logs via the standard library; set
SILKWORM_LOG_LEVEL=DEBUGfor verbose request/response/middleware output. - Periodic statistics with
log_stats_interval; final stats always include the close reason, elapsed time, queue size, requests/sec, seen URLs, items scraped and dropped, errors, retries, filtered requests, memory MB, and per-status/domain/error breakdowns. The same data is returned as aCrawlResultand can be served as Prometheus metrics (metrics_port).
Limitations
- By default, HTTP fetches are wreq-based without JavaScript execution; pages requiring client-side rendering can use the optional CDP integration (see "JavaScript rendering with Lightpanda" section) or external browser automation tools. Tor v3
.onionsites can use the optional OnionLink integration. - Request deduplication uses the request fingerprint (method, canonical URL with
params, and body); headers andmetaare ignored, so requests that differ only in headers are dropped unless you setdont_filter=Trueor pass a customdedup_key. - Redirects are followed inside the HTTP client, so
allowed_domainsfilters requests before they are sent but a redirect can still land on another host. - HTML parsing auto-detects encoding (BOM, HTTP headers/meta, charset detection fallback) but still enforces a
html_max_size_bytes/doc_max_size_bytescap (default 5 MB) inscraper-rsselectors, so very large pages may need a higher limit or preprocessing. - Several pipelines buffer all items in memory until close (PolarsPipeline, ExcelPipeline, YAMLPipeline, AvroPipeline, VortexPipeline, S3JsonLinesPipeline, FTPPipeline, SFTPPipeline), which can bloat RAM on long crawls; prefer streaming pipelines like JsonLines/CSV/SQLite for high-volume runs.
- Many destination pipelines rely on optional extras; CassandraPipeline is disabled on Windows because
cassandra-driverdepends on libev there.
Examples
python examples/quotes_spider.py→data/quotes.jlpython examples/quotes_spider_trio.py→data/quotes_trio.jl(demonstrates trio backend)python examples/quotes_spider_winloop.py→data/quotes_winloop.jl(demonstrates winloop backend for Windows)python examples/hackernews_spider.py --pages 5→data/hackernews.jlpython examples/lobsters_spider.py --pages 2→data/lobsters.jlpython examples/start_urls_from_file_spider.py --urls-file data/start_urls.txt --output data/start_urls_from_file.jl(reads one URL per line and schedules custom requests withawait self.follow(...)instart_requests)python examples/url_titles_spider.py --urls-file data/url_titles.jl --output data/titles.jl(includesSkipNonHTMLMiddlewareand stricter HTML size limits)python examples/exception_handling_spider.py→data/exception_handling.jl(demonstratesprocess_exceptionand requesterrback)python examples/cookie_reuse_spiders.py→data/cookie_reuse.jlanddata/cookies.txt(captures cookies in one run, saves them, then loads them for a second run)SILKWORM_LOG_LEVEL=DEBUG python examples/logging_controls_demo.py --mode noisythen--mode quiet→ demonstrates noisy pipeline/URL logging and the quieterEngineLogger+ pipelinelog_level=Nonesetuppython examples/export_formats_demo.py --pages 2→ JSONL, XML, and CSV outputs indata/python examples/taskiq_quotes_spider.py --pages 2→ demonstrates TaskiqPipeline for queue-based processingpython examples/sitemap_spider.py --sitemap-url https://example.com/sitemap.xml --pages 50→data/sitemap_meta.jl(extracts meta tags and Open Graph data from sitemap URLs)python examples/request_response_stream_spider.py --collector-url https://collector.example.com/events→ streams request/response telemetry while writing quotes outputCLOUDFLARE_ACCOUNT_ID=... CLOUDFLARE_API_TOKEN=... python examples/cloudflare_crawl_spider.py https://example.com --limit 10→ submits a Cloudflare Browser Rendering crawl jobpython examples/lightpanda_simple.py→ demonstrates CDP/Lightpanda for JavaScript rendering (requirespip install silkworm-rs[cdp]and running Lightpanda)python examples/lightpanda_spider.py→ full spider example using CDP/Lightpandapython examples/servo_spider.py→ full spider example usingServoFetchClientand aservofetchwheel
Convenience API
For one-off fetches without a full spider:
Standard HTTP fetch
import asyncio
from silkworm import fetch_html
async def main():
text, doc = await fetch_html("https://example.com")
title = await doc.select_first("title")
print(title.text if title else "No title")
asyncio.run(main())
CDP-based fetch (with JavaScript rendering)
import asyncio
from silkworm import fetch_html_cdp
async def main():
# Requires Lightpanda/Chrome running with CDP enabled
text, doc = await fetch_html_cdp("https://example.com")
title = await doc.select_first("title")
print(title.text if title else "No title")
asyncio.run(main())
Servo-based fetch
import asyncio
from silkworm import fetch_html_servo
async def main():
# Requires a compatible servofetch wheel.
text, doc = await fetch_html_servo("https://example.com", settle_ms=500)
title = await doc.select_first("title")
print(title.text if title else "No title")
asyncio.run(main())
HTML to Markdown
from silkworm import HTMLResponse, Response, Spider, html_to_markdown, stream_html_to_markdown
markdown = html_to_markdown("<h1>Hello</h1><p>World</p>", mode="minimal")
streamed = stream_html_to_markdown(["<h1>Hello</h1>", "<p>World</p>"])
class MarkdownSpider(Spider):
name = "markdown"
start_urls = ("https://example.com",)
async def parse(self, response: Response) -> None:
if not isinstance(response, HTMLResponse):
return
await self.emit(
{
"url": response.url,
"markdown": await response.to_markdown(mode="full"),
}
)
Modes are full for rich conversion, minimal for the lean Fast DOM path, and mdream for the mdream-backed converter. to_markdown_result(...) and convert_html_to_markdown(...) return fast-h2m's structured result.
Contributing
Pull requests and issues are welcome. To set up a dev environment, install uv, create a Python 3.13 virtualenv, and sync dev dependencies:
uv venv --python python3.13
uv sync --group dev
Run the checks before opening a PR:
just fmt && just lint && just typecheck && just test
Acknowledgements
Silkworm is built on top of excellent open-source projects:
- wreq - HTTP client with browser impersonation capabilities
- onionlink - Tor v3 onion-service client
- servofetch - Bindings to the Servo browser
- scraper-rs - Fast HTML parsing library
- fast-h2m - Fast HTML-to-Markdown conversion
- rxml - XML parsing and writing
We are grateful to the maintainers and contributors of these projects for their work.
License
MIT License. See LICENSE for details.
Metadata
Release files for silkworm-rs 0.12.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| silkworm_rs-0.12.0.tar.gz | 484.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| silkworm_rs-0.12.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 676.9 kB
Release files / silkworm_rs-0.12.0.tar.gz
| Download URL | silkworm_rs-0.12.0.tar.gz |
|---|---|
| Size | 484.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e02b001f4ab7b88628cb2ff4634f19e6cacb286cb352cfe184dcdbbb7310fc63
|
|
BLAKE2b-256 checksum How to use checksums |
cf2807c19d4c833808b68b2f371e74d47dc41eaa727b06d5972feb90349ebf67
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.
Transparency logRelease files / silkworm_rs-0.12.0-py3-none-any.whl
| Download URL | silkworm_rs-0.12.0-py3-none-any.whl |
|---|---|
| Size | 192.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
529cf27bfc5e182c6ea1deee3c6b7be7db265adf60aa5f296faaecf770688110
|
|
BLAKE2b-256 checksum How to use checksums |
b12b6c7c10e65660c0bc3dd3b6249ebf64d19339c8f452ea79abbc63c1986eb8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.
Transparency log