Skip to main content

SSscraper

Async web scraping library for SiteSudharo.
Built on Crawl4ai — downloads page resources to local disk, AWS S3, or Cloudflare R2 and returns structured metadata for every file.

Features

  • Full-page scraping via Crawl4ai (headless browser)
  • Downloads: images, CSS, fonts, JS, documents, video, audio
  • Storage backends: Local · S3 · R2 (drop-in swappable)
  • Per-resource metadata: URL, stored path, content-type, size, MD5 + SHA-256
  • Async & concurrent — configurable parallelism
  • Retry with backoff, per-file size cap

Quick start

import asyncio
from ssscraper import SScraper, LocalStorage, ScraperConfig

async def main():
    scraper = SScraper(
        storage=LocalStorage("./downloads"),
        config=ScraperConfig(download_images=True, download_css=True, download_fonts=True),
    )
    result = await scraper.scrape("https://example.com")
    print(f"Downloaded {len(result.succeeded)} resources")
    for r in result.resources:
        print(r.model_dump_json())

asyncio.run(main())

Storage backends

Local

from ssscraper import LocalStorage
storage = LocalStorage(base_dir="./downloads")

AWS S3

from ssscraper import S3Storage
storage = S3Storage(
    bucket="my-bucket",
    prefix="sitesudharo",
    region="us-east-1",
    aws_access_key_id="...",
    aws_secret_access_key="...",
)

Cloudflare R2

from ssscraper import R2Storage
storage = R2Storage(
    bucket="my-bucket",
    account_id="<cf-account-id>",
    access_key_id="...",
    secret_access_key="...",
    prefix="sitesudharo",
    public_domain="assets.yourdomain.com",  # optional
)

ScrapeResult shape

result.page            # PageMetadata — url, title, description, scraped_at
result.html            # raw HTML
result.markdown        # Crawl4ai markdown
result.resources       # List[ResourceMetadata]
result.succeeded       # filter: status == SUCCESS
result.failed          # filter: status == FAILED
result.images          # filter: type == image
result.stylesheets     # filter: type == css
result.fonts           # filter: type == font

ResourceMetadata fields

Field Type Description
original_url str Source URL
storage_key str Relative key inside the storage backend
stored_path str Absolute local path or full cloud URL
resource_type ResourceType image / css / javascript / font / …
content_type str HTTP Content-Type
size_bytes int File size
checksum_md5 str MD5 hex digest
checksum_sha256 str SHA-256 hex digest
status ResourceStatus success / failed / skipped
error str Error message if failed
downloaded_at datetime UTC timestamp

Install

pip install -e .          # from source
pip install ssscraper     # once published

Python ≥ 3.11 required.

Config options

See ssscraper/config.py — all fields are optional with sensible defaults.

Metadata

Release files for ssscraper 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ssscraper 0.1.0
File Size Uploaded
ssscraper-0.1.0.tar.gz 21.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ssscraper 0.1.0
File Interpreter ABI Platform
ssscraper-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 40.7 kB

Release files / ssscraper-0.1.0.tar.gz

Download URL ssscraper-0.1.0.tar.gz
Size 21.4 kB
Tags Source
SHA-256 checksum
How to use checksums
7e5b519dd382ac5b78f61e5b024afe974409bd120f99d0abb99cbfb9c3bc59d4
BLAKE2b-256 checksum
How to use checksums
c6a95de2cbc9eebfd441a15a4f642d5dcb9d972b5941dcaeb3f27bba3c8f1158
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.7.5

Release files / ssscraper-0.1.0-py3-none-any.whl

Download URL ssscraper-0.1.0-py3-none-any.whl
Size 19.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f497ddc0761252135a51277f39235e0ed9e0d0dc13b3aa65c18a2636f1287a12
BLAKE2b-256 checksum
How to use checksums
bbb986902b1136b4ec2b5e0c525c3e6a35d8b34283f23ed4e712b2821d160928
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.7.5

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page