Skip to main content

Typed Python SDK for Supacrawler API (scrape, jobs, screenshots, watch)

Project description

supacrawler-py

Typed Python SDK for Supacrawler API.

Install

  • From PyPI (when published):
pip install supacrawler-py
  • With uv (PEP 621 aware):
uv add supacrawler-py
# or for local dev
uv pip install -e .
  • From local path (development):
pip install -e .
  • From GitHub (direct):
pip install git+https://github.com/Supacrawler/supacrawler-py.git#subdirectory=sdk/supacrawler-py

Usage

from supacrawler import SupacrawlerClient, ScrapeParams, JobCreateRequest, ScreenshotRequest, WatchCreateRequest

client = SupacrawlerClient(api_key="YOUR_API_KEY")

# Scrape
scrape = client.scrape(ScrapeParams(url="https://example.com", format="markdown"))
print(scrape)

# Create crawl job
job = client.create_job(JobCreateRequest(url="https://supabase.com/docs", type="crawl", depth=2, link_limit=10, format="markdown"))
status = client.wait_for_job(job.job_id)
print(status.status)

# Screenshot job
sjob = client.create_screenshot_job(ScreenshotRequest(url="https://example.com", device="desktop", full_page=True))
print(sjob.job_id)

# Watch
watch = client.watch_create(WatchCreateRequest(url="https://example.com/pricing", frequency="daily", notify_email="me@example.com"))
print(watch.watch_id)

Advanced

  • See examples/*.ipynb for full parameter coverage:
    • Scrape: format, render_js, wait, device, depth, max_links, fresh
    • Jobs: format, link_limit, depth, include_subdomains, render_js, patterns
    • Screenshots: device/full_page/format/quality/viewport/device_scale, waits, selectors, blocking, modes
    • Watch: frequency, selector, include_html/image, full_page, quality, pause/resume/check/delete

API coverage

  • GET /v1/scrape (all params)
  • POST /v1/jobs, GET /v1/jobs/{id}
  • POST /v1/screenshots
  • POST /v1/watch, GET /v1/watch, GET /v1/watch/{id}, DELETE /v1/watch/{id}, PATCH /v1/watch/{id}/pause, PATCH /v1/watch/{id}/resume, POST /v1/watch/{id}/check

Development

  • Env var: SUPACRAWLER_API_KEY for examples.
  • Not needed: requirements.txt for end-users (deps declared in pyproject.toml). Optional requirements-dev.txt if you add tests or notebooks tooling.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

supacrawler_py-0.1.1.tar.gz (5.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

supacrawler_py-0.1.1-py3-none-any.whl (5.5 kB view details)

Uploaded Python 3

File details

Details for the file supacrawler_py-0.1.1.tar.gz.

File metadata

  • Download URL: supacrawler_py-0.1.1.tar.gz
  • Upload date:
  • Size: 5.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.11.13

File hashes

Hashes for supacrawler_py-0.1.1.tar.gz
Algorithm Hash digest
SHA256 f8b1879923b587b83ae0fa16ea175af434896fbf077d9dfe031262bc04bd493f
MD5 c97784016d001b21b5eaab9ced20172a
BLAKE2b-256 25aae3e06098652266afb9c2de5e54387f9834d9ca3f683b041be5dcdf18599f

See more details on using hashes here.

File details

Details for the file supacrawler_py-0.1.1-py3-none-any.whl.

File metadata

  • Download URL: supacrawler_py-0.1.1-py3-none-any.whl
  • Upload date:
  • Size: 5.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.11.13

File hashes

Hashes for supacrawler_py-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 7cf1da6ba349cc929a43d9b1bd053d996e69e4da08877aaff524cff6b9620fec
MD5 dd4ba50a055e0750c598c8d9cd9df147
BLAKE2b-256 a7e5dfb0d029b6c46e8552179a5088a7cf6eb2b49e25d12b93f77b2ec238be7f

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page