Typed Python SDK for Supacrawler API (scrape, jobs, screenshots, watch)
Project description
supacrawler-py
Typed Python SDK for Supacrawler API.
Install
- From PyPI (when published):
pip install supacrawler-py
- With uv (PEP 621 aware):
uv add supacrawler-py
# or for local dev
uv pip install -e .
- From local path (development):
pip install -e .
- From GitHub (direct):
pip install git+https://github.com/Supacrawler/supacrawler-py.git#subdirectory=sdk/supacrawler-py
Usage
from supacrawler import SupacrawlerClient, ScrapeParams, JobCreateRequest, ScreenshotRequest, WatchCreateRequest
client = SupacrawlerClient(api_key="YOUR_API_KEY")
# Scrape
scrape = client.scrape(ScrapeParams(url="https://example.com", format="markdown"))
print(scrape)
# Create crawl job
job = client.create_job(JobCreateRequest(url="https://supabase.com/docs", type="crawl", depth=2, link_limit=10, format="markdown"))
status = client.wait_for_job(job.job_id)
print(status.status)
# Screenshot job
sjob = client.create_screenshot_job(ScreenshotRequest(url="https://example.com", device="desktop", full_page=True))
print("job:", sjob.job_id)
# Wait and fetch a fresh signed URL (recommended)
signed = client.wait_for_screenshot(sjob.job_id)
print("screenshot:", signed.screenshot)
# Watch
watch = client.watch_create(WatchCreateRequest(url="https://example.com/pricing", frequency="daily", notify_email="me@example.com"))
print(watch.watch_id)
Advanced
- See
examples/*.ipynbfor full parameter coverage:- Scrape: format, wait, device, depth, max_links, fresh
- Jobs: format, link_limit, depth, include_subdomains, patterns
- Screenshots: device/full_page/format/quality/viewport/device_scale, waits, selectors, blocking, modes
- Watch: frequency, selector, include_html/image, full_page, quality, pause/resume/check/delete
API coverage
- GET
/v1/scrape(all params) - POST
/v1/crawl, GET/v1/crawl/{id} - POST
/v1/screenshots - POST
/v1/watch, GET/v1/watch, GET/v1/watch/{id}, DELETE/v1/watch/{id}, PATCH/v1/watch/{id}/pause, PATCH/v1/watch/{id}/resume, POST/v1/watch/{id}/check
Development
- Env var:
SUPACRAWLER_API_KEYfor examples. - Not needed:
requirements.txtfor end-users (deps declared inpyproject.toml). Optionalrequirements-dev.txtif you add tests or notebooks tooling.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
supacrawler_py-0.1.8.tar.gz
(40.9 kB
view details)
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file supacrawler_py-0.1.8.tar.gz.
File metadata
- Download URL: supacrawler_py-0.1.8.tar.gz
- Upload date:
- Size: 40.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5ea6cc6643e263ed7ff1bc5c3e80c3ce42d7bf45bc7cfa504683a10e37bbc1d0
|
|
| MD5 |
ebbd6a7a290e6305412aaf66e71e9a38
|
|
| BLAKE2b-256 |
25a6a38948aa877849d21273de7e362beb07d8f36f469aaf57eb99215ef3c035
|
File details
Details for the file supacrawler_py-0.1.8-py3-none-any.whl.
File metadata
- Download URL: supacrawler_py-0.1.8-py3-none-any.whl
- Upload date:
- Size: 69.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ab29fae077c9ecad0009816a489e857ce0bd49653813d33dc3ddfca272b8ee29
|
|
| MD5 |
f187455f50ef98682d079b50ec89069e
|
|
| BLAKE2b-256 |
f3708d46da49c9975c185a53e3171f8139ccc30e8d260838953a4178b508d146
|