Skip to main content

Scrapely API client for Python

The official Python client for the Scrapely REST API.

PyPI version PyPI downloads Python versions License

scrapely-client-python lets you talk to the Scrapely platform from Python — run Actors, manage storages (datasets, key-value stores, request queues), schedule tasks, configure webhooks, and use everything else exposed by the Scrapely API. It ships both synchronous and asynchronous clients, fully typed responses, automatic retries with exponential backoff, tiered timeouts, pagination helpers, streaming, and a pluggable HTTP layer.

If you want to build Actors in Python rather than consume the API, use the Scrapely SDK for Python instead — it bundles this client and adds Actor-side primitives.

Table of contents

Installation

scrapely-client-python requires Python 3.11 or higher and is published on PyPI.

  • From PyPI, it can be installed for example with pip:

    pip install scrapely-client-python
    

    or with uv:

    uv add scrapely-client-python
    

    or any other Python package manager that consumes PyPI.

    The client compresses request bodies with gzip by default (no extra dependencies required). To opt in to brotli (better compression ratio), install the optional extra and pass compression='brotli':

    pip install "scrapely-client-python[brotli]"
    # or
    uv add "scrapely-client-python[brotli]"
    

Quick start

You'll need a Scrapely API token — find yours in the integrations section of the Scrapely Console. Pass it to the client and you're ready to go.

Synchronous client

from scrapely_client import ScrapelyClient

client = ScrapelyClient('MY-SCRAPELY-TOKEN')

# Start an Actor and wait for it to finish.
run = client.actor('my-user/hello-world').call(
    run_input={'message': 'Hello, Scrapely!'},
)
if run is None:
    raise RuntimeError('Actor run was not found.')

# Iterate items from the run's default dataset.
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

Asynchronous client

import asyncio

from scrapely_client import ScrapelyClientAsync


async def main() -> None:
    client = ScrapelyClientAsync('MY-SCRAPELY-TOKEN')

    run = await client.actor('my-user/hello-world').call(
        run_input={'message': 'Hello, Scrapely!'},
    )
    if run is None:
        raise RuntimeError('Actor run was not found.')

    # Iterate items from the run's default dataset.
    async for item in client.dataset(run.default_dataset_id).iterate_items():
        print(item)


asyncio.run(main())

Keep your token secret. It authorizes requests on your behalf and can incur usage costs. Never commit it to source control or expose it to client-side code.

Features

  • Synchronous and asynchronous clients — pick ScrapelyClient or ScrapelyClientAsync to match your codebase; both expose the same API (Asyncio support).
  • Fully typed responses — every method returns a Pydantic model generated from the platform's OpenAPI spec, with IDE autocomplete and runtime validation (Typed models).
  • Automatic retries — exponential backoff for network errors, HTTP 429, and 5xx responses, configurable per client (Retries).
  • Tiered timeouts — short / medium / long tiers picked per endpoint, overridable per call (Timeouts).
  • Pagination and streaming — iterate datasets, key-value store keys, or live logs without manual paging or buffering (Pagination, Streaming).
  • Convenience methodscall(), wait_for_finish(), nested resource access, and other shortcuts that hide platform quirks (Convenience methods).
  • Pluggable HTTP layer — swap the default Impit-based HTTP client for httpx, requests, aiohttp, or any custom implementation (Custom HTTP clients).
  • Structured errors — every API error surfaces as a ScrapelyApiError with HTTP-specific subclasses for precise handling (Error handling).
  • Debug logging — opt-in structured logging on the scrapely_client logger captures request URLs, status codes, retry attempts, and more (Logging).

Usage examples

The client mirrors the platform's resource model. Each entry point returns either a single-resource client for an individual item or a collection client for listing and creating items (Single and collection clients).

List Actors and create one

actors = client.actors()
print(actors.list(limit=10).items)

new_actor = actors.create(name='my-actor')

Stream live logs while a run is in progress

run = client.actor('my-user/web-scraper').start(run_input={...})

with client.run(run.id).log().stream() as log_stream:
    for chunk in log_stream.iter_bytes():
        print(chunk.decode(), end='')

Read and write key-value store records

store = client.key_value_store('STORE-ID')
store.set_record('greeting', {'message': 'Hello!'})
record = store.get_record('greeting')

Iterate dataset items with automatic pagination

for item in client.dataset('DATASET-ID').iterate_items(fields=['title', 'url']):
    process(item)

Tune retries and timeouts

from datetime import timedelta

from scrapely_client import ScrapelyClient

client = ScrapelyClient(
    token='MY-SCRAPELY-TOKEN',
    max_retries=8,
    min_delay_between_retries=timedelta(milliseconds=500),
    timeout_long=timedelta(minutes=10),
)

For end-to-end recipes — passing input, managing tasks for reusable input, retrieving and merging Actor data, integrating with Pandas, plugging in a custom HTTP client — see the Guides.

Documentation

The full documentation lives at https://api.scrape.ly.

Section What you'll find
Introduction Overview, prerequisites, and a tour of the client.
Quick start Authenticate, run an Actor, and fetch its results step by step.
Concepts Asyncio, single vs. collection clients, nested clients, error handling, retries, logging, convenience methods, pagination, streaming, custom HTTP clients, timeouts.
Guides Pass input to an Actor, manage tasks for reusable input, retrieve Actor data, integrate with data libraries (e.g. Pandas), use HTTPX as the HTTP client.
API reference Generated reference for every class, method, and model.
Changelog Release history and breaking changes.

Related projects

  • Scrapely SDK for Python — toolkit for building Actors in Python (this client is bundled with it).
  • Crawlee for Python — high-level web scraping and browser automation framework that powers many Actors.
  • Scrapely API client for JavaScript / TypeScript — equivalent Scrapely API client for Node.js.
  • Scrapely SDK for JavaScript / TypeScript — equivalent Scrapely SDK for Node.js.
  • Crawlee for JavaScript / TypeScript — the Node.js implementation of the Crawlee framework.
  • Scrapely CLI — command-line tool for interacting with the Scrapely platform: managing Actors, runs, storages, local development, and deployment.

Support and community

  • GitHub issues — report a bug or request a feature in the repository's issue tracker.

Contributing

Bug reports, fixes, and improvements are welcome! See CONTRIBUTING.md for the development setup, coding standards, testing, and the release process. The repo uses uv for project management and Poe the Poet as a task runner; the typical loop is:

uv run poe install-dev   # install dev deps and git hooks
uv run poe check-code    # lint, type-check, unit tests, docstring check

License

Released under the Apache License 2.0.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

scrapely_client_python-1.0.0.tar.gz (126.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

scrapely_client_python-1.0.0-py3-none-any.whl (154.5 kB view details)

Uploaded Python 3

File details

Details for the file scrapely_client_python-1.0.0.tar.gz.

File metadata

  • Download URL: scrapely_client_python-1.0.0.tar.gz
  • Upload date:
  • Size: 126.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.12.10 {"installer":{"name":"uv","version":"0.12.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for scrapely_client_python-1.0.0.tar.gz
Algorithm Hash digest
SHA256 68b76c442a92908f5512e008763c6bf2f9a980858d97787bad63143ab554aa5f
MD5 8342bef4189e68b3fa0137f0b4d94541
BLAKE2b-256 2caf1d15dedce32e85f870a19851bca989dc699b1704d0edd60f6604012b5e9c

See more details on using hashes here.

File details

Details for the file scrapely_client_python-1.0.0-py3-none-any.whl.

File metadata

  • Download URL: scrapely_client_python-1.0.0-py3-none-any.whl
  • Upload date:
  • Size: 154.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.12.10 {"installer":{"name":"uv","version":"0.12.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for scrapely_client_python-1.0.0-py3-none-any.whl
Algorithm Hash digest
SHA256 ca362ffbdd7b056febfd3c0f240ddd1043172e8b18b0af3a2869e30f21d8a516
MD5 2a752d64f0fcad89107f545f2543f8e6
BLAKE2b-256 01dbaa7d52a12f69d51d808b5021c378b051ed6e8a9c99b494cf553c753df394

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

1.0.0 This release

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page