Skip to main content

Official Python SDK for the Berrycrawl API

Project description

Berrycrawl Python SDK

pypi

The official Python SDK for scraping, crawling, searching, mapping, structured extraction, screenshots, and brand profiles.

Documentation · Dashboard · GitHub

Table of Contents

Installation

pip install berrycrawl

Reference

A full reference for this library is available here.

Usage

Set BERRYCRAWL_API_KEY to an API key from the Berrycrawl dashboard.

import os
from berrycrawl import Berrycrawl

client = Berrycrawl(api_key=os.environ["BERRYCRAWL_API_KEY"])
page = client.scrape(url="https://example.com/pricing")
print(page.data["markdown"])

Crawl a website

job = client.crawl(url="https://example.com/docs", limit=50)
print(job.id)

Search and map

results = client.search(query="best headless browser libraries", limit=10)
site_map = client.map_(url="https://example.com", search="documentation")

Retrieve a brand profile

brand = client.brand.retrieve(url="https://stripe.com")
print(brand.data)

Environments

This SDK allows you to configure different environments for API requests.

from berrycrawl import Berrycrawl
from berrycrawl.environment import BerrycrawlEnvironment

client = Berrycrawl(
    environment=BerrycrawlEnvironment.PRODUCTION,
)

Async Client

The SDK also exports an async client so that you can make non-blocking calls to our API. Note that if you are constructing an Async httpx client class to pass into this client, use httpx.AsyncClient() instead of httpx.Client() (e.g. for the httpx_client parameter of this client).

import asyncio

from berrycrawl import AsyncBerrycrawl

client = AsyncBerrycrawl(
    api_key="<token>",
)


async def main() -> None:
    await client.brand.retrieve(
        url="https://stripe.com",
    )


asyncio.run(main())

Exception Handling

When the API returns a non-success status code (4xx or 5xx response), a subclass of the following error will be thrown.

from berrycrawl.core.api_error import ApiError

try:
    client.brand.retrieve(...)
except ApiError as e:
    print(e.status_code)
    print(e.body)

Advanced

Access Raw Response Data

The SDK provides access to raw response data, including headers, through the .with_raw_response property. The .with_raw_response property returns a "raw" client that can be used to access the .headers and .data attributes.

from berrycrawl import Berrycrawl

client = Berrycrawl(...)
response = client.brand.with_raw_response.retrieve(...)
print(response.headers)  # access the response headers
print(response.status_code)  # access the response status code
print(response.data)  # access the underlying object

Retries

The SDK is instrumented with automatic retries with exponential backoff. A request will be retried as long as the request is deemed retryable and the number of retry attempts has not grown larger than the configured retry limit (default: 2).

Which status codes are retried depends on the retryStatusCodes generator configuration:

legacy (current default): retries on

  • 408 (Timeout)
  • 409 (Conflict)
  • 429 (Too Many Requests)
  • 5XX (All server errors, including 500)

recommended: retries on

  • 408 (Timeout)
  • 409 (Conflict)
  • 429 (Too Many Requests)
  • 502 (Bad Gateway)
  • 503 (Service Unavailable)
  • 504 (Gateway Timeout)

Use the max_retries request option to configure this behavior.

client.brand.retrieve(..., request_options={
    "max_retries": 1
})

Timeouts

The SDK defaults to a 60 second timeout. You can configure this with a timeout option at the client or request level.

from berrycrawl import Berrycrawl

client = Berrycrawl(..., timeout=20.0)

# Override timeout for a specific method
client.brand.retrieve(..., request_options={
    "timeout": 1
})

Custom Client

You can override the httpx client to customize it for your use-case. Some common use-cases include support for proxies and transports.

import httpx
from berrycrawl import Berrycrawl

client = Berrycrawl(
    ...,
    httpx_client=httpx.Client(
        proxy="http://my.test.proxy.example.com",
        transport=httpx.HTTPTransport(local_address="0.0.0.0"),
    ),
)

Contributing

While we value open-source contributions to this SDK, this library is generated programmatically. Additions made directly to this library would have to be moved over to our generation code, otherwise they would be overwritten upon the next generated release. Feel free to open a PR as a proof of concept, but know that we will not be able to merge it as-is. We suggest opening an issue first to discuss with us!

On the other hand, contributions to the README are always very welcome!

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

berrycrawl-0.1.2.tar.gz (69.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

berrycrawl-0.1.2-py3-none-any.whl (122.9 kB view details)

Uploaded Python 3

File details

Details for the file berrycrawl-0.1.2.tar.gz.

File metadata

  • Download URL: berrycrawl-0.1.2.tar.gz
  • Upload date:
  • Size: 69.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for berrycrawl-0.1.2.tar.gz
Algorithm Hash digest
SHA256 fb079d80429d645ed0cf87aa7cb0c1c4b9ff301b852610ff627c8dc514f1eb85
MD5 4c0f4890d9445edd0f3c75aaa7997804
BLAKE2b-256 68a0f9c6f3c20244fa31c5296994104e0660b22dece41b99b2981c18f4ff1927

See more details on using hashes here.

Provenance

The following attestation bundles were made for berrycrawl-0.1.2.tar.gz:

Publisher: ci.yml on strawberry-labs/berrycrawl-python

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file berrycrawl-0.1.2-py3-none-any.whl.

File metadata

  • Download URL: berrycrawl-0.1.2-py3-none-any.whl
  • Upload date:
  • Size: 122.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for berrycrawl-0.1.2-py3-none-any.whl
Algorithm Hash digest
SHA256 d46529d5ecd5a6a3289e4b4307d0e99302989cbc41014eef848f1782c25ef8cb
MD5 29dda9818e3d3b48db85c037f56f168a
BLAKE2b-256 88558657f8f86a464a1752ae93f4b458425893d97b9663f4d50642b9d9377310

See more details on using hashes here.

Provenance

The following attestation bundles were made for berrycrawl-0.1.2-py3-none-any.whl:

Publisher: ci.yml on strawberry-labs/berrycrawl-python

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page