Skip to main content

Official Python SDK for the Berrycrawl API

Project description

Berrycrawl Python SDK

pypi

The official Python SDK for scraping, crawling, searching, mapping, structured extraction, screenshots, and brand profiles.

Documentation · Dashboard · GitHub

Table of Contents

Installation

pip install berrycrawl

Reference

A full reference for this library is available here.

Usage

Set BERRYCRAWL_API_KEY to an API key from the Berrycrawl dashboard.

import os
from berrycrawl import Berrycrawl

client = Berrycrawl(api_key=os.environ["BERRYCRAWL_API_KEY"])
page = client.scrape(url="https://example.com/pricing")
print(page.data["markdown"])

Crawl a website

job = client.crawl(url="https://example.com/docs", limit=50)
print(job.id)

Search and map

results = client.search(query="best headless browser libraries", limit=10)
site_map = client.map_(url="https://example.com", search="documentation")

Retrieve a brand profile

brand = client.brand.retrieve(url="https://stripe.com")
print(brand.data)

Brand design system

Brand responses include an optional branding object for compatibility with older API deployments. When available, it contains the rendered light/dark scheme, semantic colors, typography, spacing, representative input and button styles, and semantic image roles. Use branding.images.favicon for the square icon and branding.images.logo for the wordmark.

Environments

This SDK allows you to configure different environments for API requests.

from berrycrawl import Berrycrawl
from berrycrawl.environment import BerrycrawlEnvironment

client = Berrycrawl(
    environment=BerrycrawlEnvironment.PRODUCTION,
)

Async Client

The SDK also exports an async client so that you can make non-blocking calls to our API. Note that if you are constructing an Async httpx client class to pass into this client, use httpx.AsyncClient() instead of httpx.Client() (e.g. for the httpx_client parameter of this client).

import asyncio

from berrycrawl import AsyncBerrycrawl

client = AsyncBerrycrawl(
    api_key="<token>",
)


async def main() -> None:
    await client.brand.retrieve(
        url="https://stripe.com",
    )


asyncio.run(main())

Exception Handling

When the API returns a non-success status code (4xx or 5xx response), a subclass of the following error will be thrown.

from berrycrawl.core.api_error import ApiError

try:
    client.brand.retrieve(...)
except ApiError as e:
    print(e.status_code)
    print(e.body)

Advanced

Access Raw Response Data

The SDK provides access to raw response data, including headers, through the .with_raw_response property. The .with_raw_response property returns a "raw" client that can be used to access the .headers and .data attributes.

from berrycrawl import Berrycrawl

client = Berrycrawl(...)
response = client.brand.with_raw_response.retrieve(...)
print(response.headers)  # access the response headers
print(response.status_code)  # access the response status code
print(response.data)  # access the underlying object

Retries

The SDK is instrumented with automatic retries with exponential backoff. A request will be retried as long as the request is deemed retryable and the number of retry attempts has not grown larger than the configured retry limit (default: 2).

Which status codes are retried depends on the retryStatusCodes generator configuration:

legacy (current default): retries on

  • 408 (Timeout)
  • 409 (Conflict)
  • 429 (Too Many Requests)
  • 5XX (All server errors, including 500)

recommended: retries on

  • 408 (Timeout)
  • 409 (Conflict)
  • 429 (Too Many Requests)
  • 502 (Bad Gateway)
  • 503 (Service Unavailable)
  • 504 (Gateway Timeout)

Use the max_retries request option to configure this behavior.

client.brand.retrieve(..., request_options={
    "max_retries": 1
})

Timeouts

The SDK defaults to a 60 second timeout. You can configure this with a timeout option at the client or request level.

from berrycrawl import Berrycrawl

client = Berrycrawl(..., timeout=20.0)

# Override timeout for a specific method
client.brand.retrieve(..., request_options={
    "timeout": 1
})

Custom Client

You can override the httpx client to customize it for your use-case. Some common use-cases include support for proxies and transports.

import httpx
from berrycrawl import Berrycrawl

client = Berrycrawl(
    ...,
    httpx_client=httpx.Client(
        proxy="http://my.test.proxy.example.com",
        transport=httpx.HTTPTransport(local_address="0.0.0.0"),
    ),
)

Contributing

While we value open-source contributions to this SDK, this library is generated programmatically. Additions made directly to this library would have to be moved over to our generation code, otherwise they would be overwritten upon the next generated release. Feel free to open a PR as a proof of concept, but know that we will not be able to merge it as-is. We suggest opening an issue first to discuss with us!

On the other hand, contributions to the README are always very welcome!

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

berrycrawl-0.2.0.tar.gz (71.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

berrycrawl-0.2.0-py3-none-any.whl (129.4 kB view details)

Uploaded Python 3

File details

Details for the file berrycrawl-0.2.0.tar.gz.

File metadata

  • Download URL: berrycrawl-0.2.0.tar.gz
  • Upload date:
  • Size: 71.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for berrycrawl-0.2.0.tar.gz
Algorithm Hash digest
SHA256 4c0330ee9f82d453cb0e12335a0883dbb3b3b4a436bc834d9e6fe29a2d0ee990
MD5 2ea1e4750682b4da4eacf284f285f10d
BLAKE2b-256 691401a29d3f0fae1f273f796e6a977d9b7b3bd0eb2ccda62d02b5f3f7e33434

See more details on using hashes here.

Provenance

The following attestation bundles were made for berrycrawl-0.2.0.tar.gz:

Publisher: ci.yml on strawberry-labs/berrycrawl-python

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file berrycrawl-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: berrycrawl-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 129.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for berrycrawl-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 304d4d81c6266990f8d8e1202ead1db1adde8d5977393a65a29cd1191efbcffb
MD5 775ac457e3233a51837e0cc16510633b
BLAKE2b-256 6240a7d166b2215876ef4dacdc640b547d29909bdbffb02505f191037f1b123d

See more details on using hashes here.

Provenance

The following attestation bundles were made for berrycrawl-0.2.0-py3-none-any.whl:

Publisher: ci.yml on strawberry-labs/berrycrawl-python

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page