Skip to main content

litescrape-sdk

Python SDK for the Litescrape API: validated, concurrent, retrying calls to the Google, Bing, DuckDuckGo, Yelp, Tripadvisor, Apple Maps, Google Play, and Apple App Store endpoints, with results returned in input order.

pip install litescrape-sdk
import os
from litescrape_sdk import GoogleMaps, scrape

os.environ["LITESCRAPE_API_KEY"] = "ls_live_..."

results = scrape(
    [
        {"endpoint": "google_search", "q": "coffee grinders", "gl": "us"},
        GoogleMaps(q="coffee", type="search", ll="@40.745,-73.988,14z"),
    ]
)
for result in results:
    print(result.ok, result.data or result.error)

help(litescrape_sdk.scrape) documents every argument, the retry and concurrency rules, and the error types. ascrape is the same function for asyncio code, and litescrape_sdk.REQUEST_TYPES maps each endpoint slug to its request class.

Development: uv sync, then uv run pytest, uv run ruff check ., and uv run ruff format ..

Durable batches and recovery

Requires an API deployment with durable batch support. Set batched=True to submit every query before polling for completion. The server continues processing accepted jobs if your computer disconnects or the Python process exits. Responses are retained for 24 hours after completion.

from litescrape_sdk import GoogleSearch, scrape

queries = [GoogleSearch(q="coffee grinders"), GoogleSearch(q="espresso machines")]
results = scrape(queries, batched=True, cache_path="./my-batch.sqlite3")

# After a crash or disconnect, rerun with the same inputs, key and cache path:
results = scrape(queries, batched=True, useCache=True, cache_path="./my-batch.sqlite3")
for result in results:
    print(result.job_id, result.data or result.error)

The SDK always saves submission intentions and job IDs locally in batch mode. useCache=False is the default and creates fresh jobs. useCache=True retrieves matching saved jobs; it creates replacements only when the server confirms a job is missing or expired. A polling outage does not trigger a fresh charge. Reordering distinct queries preserves their cache matches; identical queries are matched by occurrence. Keep separate cache files for independently resumable batches. The default path is ~/.cache/litescrape/jobs.sqlite3, configurable with LITESCRAPE_JOB_CACHE. The cache stores job IDs and request fingerprints, without raw API keys or query text.

New jobs reserve one credit; terminal failures refund that reservation. Polling and retrying the same submission do not consume more credits. Batch mode skips the synchronous whole-list balance check so previously paid results can be recovered even with no credits remaining. Check every Result for submission errors, including insufficient credits for new jobs.

concurrency controls simultaneous submission/poll requests (default 32); execution uses separate server batch capacity. Network retries and polling use exponential backoff with jitter, capped at 30 seconds; retries honor Retry-After. All results are returned in input order. ascrape accepts the same options. A server request_timeout applies to each worker attempt after queueing.

Google Search fast mode

Google Search supports GoogleSearch(q="coffee grinders", fast_mode=True) to return only organic results plus search metadata and parameters. This skips AI Overview and other result groups. The default is the full response; fast mode is unavailable on the dedicated AI Overview endpoint. Requires an API deployment that supports fast_mode.

Request deadlines

Set request_timeout on scrape or ascrape to apply a server deadline to every item. Set timeout on an individual request to override that default. Both accept seconds greater than 0 and at most 90, including fractional seconds; omitting them preserves the API's standard behavior.

from litescrape_sdk import GoogleSearch, RequestDeadlineExceededError, scrape

(result,) = scrape(
    [GoogleSearch(q="coffee grinders")],
    request_timeout=15,
    attempts=1,
)
if isinstance(result.error, RequestDeadlineExceededError):
    print(result.error.request_id, result.error.retryable, result.error.message)
else:
    print(result.raise_for_error())

# A per-item deadline also works in dictionaries.
results = scrape(
    [
        GoogleSearch(q="coffee", timeout=10),
        {"endpoint": "google_search", "q": "tea", "timeout": 20},
    ],
    attempts=1,
)

The deadline covers the entire API request from gateway receipt, including admission, scraping, and billing. On expiry, the API returns HTTP 503 with the new request_deadline_exceeded code, retryable: true, a request ID, and a description. That attempt is not charged. The API cancels the underlying scrape to release concurrency; credit and concurrency cleanup can finish shortly after the response.

The SDK retries this response using its usual attempts and Retry-After rules. Each attempt gets its own deadline; use attempts=1 as above to receive the first timeout immediately. SDK queueing, the status check, network transit, retry waits, and the full batch are outside the server deadline.

The existing scrape(..., timeout=120) argument remains the HTTP transport timeout. Keep it longer than the server deadline so the API can return its error. A client-side TransportError alone does not guarantee that the request was unbilled.

Store APIs (Alpha)

All nine Store operations use your existing key. Alpha fields depend on what the storefront supplies.

Request class Endpoint slug
GooglePlayApps google_play_apps
GooglePlayGames google_play_games
GooglePlayBooks google_play_books
GooglePlayMovies google_play_movies
GooglePlayProduct google_play_product
GooglePlayReviews google_play_reviews
AppleAppStoreSearch apple_app_store_search
AppleAppStoreProduct apple_app_store_product
AppleAppStoreReviews apple_app_store_reviews
from litescrape_sdk import GooglePlayApps, GooglePlayProduct, AppleAppStoreSearch, scrape

results = scrape(
    [
        GooglePlayApps(q="coffee", hl="en", gl="us"),
        GooglePlayProduct(product_id="com.duolingo"),
        AppleAppStoreSearch(term="coffee", country="us", num=10),
    ]
)
for result in results:
    print(result.raise_for_error())

Follow litescrape_pagination.next or use the returned continuation with the same operation and parameters. Each successful page consumes one call; failures do not. For Google Play, chart, next_page_token, section_page_token, and see_more_token are mutually exclusive. Queries exclude category filters and charts. Search text is limited to 2,048 UTF-8 bytes, so non-ASCII characters can consume more than one byte. Apple search also limits URL-encoded terms to 4,096 bytes. Apps/games charts require store_device="phone" or an omitted device; other device storefronts are browsable without a chart. Omit store_device when using an apps or games query or category. GooglePlayGames(q=...) uses the shared Android app search; omit q or choose games_category to browse games. Apple review pages are one-based; exhausted pages return an empty list. Mac reviews use newest-first ordering.

search_metadata.raw_file and prettify_file, when returned, link to authenticated response artifacts retained for at least seven days (today and the previous seven UTC date buckets). Download them with the same bearer key; downloads are unbilled. Apple search applies category and case-insensitive developer-name filters before the num ceiling. The native search window can contain fewer matching results than that ceiling. See the API reference for every parameter and response group.

Release files for litescrape-sdk 0.5.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for litescrape-sdk 0.5.0
File Size Uploaded
litescrape_sdk-0.5.0.tar.gz 82.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for litescrape-sdk 0.5.0
File Interpreter ABI Platform
litescrape_sdk-0.5.0-py3-none-any.whl Python 3 none any Details

Total release size: 111.9 kB

Release files / litescrape_sdk-0.5.0.tar.gz

Download URL litescrape_sdk-0.5.0.tar.gz
Size 82.3 kB
Tags Source
SHA-256 checksum
How to use checksums
e0a2baa2d1616b784588015907473f1a4a32b60b57804813f931cc8d1d5c55d3
BLAKE2b-256 checksum
How to use checksums
5b6a15ce931aa7a9038dec8e542095c2fea6212588d79bdeac4ad5d2da583577
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 12, 2026.

Transparency log

Release files / litescrape_sdk-0.5.0-py3-none-any.whl

Download URL litescrape_sdk-0.5.0-py3-none-any.whl
Size 29.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9e38bb7ddea7f19355594927f400da29d58631a98c24764dc417acb47b42caf6
BLAKE2b-256 checksum
How to use checksums
1cd689b1e76427f6baf057e1a48ca7b4b1c7f7332c80a69c1f0620c04b1aee05
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 12, 2026.

Transparency log

Release history Release notifications | RSS feed

0.5.2

2 release files

0.5.1

2 release files

This release

0.5.0 This release

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page