The official Python SDK for Spidra
Project description
Spidra Python SDK
The official Python SDK for Spidra that allows you to scrape pages, run browser actions, batch-process URLs, and crawl entire sites. All results come back as structured data ready to feed into your LLM pipelines or store directly.
Installation
pip install spidra
Get your API key at app.spidra.io under Settings > API Keys.
Quick start
Synchronous (simplest)
from spidra import Spidra, ScrapeParams, ScrapeUrl
client = Spidra(api_key="spd_YOUR_API_KEY")
result = client.scrape(ScrapeParams(
urls=[ScrapeUrl(url="https://news.ycombinator.com")],
prompt="List the top 5 stories with title, points, and comment count",
output="json",
))
print(result.content)
In an async function
import asyncio
from spidra import AsyncSpidra, ScrapeParams, ScrapeUrl
async def main():
client = AsyncSpidra(api_key="spd_YOUR_API_KEY")
result = await client.scrape(ScrapeParams(
urls=[ScrapeUrl(url="https://news.ycombinator.com")],
prompt="List the top 5 stories with title, points, and comment count",
output="json",
))
print(result.content)
asyncio.run(main())
In a Jupyter Notebook
Jupyter already runs its own event loop, so await directly in a cell:
from spidra import AsyncSpidra, ScrapeParams, ScrapeUrl
client = AsyncSpidra(api_key="spd_YOUR_API_KEY")
result = await client.scrape(ScrapeParams(
urls=[ScrapeUrl(url="https://news.ycombinator.com")],
prompt="List the top 5 stories with title, points, and comment count",
output="json",
))
print(result.content)
Why? Jupyter's event loop is already running. Using bare
awaitin a cell works because Jupyter patches its loop to accept top-level awaits. Callingasyncio.run()would start a second loop, which Python does not allow.
Two clients
| Client | When to use |
|---|---|
Spidra |
Scripts, Django, Flask, CLI tools — anywhere you don't have an event loop |
AsyncSpidra |
FastAPI, async libraries, Jupyter notebooks |
Both have identical method signatures. All code examples below use AsyncSpidra with await. For sync usage, swap AsyncSpidra → Spidra and drop the await.
Table of contents
- Spidra Python SDK
Scraping
All scrape jobs run asynchronously on the server side. The scrape() method submits a job, polls until it finishes, and returns the result directly. If you need more control, use start_scrape() and get_scrape().
Up to 3 URLs can be passed per request and they are processed in parallel.
Basic scrape
from spidra import AsyncSpidra, ScrapeParams, ScrapeUrl
client = AsyncSpidra(api_key="spd_YOUR_API_KEY")
result = await client.scrape(ScrapeParams(
urls=[ScrapeUrl(url="https://example.com/pricing")],
prompt="Extract all pricing plans with name, price, and included features",
output="json",
))
print(result.content)
# { "plans": [{ "name": "Starter", "price": "$9/mo", "features": [...] }, ...] }
Structured output with JSON schema
When you need a guaranteed shape, pass a schema. The API will enforce the structure and return None for any missing fields rather than hallucinating values.
result = await client.scrape(ScrapeParams(
urls=[ScrapeUrl(url="https://jobs.example.com/senior-engineer")],
prompt="Extract the job listing details",
output="json",
schema={
"type": "object",
"required": ["title", "company", "remote"],
"properties": {
"title": { "type": "string" },
"company": { "type": "string" },
"remote": { "type": ["boolean", "null"] },
"salary_min": { "type": ["number", "null"] },
"salary_max": { "type": ["number", "null"] },
"skills": { "type": "array", "items": { "type": "string" } },
},
},
))
Geo-targeted scraping
Pass use_proxy=True and a proxy_country code to route the request through a specific country. Useful for geo-restricted content or localized pricing.
result = await client.scrape(ScrapeParams(
urls=[ScrapeUrl(url="https://www.amazon.de/gp/bestsellers")],
prompt="List the top 10 products with name and price",
use_proxy=True,
proxy_country="de",
))
Supported country codes include: us, gb, de, fr, jp, au, ca, br, in, nl, sg, es, it, mx, and 40+ more. Use "global" or "eu" for regional routing.
Authenticated pages
Pass cookies as a string to scrape pages that require a login session.
result = await client.scrape(ScrapeParams(
urls=[ScrapeUrl(url="https://app.example.com/dashboard")],
prompt="Extract the monthly revenue and active user count",
cookies="session=abc123; auth_token=xyz789",
))
Browser actions
Actions let you interact with the page before the scrape runs. They execute in order, and the scrape happens after all actions complete.
from spidra import BrowserAction
result = await client.scrape(ScrapeParams(
urls=[
ScrapeUrl(
url="https://example.com/products",
actions=[
BrowserAction(type="click", selector="#accept-cookies"),
BrowserAction(type="wait", duration=1000),
BrowserAction(type="scroll", to="80%"),
],
),
],
prompt="Extract all product names and prices",
))
Available actions:
| Action | Required fields | Description |
|---|---|---|
click |
selector or value |
Click a button, link, or any element |
type |
selector, value |
Type text into an input or textarea |
check |
selector or value |
Check a checkbox |
uncheck |
selector or value |
Uncheck a checkbox |
wait |
duration (ms) |
Pause execution for a set number of milliseconds |
scroll |
to (0–100%) |
Scroll the page to a percentage of its height |
forEach |
observe |
Loop over every matched element and process each one |
For selector, use a CSS selector or XPath. For value, use a plain English description and Spidra will locate the element using AI.
# CSS selector
BrowserAction(type="click", selector="button[data-testid='submit']")
# Plain English
BrowserAction(type="click", value="Accept all cookies button")
# Type into a field
BrowserAction(type="type", selector="input[name='q']", value="wireless headphones")
# Wait for content to load
BrowserAction(type="wait", duration=2000)
# Scroll to bottom
BrowserAction(type="scroll", to="100%")
forEach: process every element on a page
forEach finds a set of elements on the page and processes each one individually. It is the right tool when you need to collect data from a list of items, paginate through multiple pages, or click into each item's detail page.
You don't need
forEachif the data fits on a single page and is short — a plainpromptis simpler and works just as well.
Use forEach when:
- The list spans multiple pages and you need
pagination - You need to click into each item's detail page (
navigatemode) - You have 20+ items and want per-item AI extraction to stay consistent (
item_prompt)
inline mode
Read each element's content directly without navigating. Best for product cards, search results, table rows.
from spidra import BrowserAction
result = await client.scrape(ScrapeParams(
urls=[
ScrapeUrl(
url="https://books.toscrape.com/catalogue/category/books/mystery_3/index.html",
actions=[
BrowserAction(
type="forEach",
observe="Find all book cards in the product grid",
mode="inline",
capture_selector="article.product_pod",
max_items=20,
item_prompt="Extract title, price, and star rating. Return as JSON: {title, price, star_rating}",
),
],
),
],
prompt="Return a clean JSON array of all books",
output="json",
))
navigate mode
Follow each element's link to its destination page and capture content there. Best for product listings where the full detail is only on the individual page.
BrowserAction(
type="forEach",
observe="Find all book title links in the product grid",
mode="navigate",
capture_selector="article.product_page",
max_items=10,
wait_after_click=800,
item_prompt="Extract title, price, star rating, and availability. Return as JSON.",
)
click mode
Click each element, capture the content that appears (a modal, drawer, or expanded section), then move on. Best for hotel room cards, FAQ accordions, or any UI where clicking reveals hidden content.
BrowserAction(
type="forEach",
observe="Find all room type cards",
mode="click",
capture_selector="[role='dialog']",
max_items=8,
wait_after_click=1200,
item_prompt="Extract room name, bed type, price per night, and amenities. Return as JSON.",
)
Pagination
After processing all elements on the current page, follow the next-page link and continue collecting.
BrowserAction(
type="forEach",
observe="Find all book title links",
mode="navigate",
max_items=40,
pagination={
"next_selector": "li.next > a",
"max_pages": 3, # 3 additional pages beyond the first
},
)
max_items applies across all pages combined. The loop stops when you hit max_items, run out of pages, or reach max_pages.
Per-element actions
Run additional browser actions on each item after navigating or clicking into it, before the content is captured.
BrowserAction(
type="forEach",
observe="Find all book title links",
mode="navigate",
capture_selector="article.product_page",
max_items=5,
wait_after_click=1000,
actions=[
BrowserAction(type="scroll", to="50%"),
],
item_prompt="Extract title, price, and full description. Return as JSON.",
)
item_prompt vs top-level prompt
Both are optional and serve different purposes.
item_prompt |
prompt |
|
|---|---|---|
| When it runs | During scraping, once per item | After all items are collected |
| What it sees | One item's content | All items combined |
| Output location | Feeds into the top-level prompt |
result.content |
Manual job control
Use start_scrape() and get_scrape() when you want to manage polling yourself, or fire-and-forget and check back later.
# Submit a job and get the job_id immediately
queued = await client.start_scrape(ScrapeParams(
urls=[ScrapeUrl(url="https://example.com")],
prompt="Extract the main headline",
))
# Check status at any point
status = await client.get_scrape(queued.job_id)
if status.status == "completed":
print(status.result.content)
elif status.status == "failed":
print(status.error)
Job statuses: queued, waiting, active, completed, failed.
Poll options
scrape(), batch_scrape(), and crawl() accept poll_interval and timeout keyword arguments to control polling behaviour.
result = await client.scrape(
params,
poll_interval=3.0,
timeout=120.0,
)
Batch scraping
Submit up to 50 URLs in a single request. All URLs are processed in parallel. Each URL is a plain string.
from spidra import BatchScrapeParams
batch = await client.batch_scrape(BatchScrapeParams(
urls=[
"https://shop.example.com/product/1",
"https://shop.example.com/product/2",
"https://shop.example.com/product/3",
],
prompt="Extract product name, price, and availability",
output="json",
use_proxy=True,
))
for item in batch.items:
if item.status == "completed":
print(item.url, item.result)
elif item.status == "failed":
print(item.url, item.error)
Item statuses: pending, running, completed, failed.
Retry failed items:
queued = await client.start_batch_scrape(BatchScrapeParams(
urls=["https://example.com/1", "https://example.com/2"],
prompt="Extract the page title",
))
# Later, after checking status
result = await client.get_batch_scrape(queued.batch_id)
if result.failed_count > 0:
await client.retry_batch_scrape(queued.batch_id)
Cancel a running batch:
response = await client.cancel_batch_scrape(batch_id)
print(f"Cancelled {response.cancelled_items} items, refunded {response.credits_refunded} credits")
List past batches:
response = await client.list_batch_scrapes(page=1, limit=20)
for job in response.jobs:
print(job.uuid, job.status, f"{job.completed_count}/{job.total_urls}")
Crawling
Given a starting URL, Spidra discovers pages automatically according to your instruction and extracts structured data from each one.
from spidra import CrawlParams
job = await client.crawl(CrawlParams(
base_url="https://competitor.com/blog",
crawl_instruction="Find all blog posts published in 2024",
transform_instruction="Extract the title, author, publish date, and a one-sentence summary",
max_pages=30,
use_proxy=True,
))
for page in job.result:
print(page.url, page.data)
transform_instruction is optional. When omitted (and no schema is set), each page's data field contains the raw page markdown — no AI extraction is called and no token credits are charged for extraction.
All CrawlParams options:
| Parameter | Type | Description |
|---|---|---|
base_url |
str |
Required. Starting URL for the crawl. |
crawl_instruction |
str |
Which pages to discover. Defaults to "Find all pages on the website". |
transform_instruction |
str | None |
What to extract from each page. Omit to get raw markdown with no AI charges. |
schema |
dict | None |
JSON Schema for structured per-page output. Root must be type: object. |
max_pages |
int | None |
Cap on pages crawled. |
max_depth |
int | None |
Max link depth from the base URL. 0 = base URL only. |
include_paths |
list[str] | None |
Only crawl pages whose path matches one of these patterns. |
exclude_paths |
list[str] | None |
Skip pages whose path matches any of these patterns. |
allow_subdomains |
bool | None |
Follow links to subdomains of the base domain. |
crawl_entire_domain |
bool | None |
Follow any link on the same root domain regardless of path. |
ignore_query_params |
bool | None |
Treat URLs differing only by query string as the same page. |
webhook_url |
str | None |
URL that receives POST notifications as the job progresses. |
use_proxy |
bool | None |
Route requests through a residential proxy. |
proxy_country |
str | None |
Two-letter country code for geo-targeted proxy routing. |
cookies |
str | None |
Cookie string for authenticated crawls. |
scrape_mode |
"default" | "fast" | None |
"fast" skips JavaScript rendering for static pages. |
Submit without waiting:
queued = await client.start_crawl(CrawlParams(
base_url="https://example.com/docs",
crawl_instruction="Find all documentation pages",
max_pages=50,
))
# Check status later
status = await client.get_crawl(queued.job_id)
Limit depth and scope:
job = await client.crawl(CrawlParams(
base_url="https://example.com/blog",
crawl_instruction="Find all blog posts",
max_depth=2,
include_paths=["/blog/"],
exclude_paths=["/blog/tag/", "/blog/author/"],
ignore_query_params=True,
))
Get signed download URLs for all crawled pages:
Each page includes html and markdown fields with S3-signed URLs that expire after 1 hour.
response = await client.crawl_pages(job_id)
for page in response.pages:
print(page.url, page.status)
# Download raw HTML: page.html
# Download markdown: page.markdown
Re-extract with a new instruction:
Runs a new AI transformation over an existing completed crawl without re-crawling any pages. Charges credits for the transformation only.
queued = await client.crawl_extract(source_job_id, "Extract only the product SKUs and prices as a CSV")
# Poll the new job manually
result = await client.get_crawl(queued.job_id)
Crawl history and stats:
response = await client.crawl_history(page=1, limit=10)
stats = await client.crawl_stats()
print(f"Total crawls: {stats.total}")
Logs
Scrape logs are stored for every job that runs through the API.
# List logs with optional filters
response = await client.scrape_logs(
status="failed",
search_term="amazon.com",
start_date="2024-01-01",
end_date="2024-12-31",
page=1,
limit=20,
)
for log in response.logs:
print(log.urls[0].get("url"), log.status, log.credits_used)
Get a single log with full extraction result:
log = await client.get_scrape_log("log-uuid")
print(log.result_data) # the full AI output for that job
Usage statistics
Returns credit and request usage broken down by day or week.
# Range options: "7d" | "30d" | "weekly"
rows = await client.usage("30d")
for row in rows:
print(row.date, row.requests, row.credits, row.tokens)
Error handling
Every API error raises a typed exception. Catch the specific class you care about or fall back to the base SpidraError.
from spidra import (
AsyncSpidra,
SpidraError,
SpidraAuthenticationError,
SpidraForbiddenError,
SpidraRateLimitError,
SpidraServerError,
ScrapeParams,
ScrapeUrl,
)
try:
result = await client.scrape(ScrapeParams(
urls=[ScrapeUrl(url="https://example.com")],
prompt="...",
))
except SpidraAuthenticationError:
# 401: No x-api-key header sent
print("Check your API key")
except SpidraForbiddenError:
# 403: Invalid API key, or monthly credit limit reached — check e.message to distinguish
print("Access denied or out of credits")
except SpidraRateLimitError:
# 429: Too many requests
print("Rate limited, back off and retry")
except SpidraServerError:
# 500: Something went wrong on Spidra's side
print("Server error, try again")
except SpidraError as e:
# Any other API error
print(f"{e.status}: {e.message}")
All error classes expose err.status (HTTP status code) and err.message.
Debugging
Enable debug logging to see every HTTP request, response, and retry attempt:
import logging
logging.getLogger("spidra").setLevel(logging.DEBUG)
Sample output:
DEBUG:spidra:POST /scrape (attempt 1/4)
DEBUG:spidra:Response 200 in 1.23s
Context manager
Use AsyncSpidra as an async context manager to ensure the HTTP connection pool is properly closed.
async with AsyncSpidra(api_key="spd_YOUR_API_KEY") as client:
result = await client.scrape(ScrapeParams(
urls=[ScrapeUrl(url="https://example.com")],
prompt="Extract the page title",
))
print(result.content)
Requirements
- Python 3.9 or later
- A Spidra API key (sign up free)
License
MIT
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file spidra-0.3.0.tar.gz.
File metadata
- Download URL: spidra-0.3.0.tar.gz
- Upload date:
- Size: 42.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: Hatch/1.17.0 {"ci":null,"cpu":"arm64","distro":{"name":"macOS","version":"26.5.1"},"implementation":{"name":"CPython","version":"3.12.7"},"installer":{"name":"hatch","version":"1.17.0"},"openssl_version":"OpenSSL 3.0.15 3 Sep 2024","python":"3.12.7","system":{"name":"Darwin","release":"25.5.0"}} HTTPX2/2.5.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
af142aafded6e170e757e50c65dc12a533883e34d69a85e827e911274f481b26
|
|
| MD5 |
f86acfdc9bf57a7534be83443e85265c
|
|
| BLAKE2b-256 |
b4cffc06801d11057858bf0a530baf16c23739a52b88f2bc1946547747550643
|
File details
Details for the file spidra-0.3.0-py3-none-any.whl.
File metadata
- Download URL: spidra-0.3.0-py3-none-any.whl
- Upload date:
- Size: 38.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: Hatch/1.17.0 {"ci":null,"cpu":"arm64","distro":{"name":"macOS","version":"26.5.1"},"implementation":{"name":"CPython","version":"3.12.7"},"installer":{"name":"hatch","version":"1.17.0"},"openssl_version":"OpenSSL 3.0.15 3 Sep 2024","python":"3.12.7","system":{"name":"Darwin","release":"25.5.0"}} HTTPX2/2.5.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
554fcb05bbae2ecdba2ec3a9ecae0cfdb49d4baa004711e2815eb162dee494bb
|
|
| MD5 |
2471939d55d16dc8d13082878a54960b
|
|
| BLAKE2b-256 |
2b0366dd8e803c2096764f684fc6021090e8d7e32b91b121ac788d5d417524d6
|