Skip to main content

Scraping AI Python SDK (scraping-ai)

A modern, type-safe Python SDK for interacting with Scraping AI service API endpoints. It supports both synchronous and asynchronous operations using httpx and pydantic.

This SDK focuses strictly on core service execution endpoints (pipeline orchestration, web extraction, crawling, finder, keywords, etc.), excluding account management, billing, payments, and administrative routes.

Installation

pip install scraping-ai

Quick Start (High-Level 1-Line Helpers)

1. Web Extraction in 1 Line

from scraping_ai import ScrapingAIClient

client = ScrapingAIClient(api_key="your_api_key_here")

# Extract web data directly in one step (polls until finished)
data = client.extract(url="https://example.com", schema={"title": "string", "price": "number"})
print(data.results)

2. Page Crawling in 1 Line

# Crawl a web page and get extracted content & markdown
crawl_result = client.crawl(url="https://example.com", use_browser=True)
print(crawl_result.results)

3. Discover URLs in 1 Line

# Discover URLs on a target site up to a specific depth
found_urls = client.find_urls(base_url="https://example.com", max_depth=2)
print(found_urls.results)

4. Keyword Generation in 1 Line

# Generate relevant search keywords for a topic
keywords_res = client.generate_keywords(context="dummy_context", num_keywords=15)
print(keywords_res.keywords)

5. Schema Generation in 1 Line

# Generate JSON Schema for an extraction prompt
schema_res = client.generate_schema(schema_instruction="Extract title, author, date")
print(schema_res.schema_)

Asynchronous Usage

import asyncio
from scraping_ai import AsyncScrapingAIClient


async def main():
    async with AsyncScrapingAIClient(api_key="your_api_key_here") as client:
        # Async one-liner extraction
        data = await client.extract(url="https://example.com")
        print(data.results)


asyncio.run(main())

Service Modules Overview

The SDK exposes low-level module controllers under clean, intuitive client attributes:

  • 1-Line Helpers: client.extract(...), client.crawl(...), client.find_urls(...), client.generate_keywords(...), client.generate_schema(...).
  • client.extractor: Data extraction module (run, status, result).
  • client.crawler: Page crawling module (run, status, result).
  • client.finder: URL discovery module (run, status, result).
  • client.keywords: Keyword generation module (run, status, result).
  • client.schema: JSON Schema generation module (run, status, result).
  • client.ranker: Relevance ranking module (run, status, result).
  • client.search_url: Search URL generator module (run, status, result).
  • client.search_results: Search results discovery module (run, status, result).
  • client.exports: S3 exports management (create, list, get).
  • client.pipeline: Multi-step pipeline orchestration (create, list, get, patch, run_flow, wait_for_task).
  • client.data: Query extracted records (get_by_state, export_by_state).

Metadata

Release files for scraping-ai 0.4.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for scraping-ai 0.4.3
File Size Uploaded
scraping_ai-0.4.3.tar.gz 13.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for scraping-ai 0.4.3
File Interpreter ABI Platform
scraping_ai-0.4.3-py3-none-any.whl Python 3 none any Details

Total release size: 24.4 kB

Release files / scraping_ai-0.4.3.tar.gz

Download URL scraping_ai-0.4.3.tar.gz
Size 13.0 kB
Tags Source
SHA-256 checksum
How to use checksums
ffa367c9aba11d7de102d0eba57d70d78055569c3fed9872c6245e81273452c5
BLAKE2b-256 checksum
How to use checksums
d28622ebe5cb4a1a7020b357df33780272c6d85a9631be6289bca8e47709a4e4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / scraping_ai-0.4.3-py3-none-any.whl

Download URL scraping_ai-0.4.3-py3-none-any.whl
Size 11.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8f51708a577cfc5ce2789aba39afc186ff15d5f8cb9d441dbe1cb88811255992
BLAKE2b-256 checksum
How to use checksums
38416056e44964ccd0f0b95d7195cbb04aeb430956a931006d60a146280ffefa
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.4.3 This release

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.2.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page