Skip to main content

scrapedatshi-py

Official Python SDK for the scrapedatshi RAG pipeline API.

Scrape URLs, chunk documents, embed content, inject into vector databases, and extract structured data — all from a clean, typed Python interface.


Installation

pip install scrapedatshi

Requires Python 3.10+.


Quick Start

from scrapedatshi import ScrapedatshiClient

client = ScrapedatshiClient(api_key="sds_...")

# Chunk a URL to JSON (no embedding required)
result = client.pipeline.chunk_url("https://docs.example.com")

print(f"Got {result.total_chunks} chunks")
print(f"Cost: ${result.credits_used:.4f} | Remaining: ${result.credits_remaining:.4f}")
for chunk in result.chunks:
    print(chunk.content[:80])

Authentication

Pass your API key directly or set the SCRAPEDATSHI_API_KEY environment variable:

export SCRAPEDATSHI_API_KEY="sds_..."
# Explicit key
client = ScrapedatshiClient(api_key="sds_...")

# From environment variable
client = ScrapedatshiClient()

Get your API key at scrapedatshi.com/portal/register. New accounts receive $1.00 free credits — no credit card required.


Pricing

scrapedatshi uses a pay-per-use credit wallet — no subscriptions, no monthly fees. Credits are deducted after each successful API call. Failed requests are never charged.

Operation Rate Applies To
URL Fetch $0.0020 / URL /v1/rag-chunk, /v1/crawl-chunk, /v1/sync, /v1/ingest
Spider Fetch $0.0050 / URL /v1/spider (replaces URL fetch)
Chunk Fee $0.0005 / chunk All routes — per chunk generated
Injection Fee $0.0030 / chunk /v1/sync, /v1/ingest — per chunk upserted to vector DB
Contextual Retrieval $0.0030 / URL When contextual_retrieval=True
JS Render $0.0050 / URL When js_render=True (Playwright headless browser)
Schema Extract $0.0020 + $0.0030 + ($0.0001 × fields) /v1/extract

Top up your balance at scrapedatshi.com/portal/billing.


Pipeline Methods

Chunk to JSON

No embedding or vector DB required. Returns structured JSON chunks from any source.

Chunk a URL

result = client.pipeline.chunk_url("https://docs.example.com")

# result.chunks              → list[Chunk]
# result.total_chunks        → int
# result.source              → str (the URL)
# result.credits_used        → float
# result.credits_remaining   → float
# result.content_truncated   → bool (True if content exceeded ~75,000 words)

Chunk a URL with JS rendering

For JavaScript-heavy pages and SPAs that require a browser to render:

result = client.pipeline.chunk_url(
    "https://spa.example.com/dashboard",
    js_render=True,
)

Chunk a local file

Supports PDF, MD, TXT, YAML, YML, and JSON.

result = client.pipeline.chunk_file("./docs/manual.pdf")

print(f"Got {result.total_chunks} chunks from {result.source}")
print(f"Cost: ${result.credits_used:.4f}")

Crawl a website

Crawls via sitemap or spider and chunks all pages.

# Sitemap crawl (default) — reads sitemap.xml
result = client.pipeline.crawl("https://example.com", max_pages=10)

# Spider crawl — follows links, works on any site
result = client.pipeline.crawl(
    "https://example.com",
    crawl_mode="spider",
    max_pages=5,
    include_pattern="/docs/",
    exclude_pattern="/blog/",
)

print(f"Crawled {result.pages_crawled} pages → {result.total_chunks} chunks")
print(f"Cost: ${result.credits_used:.4f}")

Full Pipeline — Embed + Inject

Scrape, embed, and inject directly into your vector database in one call.

Sync a URL

result = client.pipeline.sync(
    url="https://docs.example.com",
    embedding_provider="openai",
    embedding_api_key="sk-...",
    vector_db="pinecone",
    vector_db_config={
        "api_key": "pc-...",
        "index_host": "https://my-index-abc123.svc.pinecone.io",
    },
)

print(f"Upserted {result.vectors_upserted} vectors ({result.total_tokens} tokens)")
print(f"Cost: ${result.credits_used:.4f}")

Ingest a local file

result = client.pipeline.ingest(
    file_path="./docs/manual.pdf",
    embedding_provider="openai",
    embedding_api_key="sk-...",
    vector_db="qdrant",
    vector_db_config={
        "url": "https://your-cluster.qdrant.io",
        "collection_name": "documents",
        "api_key": "qdrant-key",  # optional for local Qdrant
    },
)

Schema Extraction

Extract structured data from any URL using your own LLM key. Define a schema and the API returns a typed JSON object — or a list of objects for pages with multiple items.

Extract a single object

result = client.pipeline.extract(
    url="https://example.com/products/widget-pro",
    schema={
        "title": "string — the product name",
        "price": "number — the price in USD",
        "in_stock": "boolean — whether the item is in stock",
        "description": "string — the product description",
    },
    llm_provider="openai",
    llm_api_key="sk-...",
)

print(result.extracted)
# → {"title": "Widget Pro", "price": 29.99, "in_stock": True, "description": "..."}
print(f"Cost: ${result.credits_used:.4f}")

Extract a list of items

Use extract_as_list=True for pages with multiple matching items (product listings, article feeds, search results):

result = client.pipeline.extract(
    url="https://example.com/products",
    schema={
        "title": "string — the product name",
        "price": "number — the price in USD",
    },
    llm_provider="openai",
    llm_api_key="sk-...",
    extract_as_list=True,
)

print(f"Extracted {result.item_count} products")
for product in result.extracted:
    print(f"  {product['title']}: ${product['price']}")

Extract from a JS-rendered page

result = client.pipeline.extract(
    url="https://spa.example.com/data",
    schema={"value": "string — the data value"},
    llm_provider="anthropic",
    llm_api_key="sk-ant-...",
    js_render=True,
)

Contextual Retrieval (RAG 2.0)

Prepend an LLM-generated document summary to every chunk before embedding, dramatically improving retrieval accuracy.

result = client.pipeline.chunk_url(
    "https://docs.example.com",
    contextual_retrieval=True,
    llm_provider="openai",
    llm_api_key="sk-...",
    llm_model="gpt-4o-mini",
)

Available on all pipeline methods: chunk_url(), chunk_file(), crawl(), sync(), ingest().


Supported Providers

Discover all supported providers programmatically:

from scrapedatshi.providers import (
    EMBEDDING_PROVIDERS,
    VECTOR_DB_PROVIDERS,
    LLM_PROVIDERS,
)

# List all embedding providers
for key, info in EMBEDDING_PROVIDERS.items():
    print(f"{key}: {info['label']} — default model: {info['default_model']}")
    print(f"  {info['notes']}")

# Check required fields for a vector DB
print(VECTOR_DB_PROVIDERS["pinecone"]["required_fields"])
# → ["api_key", "index_host"]

Embedding Providers

Key Provider Default Model Notes
openai OpenAI text-embedding-3-small 1536 dims. Also: text-embedding-3-large (3072), text-embedding-ada-002 (1536)
cohere Cohere embed-english-v3.0 1024 dims. Also: embed-multilingual-v3.0, embed-english-light-v3.0 (384)
gemini Google Gemini gemini-embedding-001 3072 dims. Also: text-embedding-004 (768)
mistral Mistral mistral-embed 1024 dims
voyage Voyage AI voyage-3 1024 dims. Also: voyage-3-lite (512), voyage-code-3, voyage-finance-2, voyage-law-2

Models are discovered dynamically after key verification — you always get the latest available models for your account.

Vector Database Providers

Key Provider Required Fields
pinecone Pinecone api_key, index_host
qdrant Qdrant url, collection_name
supabase Supabase (pgvector) connection_string, table_name
weaviate Weaviate url, class_name
mongodb MongoDB Atlas connection_string, database_name, collection_name
azure_cosmos Azure Cosmos DB (NoSQL) connection_string, database_name, container_name
azure_cosmos_mongo Azure Cosmos DB (MongoDB API) connection_string, database_name, collection_name

LLM Providers (for Contextual Retrieval & Schema Extraction)

Key Provider Default Model
openai OpenAI gpt-4o-mini
anthropic Anthropic claude-3-haiku-20240307
gemini Google Gemini gemini-1.5-flash

Async Support

All methods have an _async variant for use with asyncio.

import asyncio
from scrapedatshi import ScrapedatshiClient

async def main():
    async with ScrapedatshiClient(api_key="sds_...") as client:
        result = await client.pipeline.chunk_url_async("https://docs.example.com")
        print(f"Got {result.total_chunks} chunks — cost ${result.credits_used:.4f}")

asyncio.run(main())

Parallel processing with asyncio.gather

async def main():
    async with ScrapedatshiClient(api_key="sds_...") as client:
        urls = [
            "https://docs.example.com/page1",
            "https://docs.example.com/page2",
            "https://docs.example.com/page3",
        ]
        results = await asyncio.gather(
            *[client.pipeline.chunk_url_async(url) for url in urls]
        )
        total = sum(r.total_chunks for r in results)
        total_cost = sum(r.credits_used for r in results)
        print(f"Processed {len(urls)} URLs → {total} total chunks — total cost ${total_cost:.4f}")

Response Models

All methods return typed Pydantic models with full IDE autocomplete support. Every response includes credits_used and credits_remaining for programmatic spend tracking.

ChunkResult

result.chunks                  # list[Chunk]
result.total_chunks            # int
result.source                  # str
result.contextual_retrieval_used  # bool
result.content_truncated       # bool — True if content exceeded ~75,000 words
result.credits_used            # float — credits deducted for this request
result.credits_remaining       # float — account balance after this request

Chunk

chunk.content              # str — the chunk text
chunk.token_estimate       # int — estimated token count
chunk.metadata             # dict — source URL, page number, etc.

CrawlChunkResult

result.chunks              # list[Chunk]
result.total_chunks        # int
result.pages_crawled       # int
result.source_url          # str
result.credits_used        # float
result.credits_remaining   # float

SyncResult / IngestResult

result.status              # "success" | "partial" | "error"
result.chunks_created      # int
result.vectors_upserted    # int
result.total_tokens        # int
result.embedding_provider  # str
result.vector_db_provider  # str
result.credits_used        # float
result.credits_remaining   # float

ExtractResult

result.extracted           # dict | list[dict] — the extracted data
result.field_count         # int — number of schema fields
result.item_count          # int | None — number of items (list mode only)
result.is_list             # bool — True if extracted is a list
result.url                 # str — the URL that was scraped
result.llm_provider        # str
result.llm_model           # str
result.schema_fields       # list[str] — field names from your schema
result.js_render           # bool — whether JS rendering was used
result.content_warning     # str | None — warning if content may be incomplete
result.credits_used        # float
result.credits_remaining   # float

Error Handling

from scrapedatshi.exceptions import (
    AuthError,              # Invalid or missing API key (401/403)
    InsufficientCreditsError,  # Balance too low — top up at portal/billing (402)
    RateLimitError,         # Per-request hard cap or rate limit exceeded (429)
    ValidationError,        # Bad request payload (422)
    ServerError,            # API server error (5xx)
    TimeoutError,           # Request timed out
    ScrapedatshiError       # Base exception — catch-all
)

try:
    result = client.pipeline.sync(
        url="https://docs.example.com",
        embedding_provider="openai",
        embedding_api_key="sk-...",
        vector_db="pinecone",
        vector_db_config={"api_key": "pc-...", "index_host": "https://..."},
    )
except InsufficientCreditsError:
    print("Balance too low — top up at scrapedatshi.com/portal/billing")
except RateLimitError as e:
    print(f"Rate limit hit: {e.message}")
except ScrapedatshiError as e:
    print(f"API error {e.status_code}: {e.message}")

Hard Caps

Per-request hard caps protect server stability and apply to all accounts:

Cap Limit
Max pages / crawl 25
Max pages / spider 25
Max chunks / request 10,000
Max content size ~75,000 words (auto-truncated)

Exceeding a hard cap returns HTTP 400. Content exceeding the size limit is automatically truncated — check result.content_truncated to detect this.


Development

git clone https://github.com/mxchris18/scrapedatshi-py
cd scrapedatshi-py
pip install -e ".[dev]"
pytest

License

MIT — see LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

scrapedatshi-0.2.0.tar.gz (20.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

scrapedatshi-0.2.0-py3-none-any.whl (21.6 kB view details)

Uploaded Python 3

File details

Details for the file scrapedatshi-0.2.0.tar.gz.

File metadata

  • Download URL: scrapedatshi-0.2.0.tar.gz
  • Upload date:
  • Size: 20.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: Hatch/1.17.0 {"ci":null,"cpu":"AMD64","implementation":{"name":"CPython","version":"3.13.12"},"installer":{"name":"hatch","version":"1.17.0"},"openssl_version":"OpenSSL 3.0.18 30 Sep 2025","python":"3.13.12","system":{"name":"Windows","release":"11"}} HTTPX2/2.4.0

File hashes

Hashes for scrapedatshi-0.2.0.tar.gz
Algorithm Hash digest
SHA256 5e8024c137d82c4354cd5c1a85dd592d1f5db7af2e578747071985251e039cc3
MD5 db5af3aedb3c726c4b52c1a39cb838b9
BLAKE2b-256 3c18e106e42526a75d374233395e6bba772eedd4314bc6bc6cf87f411260d601

See more details on using hashes here.

File details

Details for the file scrapedatshi-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: scrapedatshi-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 21.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: Hatch/1.17.0 {"ci":null,"cpu":"AMD64","implementation":{"name":"CPython","version":"3.13.12"},"installer":{"name":"hatch","version":"1.17.0"},"openssl_version":"OpenSSL 3.0.18 30 Sep 2025","python":"3.13.12","system":{"name":"Windows","release":"11"}} HTTPX2/2.4.0

File hashes

Hashes for scrapedatshi-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 3fffd888ed8ca9d7c5fab380f86859ff382be7b48631b1a79b7fe7b3524d6895
MD5 e6ff25ce4cc766568113514e53128ff3
BLAKE2b-256 abb4fb4310aa11c4b5de6b451e0a8176250a0d06edb739184e9555a5318cfcdb

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page