Skip to main content

FireScraper Python SDK

Official Python SDK for FireScraper — web scraping built for AI pipelines.

Turn websites into clean, structured text for RAG, fine-tuning, and AI agent workflows.

Installation

pip install firescraper

With LangChain integration:

pip install firescraper langchain-firescraper

Quick Start

from firescraper import FireScraper

client = FireScraper("fsk_your_api_key")

# Start a crawl
session = client.scrape(
    name="Docs crawl",
    urls=["https://docs.example.com/"],
    max_depth=2,
    scraper="article",
)

# Wait for completion
result = client.wait_for_completion(session.id)
print(f"Scraped {result.counts.success} pages")

# Download results
download = client.get_results(session.id, format="json")
with open("results.json", "wb") as f:
    f.write(download.data)

Async Usage

from firescraper import AsyncFireScraper

async with AsyncFireScraper("fsk_your_api_key") as client:
    session = await client.scrape(
        name="Async crawl",
        urls=["https://example.com/"],
        max_depth=1,
    )
    result = await client.wait_for_completion(session.id)
    download = await client.get_results(session.id, format="markdown")

LangChain Integration

from langchain_firescraper import FireScraperLoader

loader = FireScraperLoader(
    api_key="fsk_your_api_key",
    urls=["https://docs.example.com/"],
    max_depth=2,
)

# Load all documents
docs = loader.load()
for doc in docs:
    print(doc.metadata["url"], len(doc.page_content))

# Or stream with lazy_load
for doc in loader.lazy_load():
    process(doc)

API Reference

FireScraper(api_key, *, base_url, timeout)

Parameter Type Default Description
api_key str required API key (starts with fsk_)
base_url str https://firescraper.com API base URL
timeout float 30.0 HTTP request timeout in seconds

Methods

scrape(name, urls, max_depth=1, scraper="article", **kwargs)

Start a new crawl session.

Parameter Type Default Description
name str required Human-readable session name
urls list[str] required Seed URLs
max_depth int 1 Link-hop depth (0 = seeds only)
scraper str "article" "article" or "full"
ignore_urls list[str] None URLs to exclude
webhook_url str None Callback URL on completion
extraction_schema dict None JSON Schema for structured extraction
respect_robots_txt bool None Respect robots.txt
content_selector str None CSS selector for extraction

Returns a ScrapeResponse with .id, .status, .message.

get_session(session_id)

Get current session status, including page counts and processing state.

wait_for_completion(session_id, poll_interval=5, timeout=300, on_progress=None)

Poll until the session reaches a terminal status (done, error, etc.).

def progress(status):
    print(f"{status.counts.success}/{status.counts.total} pages")

result = client.wait_for_completion(session.id, on_progress=progress)

list_results(session_id)

List available result files for a completed session.

get_results(session_id, format="json")

Download results. Supported formats: zip, csv, json, markdown, structured, manifest, documents, chunks, extracted. Use documents for page-level JSONL output.

get_partial_results(session_id, format="csv")

Download mid-crawl results while the session is still running.

Error Handling

from firescraper import FireScraperError, AuthenticationError, RateLimitError

try:
    session = client.scrape(name="Test", urls=["https://example.com"])
except AuthenticationError:
    print("Invalid API key")
except RateLimitError:
    print("Rate limited — try again later")
except FireScraperError as e:
    print(f"API error: {e.message} (code={e.code}, status={e.status})")

Advanced: Progress Tracking

session = client.scrape(name="Large crawl", urls=urls, max_depth=5)

result = client.wait_for_completion(
    session.id,
    poll_interval=3,
    timeout=600,
    on_progress=lambda s: print(
        f"[{s.session.status}] {s.counts.success} pages, "
        f"queue: {s.processing.queue_length}"
    ),
)

License

MIT

Metadata

Release files for firescraper 1.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for firescraper 1.0.1
File Size Uploaded
firescraper-1.0.1.tar.gz 16.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for firescraper 1.0.1
File Interpreter ABI Platform
firescraper-1.0.1-py3-none-any.whl Python 3 none any Details

Total release size: 30.2 kB

Release files / firescraper-1.0.1.tar.gz

Download URL firescraper-1.0.1.tar.gz
Size 16.6 kB
Tags Source
SHA-256 checksum
How to use checksums
70e937698d36dd0b111e8c1077c751675ce9168a558ece7d0f8b15d89d451e3e
BLAKE2b-256 checksum
How to use checksums
b9b4bbdc9a50df818dc13da9100ccfe258533111d5a30235014e2c7f2ef61377
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.3

Release files / firescraper-1.0.1-py3-none-any.whl

Download URL firescraper-1.0.1-py3-none-any.whl
Size 13.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
08d343c853d31f509efdcd6b3fd89894a5d4d31769584c5f1883e1a4493b0b41
BLAKE2b-256 checksum
How to use checksums
bac5cf4bc624aadb40767ea5ea3461e03060154802b8995f329b8c9fcca72ea2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.3

Release history Release notifications | RSS feed

This release

1.0.1 This release

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page