Skip to main content

crawl4ai-client

Lightweight async Python client for Crawl4AI Docker server.

No browser dependencies required. Just httpx + pydantic (~2MB vs ~500MB for the full crawl4ai package).

Install

pip install crawl4ai-client

Quick Start

import asyncio
from crawl4ai_client import Crawl4aiDockerClient

async def main():
    async with Crawl4aiDockerClient(
        base_url="http://localhost:11235",
        api_token="your-token",  # optional
    ) as client:
        result = await client.crawl(["https://example.com"])
        print(result.raw_markdown)

asyncio.run(main())

Features

  • Crawl single or multiple URLs (/crawl)
  • Stream results as they complete (/crawl/stream)
  • Markdown extraction with filters (/md)
  • Screenshots as base64 PNG (/screenshot)
  • PDF generation (/pdf)
  • HTML preprocessing for schema extraction (/html)
  • JavaScript execution on pages (/execute_js)
  • LLM Q&A — ask questions about page content (/llm)
  • Per-URL configs for batch crawling (crawler_configs list)
  • Schema retrieval (/schema)
  • Async context manager with automatic cleanup

Usage

Basic crawl

from crawl4ai_client import Crawl4aiDockerClient, CrawlerRunConfig, CacheMode

async with Crawl4aiDockerClient(base_url="http://localhost:11235") as client:
    result = await client.crawl(
        ["https://example.com"],
        crawler_config=CrawlerRunConfig(cache_mode=CacheMode.BYPASS),
    )
    print(result.raw_markdown)

Multiple URLs with per-URL configs

from crawl4ai_client import Crawl4aiDockerClient, CrawlerRunConfig

async with Crawl4aiDockerClient(base_url="http://localhost:11235") as client:
    results = await client.crawl(
        ["https://example.com", "https://httpbin.org/html"],
        crawler_configs=[
            CrawlerRunConfig(word_count_threshold=5),
            CrawlerRunConfig(word_count_threshold=50),
        ],
    )
    for r in results:
        print(f"{r.url}: {len(r.raw_markdown)} chars")

Deep crawl

from crawl4ai_client import Crawl4aiDockerClient, CrawlerRunConfig, BFSDeepCrawlStrategy

async with Crawl4aiDockerClient(base_url="http://localhost:11235") as client:
    results = await client.crawl(
        ["https://example.com"],
        crawler_config=CrawlerRunConfig(
            deep_crawl_strategy=BFSDeepCrawlStrategy(max_depth=2, max_pages=10),
        ),
    )
    for r in results:
        print(f"{r.url}: {r.success}")

Also available: DFSDeepCrawlStrategy, BestFirstCrawlingStrategy.

Streaming

async with Crawl4aiDockerClient(base_url="http://localhost:11235") as client:
    async for result in client.crawl_stream(["https://example.com", "https://httpbin.org/html"]):
        print(f"Got: {result.url}")

Markdown endpoint

async with Crawl4aiDockerClient(base_url="http://localhost:11235") as client:
    md = await client.get_markdown("https://example.com", content_filter="fit")
    print(md)

Screenshot

async with Crawl4aiDockerClient(base_url="http://localhost:11235") as client:
    screenshot_b64 = await client.screenshot("https://example.com")

PDF generation

async with Crawl4aiDockerClient(base_url="http://localhost:11235") as client:
    pdf_b64 = await client.get_pdf("https://example.com")

HTML preprocessing

async with Crawl4aiDockerClient(base_url="http://localhost:11235") as client:
    html = await client.get_html("https://example.com")

JavaScript execution

async with Crawl4aiDockerClient(base_url="http://localhost:11235") as client:
    result = await client.execute_js(
        "https://example.com",
        scripts=["document.title", "document.querySelectorAll('a').length"],
    )
    print(result.js_execution_result)

LLM Q&A

async with Crawl4aiDockerClient(base_url="http://localhost:11235") as client:
    answer = await client.llm_query(
        "https://example.com",
        query="What is this page about?",
    )
    print(answer)

Hooks (custom page interaction)

async def block_images(page, context, **kwargs):
    await context.route("**/*.{png,jpg,jpeg,gif}", lambda route: route.abort())
    return page

async def scroll_page(page, context, **kwargs):
    await page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
    await page.wait_for_timeout(2000)
    return page

async with Crawl4aiDockerClient(base_url="http://localhost:11235") as client:
    result = await client.crawl(
        ["https://example.com"],
        hooks={
            "on_page_context_created": block_images,
            "before_retrieve_html": scroll_page,
        },
        hooks_timeout=30,
    )

Hooks can also be passed as pre-stringified source code:

hooks={"on_page_context_created": 'async def hook(page, context, **kwargs):\n    return page'}

Why this package?

The full crawl4ai package installs 34+ dependencies (~500MB) including Playwright, browsers, numpy, and litellm. If you're running Crawl4AI as a Docker service and only need the client, this package gives you the same Crawl4aiDockerClient with just 2 dependencies.

Compatibility

This client is compatible with Crawl4AI Docker server v0.8.x+. The config classes (BrowserConfig, CrawlerRunConfig) produce the same serialized format as the full library.

License

Apache 2.0 — based on crawl4ai by unclecode.

Release files for crawl4ai-client 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for crawl4ai-client 0.2.0
File Size Uploaded
crawl4ai_client-0.2.0.tar.gz 12.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for crawl4ai-client 0.2.0
File Interpreter ABI Platform
crawl4ai_client-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 25.1 kB

Release files / crawl4ai_client-0.2.0.tar.gz

Download URL crawl4ai_client-0.2.0.tar.gz
Size 12.7 kB
Tags Source
SHA-256 checksum
How to use checksums
ff8161148c6f366d6fa2c13fa8b12a3614f96957b542c02df51b868b27742ace
BLAKE2b-256 checksum
How to use checksums
2741dcf0d410a6401c69a35796f4ca3297648763b456ffdd428fc56470010001
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.10.8

Release files / crawl4ai_client-0.2.0-py3-none-any.whl

Download URL crawl4ai_client-0.2.0-py3-none-any.whl
Size 12.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
cea2fa60bef89fbafac3b7472db18ae26010974fbbfd0f482020ed3a31772fb3
BLAKE2b-256 checksum
How to use checksums
954b49b928b9dad908b65c3ed5e74536b7e6704cd8bafaaa39113befc4fb80e5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.10.8

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page