Skip to main content

🔥 PyreCrawl — Web Browsing Superpowers for Your AI Agent

License: MIT MCP Python 3.10+ PyPI

One command gives any AI agent the whole web. Scrape, extract, crawl, map, and search — self-hosted, no API keys, no rate limits, no subscription.

PyreCrawl speaks MCP (Model Context Protocol), the standard tool interface for Claude, Cursor, VS Code, Codex, OpenCode, Hermes, and any MCP-compatible agent.

A smart auto-fallback ladder always picks the cheapest method that succeeds:

fast HTTP
    │  (403/503/Cloudflare challenge or empty body)
    ▼
stealth browser (real Chromium + Cloudflare solver)
    │  (still blocked, or the page needs full JS rendering)
    ▼
deep processing (LLM-ready markdown, citations, structured extraction)

⚡ Tools exposed

Tool What it does
scrape(url, prefer="auto") Single URL → LLM-ready markdown
extract(url, schema) Scrape + structured extraction (JsonCss schema)
map_site(root, include_pattern=None, limit=200) Enumerate all internal URLs
crawl(root, max_pages=5, prefer="auto", include_paths=None, exclude_paths=None, max_depth=0) Multi-page crawl with path filters + true BFS depth
document(url) PDF/DOCX/PPTX → markdown (no browser, optional [docs] extras)
search(query, limit=10) Web search via DuckDuckGo HTML (no API key)
batch_scrape(urls[], ...) Many URLs in ONE call — parallel, deduped, cache-aware
deep_research(query, limit=5, scrape_top=3) Search → evidence pack with [n] citations (no LLM synthesis — your agent does that)
monitor(url, action, css_selector=None) Change detection with persisted snapshots + unified diff
session(session, action, ...) Persistent browser session (cookies kept) — login walls, multi-step flows, screenshots
cache(action) Inspect/clear/enable/disable the HTTP response cache
health() Versions + import sanity check

MCP Resources (read-only state without a tool call): pyrecrawl://cache/stats · pyrecrawl://sessions · pyrecrawl://monitors

MCP Prompts (ready-made playbooks): research(topic) · rag_ingest(site) · watch_page(url)

Env flags

Variable Default Effect
PYRECRAWL_CACHE off 1 = in-memory LRU (128 pages), or a directory path (reserved for disk mode)
PYRECRAWL_CACHE_TTL 900 Cache entry lifetime in seconds
PYRECRAWL_MONITOR_DIR ~/.pyrecrawl/monitors Where monitor snapshots persist

prefer options: "auto" (default ladder) · "fast" (HTTP only) · "stealth" (CF bypass) · "llm" (deep processing).


🚀 Install & Use (one-liner)

1. Install

# Using uv (recommended — fast, isolated, no venv needed)
uv tool install pyrecrawl

# Or pipx (alternative)
pipx install pyrecrawl

# Or pip into a venv
pip install pyrecrawl

2. One-time browser engines

pyrecrawl setup

This installs Chromium + stealth browser engines (~2 min, one-time).

3. Register with your AI agent

# Auto-detect installed agents and write their MCP configs
pyrecrawl install

# Or target specific agents
pyrecrawl install claude-desktop cursor

# Dry-run to preview what would change
pyrecrawl install --dry-run

Supported agents: claude-desktop, claude-code, cursor, vscode, codex, opencode, hermes.

4. Start chatting

After installing + registering, restart your agent (or start a new session). Then ask:

"Scrape https://example.com and summarize it."

The tools appear as mcp_pyrecrawl_scrape, mcp_pyrecrawl_extract, mcp_pyrecrawl_map_site, mcp_pyrecrawl_crawl, mcp_pyrecrawl_search, mcp_pyrecrawl_health.


📚 Manual config (if pyrecrawl install doesn't match your setup)

Claude Desktop

Config file

  • Linux: ~/.config/Claude/claude_desktop_config.json
  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
  • Windows: %AppData%\Claude\claude_desktop_config.json
{
  "mcpServers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
    }
  }
}

Claude Code

Config file: project-scoped .mcp.json

{
  "mcpServers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
    }
  }
}

Cursor

Config file: ~/.cursor/mcp.json

{
  "mcpServers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
    }
  }
}

VS Code / Copilot

Config file: .vscode/mcp.json (project-scoped)

{
  "servers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"],
      "type": "stdio"
    }
  }
}

Codex CLI

Config file: ~/.codex/config.toml

[mcp_servers.pyrecrawl]
command = "uvx"
args = ["--from", "pyrecrawl", "pyrecrawl", "serve"]

OpenCode

Config file: ~/.config/opencode/opencode.json

{
  "mcp": {
    "pyrecrawl": {
      "type": "local",
      "command": ["uvx", "--from", "pyrecrawl", "pyrecrawl", "serve"],
      "enabled": true
    }
  }
}

Hermes

Config file

  • Linux/macOS: ~/.hermes/config.yaml
  • Windows: %LocalAppData%\hermes\config.yaml
mcp_servers:
  pyrecrawl:
    command: uvx
    args:
      - --from
      - pyrecrawl
      - pyrecrawl
      - serve
    enabled: true

Windows note: uvx must be on PATH. If not, use the full path to uvx.exe (e.g. C:\Users\<you>\AppData\Local\hermes\bin\uvx.exe).


🧠 How the ladder chooses

PyreCrawl runs each request through three tiers, stopping at the first one that returns a complete, LLM-ready result:

Concern Fast tier Stealth tier Deep tier
Static HTML page ✅ ~200ms — —
Cloudflare-protected ❌ ✅ Turnstile solver —
JS-heavy SPA ❌ ✅ real Chromium —
Live DOM data (input .value, JS state) ❌ ✅ js param —
LLM-ready markdown + citations — — ✅ BM25, fit-markdown
Structured extraction (CSS schema) — — ✅
Deep crawl (BFS/DFS/BestFirst) — — ✅ adaptive

The agent never has to pick. prefer="auto" does it every call.

Live DOM data with js and wait_for

Some sites keep the data you want in a DOM property (e.g. an <input>'s .value) that JS writes after an XHR — it never appears in the serialized HTML. The scrape tool accepts two stealth-tier params for exactly this:

{
  "url": "https://temp-mail.org/id",
  "prefer": "stealth",
  "wait_for": "document.getElementById('mail').value.includes('@')",
  "js": "document.getElementById('mail').value"
}
  • wait_for — a JS predicate expression polled until truthy (bounded by timeout). Use it instead of guessing a sleep for anything that arrives asynchronously.
  • js — a JS expression evaluated once the page settles; the value comes back in meta.js_result. Errors are captured in meta.js_error (the page result is still returned, never a crash).

📊 Compared to Firecrawl (hosted)

Firecrawl PyreCrawl
Cost Free 1k/mo, then $16–333/mo Free, self-hosted
Local LLM support ❌ ✅ Ollama / any LLM
Cloudflare bypass ✅ (Fire-Engine, paid) ✅ (free, built-in)
Markdown + BM25 ✅ ✅
Self-host ❌ ✅
Hosted search API ✅ /search ⚠️ DuckDuckGo HTML (no key)

🔧 Development

git clone https://github.com/SanggonBoy/PyreCrawl.git
cd PyreCrawl
uv venv --python 3.12 .venv
source .venv/Scripts/activate  # Windows; or .venv/bin/activate on macOS/Linux
uv pip install -e ".[dev]"
python -m playwright install chromium
scrapling install

Run tests

python scripts/selfcheck.py   # real-network smoke test
python scripts/probe_stdio.py # stdio JSON-RPC probe

📦 Publish

Maintainers only:

git tag v0.2.1
git push origin v0.2.1

GitHub Actions builds + uploads to PyPI via trusted publishing.


📜 Uninstall

# Remove from all agent configs
pyrecrawl uninstall

# Remove the package
uv tool uninstall pyrecrawl

🛡️ License

MIT — see LICENSE.

🙏 Credits

Built on the shoulders of Scrapling and Crawl4AI — both MIT, both excellent.

Metadata

Release files for pyrecrawl 0.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pyrecrawl 0.6.0
File Size Uploaded
pyrecrawl-0.6.0.tar.gz 38.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pyrecrawl 0.6.0
File Interpreter ABI Platform
pyrecrawl-0.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 73.8 kB

Release files / pyrecrawl-0.6.0.tar.gz

Download URL pyrecrawl-0.6.0.tar.gz
Size 38.9 kB
Tags Source
SHA-256 checksum
How to use checksums
c541d689faee895d3f745d4fe750192ffc2085f318da795d9cc79faa0c37a303
BLAKE2b-256 checksum
How to use checksums
15ea70aa5d825337cfb18543c8fc0e5ad69ecdbdd33ad479086c14c33282179f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.7 {"installer":{"name":"uv","version":"0.12.7","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / pyrecrawl-0.6.0-py3-none-any.whl

Download URL pyrecrawl-0.6.0-py3-none-any.whl
Size 34.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0aaffbb2cc0fed4bc9d30330f427a74d7aced43749d4ada31452a43f4185c94e
BLAKE2b-256 checksum
How to use checksums
5ebd5dbaf49a53ac1423df58fbc205ac85a0c79a2ab495f1eb1d1fa6362448a4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.7 {"installer":{"name":"uv","version":"0.12.7","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

0.8.1

2 release files

0.8.0

2 release files

0.7.5

2 release files

0.7.4

2 release files

0.7.3

2 release files

0.7.2

2 release files

0.7.1

2 release files

0.7.0

2 release files

This release

0.6.0 This release

2 release files

0.2.2

2 release files

0.2.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page