Skip to main content

🔥 PyreCrawl — Web Browsing Superpowers for Your AI Agent

License: MIT MCP Python 3.10+ PyPI

One command gives any AI agent the whole web. Scrape, extract, crawl, map, and search — self-hosted, no API keys, no rate limits, no subscription.

PyreCrawl speaks MCP (Model Context Protocol), the standard tool interface for Claude, Cursor, VS Code, Codex, OpenCode, Hermes, and any MCP-compatible agent.

A smart auto-fallback ladder always picks the cheapest method that succeeds:

fast HTTP
    │  (403/503/Cloudflare challenge or empty body)
    ▼
stealth browser (real Chromium + Cloudflare solver)
    │  (still blocked, or the page needs full JS rendering)
    ▼
deep processing (LLM-ready markdown, citations, structured extraction)

⚡ Tools exposed

Tool What it does
scrape(url, prefer="auto") Single URL → LLM-ready markdown
extract(url, schema) Scrape + structured extraction (JsonCss schema)
map_site(root, include_pattern=None, limit=200) Enumerate all internal URLs
crawl(root, max_pages=5, prefer="auto", include_paths=None, exclude_paths=None, max_depth=0) Multi-page crawl with path filters + true BFS depth
document(url) PDF/DOCX/PPTX → markdown (no browser, optional [docs] extras)
search(query, limit=10) Web search via DuckDuckGo HTML (no API key)
search_papers(query, limit=8, source="arxiv", category=None) Academic search via arXiv + Crossref (no API key) — feed pdf_url into document
batch_scrape(urls[], ...) Many URLs in ONE call — parallel, deduped, cache-aware
deep_research(query, limit=5, scrape_top=3) Search → evidence pack with [n] citations (no LLM synthesis — your agent does that)
monitor(url, action, css_selector=None) Change detection with persisted snapshots + unified diff
session(session, action, ...) Persistent browser session (cookies kept) — login walls, multi-step flows, screenshots
cache(action) Inspect/clear/enable/disable the HTTP response cache
health() Versions + import sanity check

MCP Resources (read-only state without a tool call): pyrecrawl://cache/stats · pyrecrawl://sessions · pyrecrawl://monitors

MCP Prompts (ready-made playbooks): research(topic) · rag_ingest(site) · watch_page(url)

Env flags

Variable Default Effect
PYRECRAWL_CACHE off 1 = in-memory LRU (128 pages), or a directory path (reserved for disk mode)
PYRECRAWL_CACHE_TTL 900 Cache entry lifetime in seconds
PYRECRAWL_MONITOR_DIR ~/.pyrecrawl/monitors Where monitor snapshots persist

prefer options: "auto" (default ladder) · "fast" (HTTP only) · "stealth" (CF bypass) · "llm" (deep processing).


🚀 Install & Use (one-liner)

1. Install

# Using uv (recommended — fast, isolated, no venv needed)
uv tool install pyrecrawl

# Or pipx (alternative)
pipx install pyrecrawl

# Or pip into a venv
pip install pyrecrawl

2. One-time browser engines

pyrecrawl setup

This installs Chromium + stealth browser engines (~2 min, one-time).

3. Register with your AI agent

# Auto-detect installed agents and write their MCP configs
pyrecrawl install

# Or target specific agents
pyrecrawl install claude-desktop cursor

# Dry-run to preview what would change
pyrecrawl install --dry-run

Supported agents: claude-desktop, claude-code, cursor, vscode, codex, opencode, hermes.

4. Start chatting

After installing + registering, restart your agent (or start a new session). Then ask:

"Scrape https://example.com and summarize it."

The tools appear as mcp_pyrecrawl_scrape, mcp_pyrecrawl_extract, mcp_pyrecrawl_map_site, mcp_pyrecrawl_crawl, mcp_pyrecrawl_search, mcp_pyrecrawl_health.


📚 Manual config (if pyrecrawl install doesn't match your setup)

Claude Desktop

Config file

  • Linux: ~/.config/Claude/claude_desktop_config.json
  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
  • Windows: %AppData%\Claude\claude_desktop_config.json
{
  "mcpServers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
    }
  }
}

Claude Code

Config file: project-scoped .mcp.json

{
  "mcpServers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
    }
  }
}

Cursor

Config file: ~/.cursor/mcp.json

{
  "mcpServers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
    }
  }
}

VS Code / Copilot

Config file: .vscode/mcp.json (project-scoped)

{
  "servers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"],
      "type": "stdio"
    }
  }
}

Codex CLI

Config file: ~/.codex/config.toml

[mcp_servers.pyrecrawl]
command = "uvx"
args = ["--from", "pyrecrawl", "pyrecrawl", "serve"]

OpenCode

Config file: ~/.config/opencode/opencode.json

{
  "mcp": {
    "pyrecrawl": {
      "type": "local",
      "command": ["uvx", "--from", "pyrecrawl", "pyrecrawl", "serve"],
      "enabled": true
    }
  }
}

Hermes

Config file

  • Linux/macOS: ~/.hermes/config.yaml
  • Windows: %LocalAppData%\hermes\config.yaml
mcp_servers:
  pyrecrawl:
    command: uvx
    args:
      - --from
      - pyrecrawl
      - pyrecrawl
      - serve
    enabled: true

Windows note: uvx must be on PATH. If not, use the full path to uvx.exe (e.g. C:\Users\<you>\AppData\Local\hermes\bin\uvx.exe).


🧠 How the ladder chooses

PyreCrawl runs each request through three tiers, stopping at the first one that returns a complete, LLM-ready result:

Concern Fast tier Stealth tier Deep tier
Static HTML page ✅ ~200ms — —
Cloudflare-protected ❌ ✅ Turnstile solver —
JS-heavy SPA ❌ ✅ real Chromium —
Live DOM data (input .value, JS state) ❌ ✅ js param —
LLM-ready markdown + citations — — ✅ BM25, fit-markdown
Structured extraction (CSS schema) — — ✅
Deep crawl (BFS/DFS/BestFirst) — — ✅ adaptive

The agent never has to pick. prefer="auto" does it every call.

Live DOM data with js and wait_for

Some sites keep the data you want in a DOM property (e.g. an <input>'s .value) that JS writes after an XHR — it never appears in the serialized HTML. The scrape tool accepts two stealth-tier params for exactly this:

{
  "url": "https://temp-mail.org/id",
  "prefer": "stealth",
  "wait_for": "document.getElementById('mail').value.includes('@')",
  "js": "document.getElementById('mail').value"
}
  • wait_for — a JS predicate expression polled until truthy (bounded by timeout). Use it instead of guessing a sleep for anything that arrives asynchronously.
  • js — a JS expression evaluated once the page settles; the value comes back in meta.js_result. Errors are captured in meta.js_error (the page result is still returned, never a crash).

📊 Compared to Firecrawl (hosted)

Firecrawl PyreCrawl
Cost Free 1k/mo, then $16–333/mo Free, self-hosted
Local LLM support ❌ ✅ Ollama / any LLM
Cloudflare bypass ✅ (Fire-Engine, paid) ✅ (free, built-in)
Markdown + BM25 ✅ ✅
Self-host ❌ ✅
Academic paper search ❌ ✅ arXiv + Crossref (search_papers)
Hosted search API ✅ /search ⚠️ DuckDuckGo HTML + arXiv/Crossref (no key)

🔧 Development

git clone https://github.com/SanggonBoy/PyreCrawl.git
cd PyreCrawl
uv venv --python 3.12 .venv
source .venv/Scripts/activate  # Windows; or .venv/bin/activate on macOS/Linux
uv pip install -e ".[dev]"
python -m playwright install chromium
scrapling install

Run tests

python scripts/selfcheck.py   # real-network smoke test
python scripts/probe_stdio.py # stdio JSON-RPC probe

📦 Publish

Maintainers only:

git tag v0.8.0
git push origin v0.8.0

GitHub Actions builds + uploads to PyPI via trusted publishing.


📜 Uninstall

# Remove from all agent configs
pyrecrawl uninstall

# Remove the package
uv tool uninstall pyrecrawl

🛡️ License

MIT — see LICENSE.

🙏 Credits

Built on the shoulders of Scrapling and Crawl4AI — both MIT, both excellent.

Metadata

Release files for pyrecrawl 0.7.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pyrecrawl 0.7.1
File Size Uploaded
pyrecrawl-0.7.1.tar.gz 41.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pyrecrawl 0.7.1
File Interpreter ABI Platform
pyrecrawl-0.7.1-py3-none-any.whl Python 3 none any Details

Total release size: 78.1 kB

Release files / pyrecrawl-0.7.1.tar.gz

Download URL pyrecrawl-0.7.1.tar.gz
Size 41.1 kB
Tags Source
SHA-256 checksum
How to use checksums
8a2013e6e7638efebacac1248cdc6c6618bb4269f956c6bfb4d873e4a92f9aaf
BLAKE2b-256 checksum
How to use checksums
365a1d9647a246f1180a17817b96712719f14838ecbab7c66b9d0ac4c2f8c4da
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.7 {"installer":{"name":"uv","version":"0.12.7","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / pyrecrawl-0.7.1-py3-none-any.whl

Download URL pyrecrawl-0.7.1-py3-none-any.whl
Size 37.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2b4c2af07b5cc3635d841a6c30942276fbfbf51b1a269baee040759d5f7c3ce9
BLAKE2b-256 checksum
How to use checksums
8933889d424a8b955906b30e80c56b3d5a4c0cc3bb8c9740d7815e2d6b27f8b4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.7 {"installer":{"name":"uv","version":"0.12.7","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

0.8.1

2 release files

0.8.0

2 release files

0.7.5

2 release files

0.7.4

2 release files

0.7.3

2 release files

0.7.2

2 release files

This release

0.7.1 This release

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.2.2

2 release files

0.2.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page