Skip to main content

🔥 PyreCrawl — Web Browsing Superpowers for Your AI Agent

License: MIT MCP Python 3.10+ PyPI

One command gives any AI agent the whole web. Scrape, extract, crawl, map, and search — self-hosted, no API keys, no rate limits, no subscription.

PyreCrawl speaks MCP (Model Context Protocol), the standard tool interface for Claude, Cursor, VS Code, Codex, OpenCode, Hermes, and any MCP-compatible agent.

A smart auto-fallback ladder always picks the cheapest method that succeeds:

fast HTTP
    │  (403/503/Cloudflare challenge or empty body)
    ▼
stealth browser (real Chromium + Cloudflare solver)
    │  (still blocked, or the page needs full JS rendering)
    ▼
deep processing (LLM-ready markdown, citations, structured extraction)

⚡ Tools exposed

Tool What it does
scrape(url, prefer="auto") Single URL → LLM-ready markdown
extract(url, schema) Scrape + structured extraction (JsonCss schema)
map_site(root, include_pattern=None, limit=200) Enumerate all internal URLs
crawl(root, max_pages=5, prefer="auto", include_paths=None, exclude_paths=None, max_depth=0) Multi-page crawl with path filters + true BFS depth
document(url) PDF/DOCX/PPTX → markdown (no browser, optional [docs] extras)
search(query, limit=10) Web search via DuckDuckGo HTML (no API key)
search_papers(query, limit=8, source="arxiv", category=None) Academic search via arXiv + Crossref (no API key) — feed pdf_url into document
batch_scrape(urls[], ...) Many URLs in ONE call — parallel, deduped, cache-aware
deep_research(query, limit=5, scrape_top=3) Search → evidence pack with [n] citations (no LLM synthesis — your agent does that)
monitor(url, action, css_selector=None) Change detection with persisted snapshots + unified diff
session(session, action, ...) Persistent browser session (cookies kept) — login walls, multi-step flows, screenshots
cache(action) Inspect/clear/enable/disable the HTTP response cache
health() Versions + import sanity check

MCP Resources (read-only state without a tool call): pyrecrawl://cache/stats · pyrecrawl://sessions · pyrecrawl://monitors

MCP Prompts (ready-made playbooks): research(topic) · rag_ingest(site) · watch_page(url)

Env flags

Variable Default Effect
PYRECRAWL_CACHE off 1 = in-memory LRU (128 pages), or a directory path (reserved for disk mode)
PYRECRAWL_CACHE_TTL 900 Cache entry lifetime in seconds
PYRECRAWL_MONITOR_DIR ~/.pyrecrawl/monitors Where monitor snapshots persist
PYRECRAWL_NO_TELEMETRY off 1 = disable the anonymous startup ping (also honors DO_NOT_TRACK=1)

prefer options: "auto" (default ladder) · "fast" (HTTP only) · "stealth" (CF bypass) · "llm" (deep processing).


🚀 Install & Use (one-liner)

1. Install

UV (recommended — one command, zero Python setup)

UV is a fast Python package manager that handles Python itself — no need to install Python separately. Get it once:

# macOS / Linux
curl -LsSf https://astral.sh/uv/install.sh | sh

# Windows (PowerShell)
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

Learn more about UV →

Then run PyreCrawl directly — no venv, no pip install, no Python download:

uvx pyrecrawl@latest

Or via uv tool install (persistent, recommended for regular use)

uv tool install pyrecrawl

Or via pipx (alternative)

pipx install pyrecrawl

Or via pip into a venv

pip install pyrecrawl

2. One-time browser engines

pyrecrawl setup

This installs Chromium + stealth browser engines (~2 min, one-time).

3. Register with your AI agent

# Auto-detect installed agents and write their MCP configs
pyrecrawl install

# Or target specific agents
pyrecrawl install claude-desktop cursor

# Dry-run to preview what would change
pyrecrawl install --dry-run

Supported agents: claude-desktop, claude-code, cursor, vscode, codex, opencode, hermes.

4. Start chatting

After installing + registering, restart your agent (or start a new session). Then ask:

"Scrape https://example.com and summarize it."

The tools appear as mcp_pyrecrawl_scrape, mcp_pyrecrawl_extract, mcp_pyrecrawl_map_site, mcp_pyrecrawl_crawl, mcp_pyrecrawl_search, mcp_pyrecrawl_health.


📚 Manual config (if pyrecrawl install doesn't match your setup)

Claude Desktop

Config file

  • Linux: ~/.config/Claude/claude_desktop_config.json
  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
  • Windows: %AppData%\Claude\claude_desktop_config.json
{
  "mcpServers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
    }
  }
}

Claude Code

Config file: project-scoped .mcp.json

{
  "mcpServers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
    }
  }
}

Cursor

Config file: ~/.cursor/mcp.json

{
  "mcpServers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
    }
  }
}

VS Code / Copilot

Config file: .vscode/mcp.json (project-scoped)

{
  "servers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"],
      "type": "stdio"
    }
  }
}

Codex CLI

Config file: ~/.codex/config.toml

[mcp_servers.pyrecrawl]
command = "uvx"
args = ["--from", "pyrecrawl", "pyrecrawl", "serve"]

OpenCode

Config file: ~/.config/opencode/opencode.json

{
  "mcp": {
    "pyrecrawl": {
      "type": "local",
      "command": ["uvx", "--from", "pyrecrawl", "pyrecrawl", "serve"],
      "enabled": true
    }
  }
}

Hermes

Config file

  • Linux/macOS: ~/.hermes/config.yaml
  • Windows: %LocalAppData%\hermes\config.yaml
mcp_servers:
  pyrecrawl:
    command: uvx
    args:
      - --from
      - pyrecrawl
      - pyrecrawl
      - serve
    enabled: true

Windows note: uvx must be on PATH. If not, use the full path to uvx.exe (e.g. C:\Users\<you>\AppData\Local\hermes\bin\uvx.exe).


🧠 How the ladder chooses

PyreCrawl runs each request through three tiers, stopping at the first one that returns a complete, LLM-ready result:

Concern Fast tier Stealth tier Deep tier
Static HTML page ✅ ~200ms — —
Cloudflare-protected ❌ ✅ Turnstile solver —
JS-heavy SPA ❌ ✅ real Chromium —
Live DOM data (input .value, JS state) ❌ ✅ js param —
LLM-ready markdown + citations — — ✅ BM25, fit-markdown
Structured extraction (CSS schema) — — ✅
Deep crawl (BFS/DFS/BestFirst) — — ✅ adaptive

The agent never has to pick. prefer="auto" does it every call.

Live DOM data with js and wait_for

Some sites keep the data you want in a DOM property (e.g. an <input>'s .value) that JS writes after an XHR — it never appears in the serialized HTML. The scrape tool accepts two stealth-tier params for exactly this:

{
  "url": "https://temp-mail.org/id",
  "prefer": "stealth",
  "wait_for": "document.getElementById('mail').value.includes('@')",
  "js": "document.getElementById('mail').value"
}
  • wait_for — a JS predicate expression polled until truthy (bounded by timeout). Use it instead of guessing a sleep for anything that arrives asynchronously.
  • js — a JS expression evaluated once the page settles; the value comes back in meta.js_result. Errors are captured in meta.js_error (the page result is still returned, never a crash).

📊 Compared to Firecrawl (hosted)

Firecrawl PyreCrawl
Cost Free 1k/mo, then $16–333/mo Free, self-hosted
Local LLM support ❌ ✅ Ollama / any LLM
Cloudflare bypass ✅ (Fire-Engine, paid) ✅ (free, built-in)
Markdown + BM25 ✅ ✅
Self-host ❌ ✅
Academic paper search ❌ ✅ arXiv + Crossref (search_papers)
Hosted search API ✅ /search ⚠️ DuckDuckGo HTML + arXiv/Crossref (no key)

🔧 Development

git clone https://github.com/SanggonBoy/PyreCrawl.git
cd PyreCrawl
uv venv --python 3.12 .venv
source .venv/Scripts/activate  # Windows; or .venv/bin/activate on macOS/Linux
uv pip install -e ".[dev]"
python -m playwright install chromium
scrapling install

Run tests

python scripts/selfcheck.py        # real-network smoke test (13 tools + engines)
python scripts/probe_stdio.py      # stdio JSON-RPC probe
python scripts/test_ladder_bug.py  # SPA-shell ladder escalation regression
python scripts/test_js_eval.py     # stealth js/wait_for params regression
python scripts/test_scope_selector.py  # crawl css_selector/max_depth wiring
python scripts/test_link_harvest.py    # map/BFS link purity regression

📦 Publish

Maintainers only:

git tag vX.Y.Z
git push origin vX.Y.Z

GitHub Actions builds + uploads to PyPI via trusted publishing.


🔔 Stay up to date

PyreCrawl checks PyPI on every startup and reports the latest version — your MCP agent sees this automatically via the health() tool response and can notify you inline.

To check manually:

pyrecrawl version

To upgrade:

pyrecrawl update   # runs: uv tool upgrade pyrecrawl

Get notified of new releases: click Watch → Releases only at the GitHub repo to receive email notifications when a new version is published.


[!NOTE] PyreCrawl sends one anonymous usage ping per 24 h at server startup — see Privacy for exactly what's sent and how to opt out.

🔒 Privacy — anonymous usage ping

PyreCrawl phones home once per 24 h with a tiny anonymous ping when the MCP server starts, so we can count real users (DAU/MAU) instead of raw downloads.

Sent (4 fields, ~100 bytes) Never sent
Hashed machine id (SHA-256 of hostname+MAC — not reversible) Your IP (not stored)
PyreCrawl version Any URL you scrape
Python version Any page content or search queries
OS family (windows / linux / darwin) Anything else

Client code: src/pyrecrawl/telemetry.py (~90 lines, stdlib only) · Collector: workers/telemetry/ — a self-hostable Cloudflare Worker + D1, no third-party analytics service.

Opt out any time:

export PYRECRAWL_NO_TELEMETRY=1   # or the industry-standard DO_NOT_TRACK=1

📜 Uninstall

# Remove from all agent configs
pyrecrawl uninstall

# Remove the package
uv tool uninstall pyrecrawl

🛡️ License

MIT — see LICENSE.

🙏 Credits

Built on the shoulders of Scrapling and Crawl4AI — both MIT, both excellent.

Metadata

Release files for pyrecrawl 0.7.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pyrecrawl 0.7.5
File Size Uploaded
pyrecrawl-0.7.5.tar.gz 372.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pyrecrawl 0.7.5
File Interpreter ABI Platform
pyrecrawl-0.7.5-py3-none-any.whl Python 3 none any Details

Total release size: 415.5 kB

Release files / pyrecrawl-0.7.5.tar.gz

Download URL pyrecrawl-0.7.5.tar.gz
Size 372.1 kB
Tags Source
SHA-256 checksum
How to use checksums
ce74bd0cded91988fd63e662d39f96e652b6e01dd340774f5ae9121980152823
BLAKE2b-256 checksum
How to use checksums
67fa4f6bbee924af3c674c8c2d678e06f4dba36914eeb13928fe1fa3d1509a76
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.7 {"installer":{"name":"uv","version":"0.12.7","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / pyrecrawl-0.7.5-py3-none-any.whl

Download URL pyrecrawl-0.7.5-py3-none-any.whl
Size 43.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
11b07d867cd50fc18b6f37bc6994e3970fe5369b8da238d3ed9897b7f2fcb5cd
BLAKE2b-256 checksum
How to use checksums
c8ec49bc04c5c199e849cb1b10a7272e49b70da4cd49eff78f0bd167d47d5081
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.7 {"installer":{"name":"uv","version":"0.12.7","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

0.8.1

2 release files

0.8.0

2 release files

This release

0.7.5 This release

2 release files

0.7.4

2 release files

0.7.3

2 release files

0.7.2

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.2.2

2 release files

0.2.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page