Skip to main content

Crawlemoon MCP Server

Crawlemoon MCP Server — free, AI-native web crawling for the agent era

python 3.10+ · pypi 1.1.0 · MIT · MCP-native · code style black

A free, open-source MCP server that gives any agent (Claude Code, Cursor, Windsurf, …) 55 production-grade tools for the full web-crawling stack: deep analysis, stealth, API discovery, session recording → runnable crawler, smart extraction. No proprietary API. No per-request fee.

Crawlemoon capabilities — deep analysis, stealth, record→crawler, smart extraction


Quick start

Three install paths — uvx, pipx, pip

The recommended path needs no install — uvx runs straight from PyPI:

{
  "mcpServers": {
    "crawlemoon": {
      "command": "uvx",
      "args": ["crawlemoon"]
    }
  }
}

Requires uv. Install once: curl -LsSf https://astral.sh/uv/install.sh | sh. Or use pipx run crawlemoon / pip install crawlemoon instead.

Where to put that JSON: Cursor → Settings → MCP. Claude Code → ~/.config/claude/mcp_settings.json. Windsurf → Settings → MCP Servers.


How it works

Agent → Crawlemoon → Browser/HTTP/Proxy → target web

Your agent talks to Crawlemoon over the Model Context Protocol. Crawlemoon owns a hardened browser pool, an HTTP stack with TLS fingerprinting, and a rotating proxy pool. While it fetches pages, it captures network traffic, reads scripts, and introspects schemas — so the agent gets clean structured data, not raw HTML.


What's in the box

A short list — see the source for the full set of 55 tools.

Group Tools
Deep analysis deep_analyze, discover_apis, introspect_graphql, analyze_websocket, analyze_auth, detect_protection, detect_technology
Stealth stealth_request, configure_proxies, configure_rate_limit, add_proxy, test_proxy
Record → crawler record_session, stop_recording, export_recording, generate_crawler
Extraction smart_extract, extract_article, extract_tables, extract_links, extract_forms, extract_metadata, convert_to_markdown
Page interaction take_screenshot, fill_form, wait_and_extract, compare_pages, measure_performance, check_accessibility, get_dom_tree
Sessions & cache save_session, load_session, get_cookies, get_storage, clear_cache, get_cache_stats
Advanced (opt-in) execute_js, execute_cdp, deobfuscate_js, extract_from_js, solve_captcha

Smart extraction — bring any LLM, including free ones

smart_extract works without any API key using pattern matching. Plug in any OpenAI-compatible endpoint for higher accuracy — including FREE tiers:

# OpenRouter (free models exist)
CRAWLEMOON_LLM_PROVIDER=openrouter
CRAWLEMOON_LLM_API_KEY=sk-or-v1-xxx
CRAWLEMOON_LLM_MODEL=meta-llama/llama-3.2-3b-instruct:free

# Groq (free, very fast)
CRAWLEMOON_LLM_PROVIDER=groq
CRAWLEMOON_LLM_API_KEY=gsk_xxx

# Local Ollama (no key needed)
CRAWLEMOON_LLM_PROVIDER=ollama
CRAWLEMOON_LLM_MODEL=llama3.2

Together, DeepSeek, Mistral, Fireworks, and standard OpenAI also work via CRAWLEMOON_LLM_BASE_URL.


Configuration

Variable Default Notes
CRAWLEMOON_HEADLESS true Run browser without UI
CRAWLEMOON_BROWSER chromium chromium / firefox / webkit
CRAWLEMOON_POOL_SIZE 5 Max concurrent browsers
CRAWLEMOON_NAV_TIMEOUT 30.0 Page-load timeout (s)
CRAWLEMOON_API_KEY unset If set, every tool call must include matching _api_key
CRAWLEMOON_ALLOW_DANGEROUS_JS false Required for execute_js / execute_cdp / deobfuscate_js
CRAWLEMOON_JS_MAX_LENGTH 50000 Length cap for JS payloads
CRAWLEMOON_JS_EXEC_TIMEOUT 10.0 Per-script timeout (s)
CRAWLEMOON_PROXIES unset Comma/newline separated proxy entries
CRAWLEMOON_PROXIES_FILE unset Local file with one proxy per line
CRAWLEMOON_PROXY_SCHEME http Scheme for entries without one: http, https, socks4, socks5
CRAWLEMOON_PROXY_ROTATION round_robin round_robin, random, sticky, least_used
CRAWLEMOON_PROXY_HEALTH_CHECK_INTERVAL 300 Proxy health-check interval (s)
CRAWLEMOON_PROXY_FAIL_CLOSED false Raise on startup proxy config errors instead of continuing direct

Proxy, V2ray & connection control

Crawlemoon uses one rotating proxy pool for browser contexts, stealth_request, and Xray/V2ray local exits. Credentials are stored separately from normalized URLs and are masked in stats/log-style responses.

Supported proxy entry formats:

http://user:pass@31.59.20.176:6754
socks5://user:pass@127.0.0.1:1080
31.59.20.176:6754:user:pass
31.59.20.176:6754

Start the MCP server with a proxy file:

{
  "mcpServers": {
    "crawlemoon": {
      "command": "uvx",
      "args": ["crawlemoon"],
      "env": {
        "CRAWLEMOON_PROXIES_FILE": "/secure/path/proxies.txt",
        "CRAWLEMOON_PROXY_SCHEME": "http",
        "CRAWLEMOON_PROXY_ROTATION": "sticky"
      }
    }
  }
}

Proxy files can use Webshare-style lines and comments:

# host:port:username:password
31.59.20.176:6754:uusmdewb:en3w097syrxh
31.56.127.193:7684:uusmdewb:en3w097syrxh

Configure or replace proxies at runtime with the configure_proxies MCP tool:

{
  "proxies_text": "31.59.20.176:6754:user:pass\n31.56.127.193:7684:user:pass",
  "default_scheme": "http",
  "rotation_strategy": "sticky",
  "health_check_interval": 300,
  "replace_existing": true
}

Use add_proxy, remove_proxy, test_proxy, and get_proxy_stats for incremental control. get_proxy_stats returns masked proxy URLs, for example http://user:***@31.59.20.176:6754.

For V2ray/Xray, call configure_xray_subscription with either subscription_url or raw_links. Crawlemoon starts local SOCKS5 exits such as socks5://127.0.0.1:10801, registers them in the same proxy pool, and can rotate or benchmark nodes with rotate_xray_node and test_xray_nodes.

Keep proxy files out of git. They contain live credentials.


Security

execute_js, execute_cdp, and deobfuscate_js are disabled by default — they execute or operate on arbitrary code in a real browser. Enable on trusted networks with CRAWLEMOON_ALLOW_DANGEROUS_JS=true. Even then, payloads are length-capped, time-bounded, and a denylist rejects eval, new Function, dynamic import(), document.write, importScripts, and WebAssembly.{compile,instantiate}. Set CRAWLEMOON_API_KEY so MCP clients must present a matching _api_key.

These are mitigations, not a sandbox: do not expose this server to untrusted clients.


Develop

git clone https://github.com/razavioo/crawlemoon.git
cd crawlemoon
make dev-install      # editable install + dev/captcha/ocr extras + pre-commit
make test             # pytest
make lint             # ruff + mypy

Releases

This project uses Trusted Publishing (OIDC) via GitHub Actions to automate publishing releases directly to PyPI.

To release a new version:

  1. Bump the version number in pyproject.toml.
  2. Commit the change and create a git tag matching the version (e.g. v1.1.8):
    git add pyproject.toml
    git commit -m "chore: bump version to 1.1.8"
    git tag v1.1.8
    
  3. Push your branch and the tag to GitHub:
    git push origin main --tags
    

GitHub Actions will automatically run tests, build the package, and publish it securely to PyPI under the crawlemoon package space.

PRs welcome. Particularly interested in: distributed mode (Redis queue), result sinks (Postgres / S3), Prometheus metrics. See MIT License.

Made by emad.dev

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

crawlemoon-1.1.9.tar.gz (170.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

crawlemoon-1.1.9-py3-none-any.whl (141.1 kB view details)

Uploaded Python 3

File details

Details for the file crawlemoon-1.1.9.tar.gz.

File metadata

  • Download URL: crawlemoon-1.1.9.tar.gz
  • Upload date:
  • Size: 170.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for crawlemoon-1.1.9.tar.gz
Algorithm Hash digest
SHA256 7d4cb158c60eaef676efe3e5902393e2d174acefd352c5d079efeb15a385d1f2
MD5 0f647bdc479f5e6998cb4ec58a09199c
BLAKE2b-256 20497a2304a46139606b4e8d8918834f610c8bfa89344bcb45ada4e074f26c5f

See more details on using hashes here.

Provenance

The following attestation bundles were made for crawlemoon-1.1.9.tar.gz:

Publisher: release.yml on razavioo/crawlemoon

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file crawlemoon-1.1.9-py3-none-any.whl.

File metadata

  • Download URL: crawlemoon-1.1.9-py3-none-any.whl
  • Upload date:
  • Size: 141.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for crawlemoon-1.1.9-py3-none-any.whl
Algorithm Hash digest
SHA256 5faa4470c94d14a2424ea676ea3b3619c7ab728a7312d76b2888c8c4e9d22df7
MD5 5221f7118bd524d9a4a5abe8f23cae44
BLAKE2b-256 8d0a66176419d55763acfd495c05031136cc00718cb42b401d06773841e9b8bc

See more details on using hashes here.

Provenance

The following attestation bundles were made for crawlemoon-1.1.9-py3-none-any.whl:

Publisher: release.yml on razavioo/crawlemoon

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

1.1.11

2 files

1.1.10

2 files

This release

1.1.9 This release

2 files

1.1.8

2 files

1.1.7

2 files

1.1.6

2 files

1.1.5

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page