This release is a pre-release and may not be stable for production use.
crawl4tools
What is crawl4tools
crawl4tools is a web crawler built on top of the crawl4ai library. It is designed to serve three purposes:
- An HTTP server,
crawl4server, that works directly as an Open WebUI external web loader, returning the body of a URL as Markdown. On the Open WebUI side, this only requires setting Web Loader Engine toexternaland the External Web Loader URL in the admin UI (or the equivalentWEB_LOADER_ENGINE=externalandEXTERNAL_WEB_LOADER_URLenvironment variables). - An MCP server that lets AI agents such as Claude Code fetch the full content of a web page, rather than a summary.
- A local CLI that downloads a given URL (or multiple URLs at once) as Markdown or other formats.
Status
Beta. This release (1.0.0b1) is the first beta of crawl4tools, published on PyPI and Docker Hub. It provides the local CLI crawl4cli, the MCP server crawl4mcp, and the combined Open WebUI web loader + MCP server crawl4server, with a Dockerfile and compose file. All three commands can show their messages in English or Japanese. Feedback from real use is welcome on the Issues page; 1.0.0 will follow once this beta has been tested.
Features
Available now (CLI):
- Download one or more URLs as Markdown, HTML, PDF, screenshot (PNG), MHTML, or the raw source
- PDFs are transcribed to Markdown; images and other non-HTML files are saved as they are
- Clear error messages for HTTP errors, unknown hosts, refused connections, timeouts, and a missing browser
- Downloading through an HTTP/HTTPS/SOCKS5 proxy, with a single retry over a direct connection when the proxy itself fails or breaks the result — including when a proxy that intercepts TLS ("SSL bumping") shows a TLS error, an error page from the proxy itself, or a bot challenge seen through it; a refusal by the proxy itself is never bypassed
- Messages in English or Japanese (see Language of messages)
Available now (MCP server):
- Two tools:
fetchreturns pages directly to the client as Markdown (default), HTML, or a PNG screenshot (images come back as images, PDFs are transcribed to Markdown);downloadsaves pages in any format (Markdown, HTML, PDF, screenshot, MHTML, or raw) into a directory on the server and returns the paths - Several URLs per call (default limit 20), fetched concurrently under a server-wide limit (default 3), sharing one headless browser
- stdio (default) and Streamable HTTP transports
- Proxy support with a direct-connection fallback, configured on the server side, including for a proxy that intercepts TLS
- Settings via command-line options,
CRAWL4MCP_*environment variables, or a YAML/JSON config file - Messages in English or Japanese (see Language of messages)
Available now (Open WebUI web loader, crawl4server):
POST /crawlaccepts{"urls": [...]}and returns Markdown plus metadata (source, URL, title, status code, content type) for each page Open WebUI can use; invalid or failing URLs are simply left out of the response instead of failing the whole batch- Serves the same MCP endpoint as
crawl4mcpon a second port, sharing one headless browser and one concurrency limit with the web loader GET /healthfor liveness checks, and an optional bearer API key on the web loader endpoint- Settings via command-line options,
CRAWL4SERVER_*environment variables, or a YAML/JSON config file - Dockerfile and compose file for running the server in a container
- Messages in English or Japanese (see Language of messages)
Planned:
- Authentication for the MCP endpoint
Installation
The CLI and the MCP server need Python 3.11 or newer and uv.
uv tool install --with-executables-from playwright crawl4tools
playwright install chromium # downloads the headless browser (once)
This installs crawl4cli, crawl4mcp, and crawl4server from PyPI.
To install the development version instead, point uv at the git repository:
uv tool install --with-executables-from playwright git+https://github.com/xhighhongo41/crawl4tools
The browser is stored in Playwright's cache directory (for example ~/Library/Caches/ms-playwright on macOS). crawl4ai also creates a ~/.crawl4ai directory for its own data.
To run crawl4server in a container instead, see Docker below.
Usage
crawl4cli https://example.com/ # Markdown to stdout
crawl4cli -o page.md https://example.com/ # save to a file
crawl4cli -d out/ URL1 URL2 URL3 # several URLs into a directory
crawl4cli -f screenshot https://example.com/ # saves example.com.png
crawl4cli --proxy http://proxy.local:8080 URL # through a proxy
| Option | Meaning |
|---|---|
-f, --format |
markdown (default), html, pdf, screenshot, mhtml, or raw |
-o, --output FILE |
Save a single URL to FILE instead of stdout |
-d, --output-dir DIR |
Directory for several URLs and binary formats (default: current directory) |
--proxy URL |
http://, https://, or socks5:// proxy; credentials as user:pass@host:port |
--no-fallback |
Do not retry over a direct connection when the proxy fails |
-j, --concurrency N |
URLs fetched at once (default: 3) |
--timeout SECONDS |
Page load timeout per URL (default: 60) |
--fit |
Keep only the main content (drops menus, footers, and the like); falls back to the full page if nothing is left |
--citations |
Turn links into numbered references listed at the end |
--no-links, --no-images |
Drop links or image references from the Markdown |
-q, --quiet / -v, --verbose |
Less or more output on stderr |
--lang en|ja |
Language of messages (default: follows the OS locale; see Language of messages) |
Only the document goes to stdout; notes, errors, and the summary go to stderr. File names are derived from the URL (https://example.com/a/b → example.com_a_b.md). The exit code is 0 when every URL succeeded, 1 when any failed, and 2 for invalid arguments. Every option can also be set with an environment variable named CRAWL4CLI_<OPTION>, for example CRAWL4CLI_PROXY. The standard HTTP_PROXY/HTTPS_PROXY variables are not used.
Open WebUI web loader (crawl4server)
crawl4server runs a single process that serves both an Open WebUI external web loader and an MCP server, on two separate ports, sharing one headless browser and one concurrency limit (-j) between them.
Running
crawl4server
By default this listens on 127.0.0.1:8766 for the web loader (POST /crawl, GET /health) and 127.0.0.1:8765 for MCP (/mcp).
Connecting Open WebUI
In Open WebUI (0.11.x or later), under Admin Settings > Web Search, set:
- Web Loader Engine:
external - External Web Loader URL:
http://<host>:8766/crawl - External Web Loader API Key: the value passed to
--loader-api-key(leave empty if none)
The equivalent environment variables on the Open WebUI side are WEB_LOADER_ENGINE=external, EXTERNAL_WEB_LOADER_URL, and EXTERNAL_WEB_LOADER_API_KEY.
Open WebUI posts {"urls": [...]} to that URL and gets back a JSON array of {"page_content": <Markdown>, "metadata": {"source", "url", "title", "status_code", "content_type"}} documents. URLs that are invalid or fail are simply left out of the response (and logged to the server's stderr), so one bad URL never drops the whole batch. Open WebUI does not time out these requests itself, so tune --timeout and -j/--concurrency if searches feel slow.
Options
| Option | Meaning |
|---|---|
--host |
Host to listen on, both ports (default: 127.0.0.1) |
--loader-port |
Web loader port (default: 8766) |
--loader-path |
HTTP path of the web loader endpoint (default: /crawl) |
--loader-api-key KEY |
Require Authorization: Bearer KEY on the web loader (Open WebUI's External Web Loader API Key) |
--loader-fit / --no-loader-fit |
Keep only the main content of each page (default: off, full-page Markdown) |
--mcp-port |
MCP Streamable HTTP port (default: 8765) |
--mcp-path |
HTTP path of the MCP endpoint (default: /mcp) |
--proxy URL |
http://, https://, or socks5:// proxy; credentials as user:pass@host:port |
--no-fallback |
Do not retry over a direct connection when the proxy fails |
--timeout SECONDS |
Default per-URL timeout, for web loader requests and MCP tool calls that omit timeout_s (default: 60) |
-j, --concurrency N |
Maximum URLs fetched at once across both ports (default: 3) |
--max-urls N |
Maximum URLs accepted per web loader request or MCP tool call; Open WebUI sends up to 20, so keep this at 20 or more (default: 20) |
--download-dir DIR |
Root directory the MCP download tool saves files into (default: current directory) |
--config FILE |
YAML or JSON config file (see Configuration below) |
-v, --verbose |
Verbose logging on stderr |
--lang en|ja |
Language of messages the server produces while running (default: English; see Language of messages) |
Configuration
Settings are resolved in this order: command-line options > CRAWL4SERVER_* environment variables (for example CRAWL4SERVER_LOADER_API_KEY) > config file (--config or CRAWL4SERVER_CONFIG) > built-in defaults. Config file keys are the same option names in snake_case:
host: 0.0.0.0
loader_port: 8766
loader_api_key: change-me
mcp_port: 8765
max_urls: 20
concurrency: 3
lang: ja
Docker
compose.yaml pulls the published image from Docker Hub (xhighhongo41/crawl4tools, built for linux/amd64 and linux/arm64):
git clone https://github.com/xhighhongo41/crawl4tools
cd crawl4tools
mkdir -p downloads
docker compose up -d
curl http://localhost:8766/health
Without compose, the same image can be pulled directly: docker pull xhighhongo41/crawl4tools:1.0.0b1. Beta versions are not tagged latest, so always use the version tag. To build the image locally instead of pulling it, run docker build -t xhighhongo41/crawl4tools:1.0.0b1 . and then docker compose up -d.
Set CRAWL4SERVER_LOADER_API_KEY in compose.yaml's environment section. The container runs as uid 1000, so downloads/ must be writable by it; shm_size: 1gb is set for Chromium. When Open WebUI runs in the same compose project, point it at http://crawl4tools:8766/crawl instead of localhost. Stopping the container (docker stop, or docker compose down) lets requests already in progress finish, for up to 5 seconds, before closing the remaining connections.
Security
The web loader checks the API key only when --loader-api-key is set; the MCP port has no authentication at all. When listening on a non-loopback host (as in the Docker setup), set a loader API key and do not expose the MCP port beyond a trusted network — crawl4server prints a warning to stderr at startup in that case. crawl4server serves plain HTTP; if you need HTTPS, terminate TLS at a reverse proxy placed in front of it.
MCP server
crawl4mcp exposes the same fetching engine as an MCP server, with two tools: fetch (return content directly) and download (save it to files). crawl4server (above) serves this same MCP endpoint alongside the Open WebUI web loader, from one process.
Connecting a client
For Claude Code, over stdio (the default transport):
claude mcp add --transport stdio crawl4tools -- crawl4mcp
Or, over Streamable HTTP: start the server, then point Claude Code at it:
crawl4mcp --transport http
claude mcp add --transport http crawl4tools http://127.0.0.1:8765/mcp
For Claude Desktop, add an entry to claude_desktop_config.json:
{
"mcpServers": {
"crawl4tools": {
"command": "crawl4mcp",
"args": ["--download-dir", "/path/to/downloads"]
}
}
}
If Claude Desktop cannot find crawl4mcp (it does not always see your shell's PATH), replace "command": "crawl4mcp" with the full path shown by which crawl4mcp.
Other clients that support the Streamable HTTP transport (for example Open WebUI's MCP support) can connect to http://<host>:<port>/mcp once the server is running with --transport http.
Tools
| Tool | Parameters |
|---|---|
fetch |
urls (required), format (markdown default, html, screenshot), fit, citations, ignore_links, ignore_images, timeout_s |
download |
urls (required), format (markdown default, html, pdf, screenshot, mhtml, raw), directory, fit, citations, ignore_links, ignore_images, timeout_s |
urls: one or morehttp/httpsURLs, up to the server's per-call limit (--max-urls, default 20); duplicate URLs are fetched once.format: forfetch, the page comes back directly asmarkdown(default),html, or ascreenshot(PNG); image URLs are returned as images, and PDFs are transcribed to Markdown. Fordownload, the file is saved asmarkdown(default),html,pdf,screenshot,mhtml, orraw.fit,citations,ignore_links,ignore_images: the same content-shaping options as the CLI's--fit,--citations,--no-links, and--no-images(Markdown only).timeout_s: page load timeout for this call, in seconds; defaults to the server's--timeout.directory(downloadonly): subdirectory of the server's download directory (--download-dir) to save into; it cannot resolve outside that directory. Defaults to the download directory itself.
When a call covers several URLs, each result starts with a <!-- crawl4tools: url=... status=... --> line; a URL that failed is reported as an error: ... line instead, and the call only fails when every URL fails. A proxy fallback or other remark about a result appears as a <!-- note: ... --> line. download overwrites files that already exist at the destination.
Options
| Option | Meaning |
|---|---|
--transport |
stdio (default) or http |
--host |
Host to listen on (http transport only; default: 127.0.0.1) |
--port |
Port to listen on (http transport only; default: 8765) |
--path |
HTTP path for the MCP endpoint (http transport only; default: /mcp) |
--proxy URL |
http://, https://, or socks5:// proxy; credentials as user:pass@host:port |
--no-fallback |
Do not retry over a direct connection when the proxy fails |
--timeout SECONDS |
Default per-URL timeout, used when a tool call omits timeout_s (default: 60) |
-j, --concurrency N |
Maximum URLs fetched at once across every tool call (default: 3) |
--max-urls N |
Maximum URLs accepted in a single tool call (default: 20) |
--download-dir DIR |
Root directory the download tool saves files into (default: current directory) |
--config FILE |
YAML or JSON config file (see Configuration below) |
-v, --verbose |
Verbose logging on stderr |
--lang en|ja |
Language of messages the server produces while running (default: English; see Language of messages) |
Configuration
Settings are resolved in this order: command-line options > CRAWL4MCP_* environment variables > config file > built-in defaults.
Every option can be set with an environment variable named CRAWL4MCP_ followed by the option's name in upper case with dashes replaced by underscores — for example CRAWL4MCP_MAX_URLS for --max-urls, or CRAWL4MCP_CONCURRENCY for --concurrency. CRAWL4MCP_CONFIG points to the config file itself, same as --config.
The config file is YAML (JSON also works, since JSON is valid YAML), with the same option names as keys, using underscores instead of dashes; unknown keys are an error:
transport: http
host: 127.0.0.1
port: 8765
max_urls: 50
concurrency: 5
download_dir: ./downloads
proxy: http://proxy.local:8080
lang: ja
Relative paths in the config file (such as download_dir) are resolved from the current directory the server is started in.
Security
The Streamable HTTP transport has no authentication. By default the server listens on 127.0.0.1 only; if you bind it to another host, anyone who can reach that port can use the server, and crawl4mcp prints a warning to stderr when it starts. In stdio mode, stdout is reserved for the MCP protocol — all logging goes to stderr.
Language of messages
All three commands can show their messages in English (en) or Japanese (ja).
Choose the language with the --lang option, an environment variable (CRAWL4CLI_LANG, CRAWL4MCP_LANG, CRAWL4SERVER_LANG), or, for the two servers, the lang key in the config file. When more than one is set, the option wins, then the environment variable, then the config file.
Defaults: crawl4cli follows the OS locale (LANGUAGE, LC_ALL, LC_MESSAGES, then LANG, whichever is set first) and uses Japanese when it starts with ja (for example LANG=ja_JP.UTF-8), otherwise English. crawl4mcp and crawl4server always default to English, regardless of the locale, so containers and MCP clients get stable output.
crawl4cli --lang ja https://example.com/
CRAWL4SERVER_LANG=ja crawl4server
This covers --help, error messages, notes and progress lines on stderr, the MCP tools' descriptions and result text, the servers' startup/warning lines, and the web loader's JSON error responses. --help always follows --lang/the environment variable (and, for crawl4cli, the locale) — the config file's lang only applies to messages produced while the server is running.
Always in English, unchanged by --lang: log output; the error:, note:, saved:, done:, failed: line prefixes; JSON keys; the <!-- crawl4tools: url=... status=... --> header lines; GET /health; --version; and anything printed by click or other libraries (for example Usage: or Error: Invalid value ...).
Each process uses one language for its whole run; there is no per-request language.
Acknowledgements
This product includes software developed by UncleCode (https://x.com/unclecode) as part of the Crawl4AI project (https://github.com/unclecode/crawl4ai). Crawl4AI is licensed under the Apache License 2.0.
License
This project is licensed under the Apache License 2.0. See the LICENSE file for details.
Release files for crawl4tools 1.0.0b1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| crawl4tools-1.0.0b1.tar.gz | 458.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| crawl4tools-1.0.0b1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 550.6 kB
Release files / crawl4tools-1.0.0b1.tar.gz
| Download URL | crawl4tools-1.0.0b1.tar.gz |
|---|---|
| Size | 458.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
59ef9ba2f50f4a81d57c99b865a3af7d156385ba12cf6b337d76309ed62354d6
|
|
BLAKE2b-256 checksum How to use checksums |
078eaeb48fb0f9d0f080efca78194f708ac59d3e18944db5bc519ec3b1da632a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 25, 2026.
Transparency logRelease files / crawl4tools-1.0.0b1-py3-none-any.whl
| Download URL | crawl4tools-1.0.0b1-py3-none-any.whl |
|---|---|
| Size | 92.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
3ec0ae1de106e7c3c3222698808b1281a5bfd518d292448b207e1ab5ccb75a3e
|
|
BLAKE2b-256 checksum How to use checksums |
141d3c426db3a2040f23e46090f2c1d42d47182b7d4b8319f0ffc78ef2cf0f4e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 25, 2026.
Transparency log