A high-performance, asynchronous MCP server for Firecrawl Search, featuring connection pooling, request retries, and intelligent input parsing.
Project description
Firecrawl MCP Toolkit
A high-performance, asynchronous MCP server that provides comprehensive Google search and web content scraping capabilities through the Firecrawl API (excluding some rarely used interfaces).
This project is built on httpx, utilizing asynchronous clients and connection pool management to offer LLMs a stable and efficient external information retrieval tool.
PyPI Package
firecrawl-toolkit: https://pypi.org/project/firecrawl-toolkit/
Key Features
- Asynchronous Architecture: Fully based on
asyncioandhttpx, ensuring high throughput and non-blocking I/O operations. - HTTP Connection Pool: Manages and reuses TCP connections through a global
httpx.AsyncClientinstance, significantly improving performance under high concurrency. - Concurrency Control: Built-in global and per-API endpoint concurrency semaphores effectively manage API request rates to prevent exceeding rate limits.
- Automatic Retry Mechanism: Integrated request retry functionality with exponential backoff strategy automatically handles temporary network fluctuations or server errors, enhancing service stability.
- Intelligent Country Code Parsing: Includes a comprehensive country name dictionary supporting inputs in Chinese, English, ISO Alpha-2/3, and other formats, with automatic normalization.
- Response Field Mapping: Search/Scrape responses are normalized into minimal, client-facing JSON schemas instead of upstream passthrough payloads.
- Noise Reduction for Scrape: Built-in
excludeTagsselector filtering removes common non-content blocks (navigation, ads, sidebars, comments, etc.) to improve signal quality. Supports returning a specified Markdown character window withstartIndexandmaxCharacters. - Flexible Environment Variable Configuration: Supports fine-tuned service configuration via environment variables.
- The Search and Scrape Endpoints perform some request pre-processing and post-processing, which can save quite a few tokens.
Available Tools
This service provides the following tools:
| Tool Name | Description |
|---|---|
firecrawl-aggregated-search |
Aggregated Search Interface, Combining Webpage, News, And Image Search Results. |
firecrawl-web-search |
Web Search Interface. |
firecrawl-news-search |
News Search Interface. |
firecrawl-image-search |
Image Search Interface. |
firecrawl-scrape |
Scrapes and returns the content of a specified URL. |
Installation Guide
It is recommended to install using pip or uv.
# Using pip
pip install firecrawl-toolkit
# Or using uv
uv pip install firecrawl-toolkit
Quick Start
Set Environment Variables
Create a .env file in the project root directory and enter your Firecrawl API key:
| Environment Variables | Default value | Description |
|---|---|---|
FIRECRAWL_API_KEY |
fc-xxx | Your Firecrawl API key. Multiple keys can be separated by commas, and one will be selected randomly for each request. |
FIRECRAWL_HTTP2 |
0 | Disable or enable HTTP2, <0/1> |
FIRECRAWL_MAX_WORKERS |
10 | Number of processes |
FIRECRAWL_MAX_CONNECTIONS |
200 | Maximum number of connections |
FIRECRAWL_MAX_CONCURRENT_REQUESTS |
200 | Maximum number of concurrent requests |
FIRECRAWL_KEEPALIVE |
20 | Maximum number of concurrent connections |
FIRECRAWL_RETRY_COUNT |
3 | Maximum number of retries |
FIRECRAWL_RETRY_BASE_DELAY |
0.5 | Base delay time for retries in seconds |
FIRECRAWL_ENDPOINT_CONCURRENCY |
{"search":10,"scrape":2} |
Set concurrency per endpoint (JSON format) |
FIRECRAWL_ENDPOINT_RETRYABLE |
{"scrape": false} |
Set retry allowance per endpoint (JSON format) |
FIRECRAWL_MCP_ENABLE_STDIO |
0 | Disable or enable STDIO, <0/1> |
FIRECRAWL_MCP_ENABLE_HTTP |
0 | Disable or enable HTTP, <0/1> |
FIRECRAWL_MCP_ENABLE_SSE |
0 | Disable or enable SSE, <0/1> |
FIRECRAWL_MCP_HTTP_HOST |
127.0.0.1 | HTTP host address |
FIRECRAWL_MCP_HTTP_PORT |
7001 | HTTP host port |
FIRECRAWL_MCP_SSE_HOST |
127.0.0.1 | SSE host address |
FIRECRAWL_MCP_SSE_PORT |
7001 | SSE host port |
FIRECRAWL_MCP_LOCK_FILE |
/tmp/firecrawl_mcp.lock |
Lock file path |
- STDIO, HTTP, and SSE can only be used one at a time. If you need to use multiple protocols, please start separate services for each.
- When using multiple services, please specify different lock files for each.
Configure MCP Client
Add the following server configuration in the MCP client configuration file:
{
"mcpServers": {
"firecrawl": {
"command": "python3",
"args": ["-m", "firecrawl-toolkit"],
"env": {
"FIRECRAWL_API_KEY": "<Your Firecrawl API key>"
}
}
}
}
{
"mcpServers": {
"firecrawl": {
"command": "uvx",
"args": ["firecrawl-toolkit"],
"env": {
"FIRECRAWL_API_KEY": "<Your Firecrawl API key>"
}
}
}
}
Tool Parameters and Usage Examples
firecrawl Search: Perform aggregated / web / news / images search
Parameters:
query(str, required): Keywords to search.country(str, optional): Specify the country/region for search results. Supports Chinese names (e.g., "China"), English names (e.g., "United States"), or ISO codes (e.g., "US"). Default is "US".search_num(int, optional): Number of results to return, range 1-100. Default is 20.search_time(str, optional): Filter results by time range. Available values: "hour", "day", "week", "month", "year".
Example:
result_json = firecrawl_web_search(
query="AI advancements 2024",
country="United States",
search_num=5,
search_time="month"
)
Response (mapped):
- Top-level fields:
success,data,creditsUsed data.web[]:title,description,urldata.news[]:title,snippet,url,datedata.images[]:title,imageUrl,urlweb/news/imagesremain arrays and may be empty ([])- Missing mapped fields are preserved as
null - Output is compact single-line JSON (no extra spaces)
Example response:
{"success":true,"data":{"web":[{"title":"Example Web","description":"Example description","url":"https://example.com"}],"news":[],"images":[]},"creditsUsed":1}
firecrawl-scrape: Scrape webpage content
Parameters:
url(str, required): URL of the target webpage.excludeTags(list[str], optional, default[]): Additional CSS selectors to exclude; merged with built-in noise-filter selectors after normalization and deduplication unlessemptyTags=True.includeTags(list[str], optional, defaultNone): Additional CSS selectors to include; no built-in defaults are applied, and the cleaned list is forwarded only when this parameter is provided.maxCharacters(int, optional, defaultNone): Truncate only the returnedmarkdownto N characters starting atstartIndex. Invalid values (non-int,<= 0) are ignored and treated as not provided.startIndex(int, optional, default0): Start offset used withmaxCharacterswhen slicing returnedmarkdown. Invalid values (non-int,< 0) are treated as0.emptyTags(bool, optional, defaultFalse): Clear the built-in exclude selector list for this request, while still keeping any user-providedexcludeTags.headers(dict[str, str], optional, defaultNone): Root-level request headers passed through to the upstream scrape request only when a non-empty object is provided.
Example:
result_json = firecrawl_scrape(
url="https://www.example.com",
includeTags=["article", ".content"],
excludeTags=["[class^=\"skip\"]", "[id*=\"disqus\"]"],
startIndex=0,
maxCharacters=1200,
headers={"Authorization": "Bearer token", "X-Trace-Id": "abc123"}
)
This returns at most 1200 characters in markdown, starting at character index 0.
To explicitly send an empty include selector list:
result_json = firecrawl_scrape(
url="https://www.example.com",
includeTags=[]
)
To disable only the built-in exclude selectors for one request:
result_json = firecrawl_scrape(
url="https://www.example.com",
emptyTags=True
)
To disable the built-in exclude selectors but keep your own:
result_json = firecrawl_scrape(
url="https://www.example.com",
excludeTags=[".nav"],
emptyTags=True
)
Built-in noise filtering:
- The tool uses an internal
excludeTagsselector set to suppress noisy DOM regions and prioritize main content quality. includeTagshas no built-in defaults and is only forwarded when explicitly provided.- Passing
emptyTags=Trueclears only the built-in exclude selector set for that request. - If the first scrape returns
data.markdown == "", the tool automatically retries once withoutincludeTags/excludeTagsas a fallback. startIndex/maxCharactersslicing is applied locally in this toolkit post-processing and is not forwarded to upstream Firecrawl payloads.
Response (mapped):
- Top-level fields:
success,proxyUsed,title,description,language,markdown,creditsUsed markdownis URL-decoded before returning to the client- When a valid
maxCharactersis provided,markdownlength is capped at that value after applyingstartIndex - Missing mapped fields are preserved as
null - Output is compact single-line JSON (no extra spaces)
Example response:
{"success":true,"proxyUsed":"auto","title":"Example Page","description":"Example summary","language":"en","markdown":"Hello world!","creditsUsed":1}
Response Contract Notes
firecrawl-searchandfirecrawl-scrapesuccess payloads are mapped to stable minimal schemas.- Missing mapped fields are preserved as
null(arrays remain arrays, and may be empty). - Both success and error responses are compact single-line JSON.
License Agreement
This project is licensed under the MIT License.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file firecrawl_toolkit-0.0.24.tar.gz.
File metadata
- Download URL: firecrawl_toolkit-0.0.24.tar.gz
- Upload date:
- Size: 33.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
012e7c839207167b51604cbd22fe6fa4cb6887ac1c9955d34d9305ce39ec652e
|
|
| MD5 |
c2b3be856da3645cb975eace2bc28243
|
|
| BLAKE2b-256 |
9ffda9508e46bd2fbf1c89c4a073e787876b0ac32f188083e2d88290c61e1a0b
|
Provenance
The following attestation bundles were made for firecrawl_toolkit-0.0.24.tar.gz:
Publisher:
publish-pypi.yml on Joey-Kot/firecrawl-toolkit
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
firecrawl_toolkit-0.0.24.tar.gz -
Subject digest:
012e7c839207167b51604cbd22fe6fa4cb6887ac1c9955d34d9305ce39ec652e - Sigstore transparency entry: 1629711061
- Sigstore integration time:
-
Permalink:
Joey-Kot/firecrawl-toolkit@e6cfbf07922a5b45801c218f6d0261f928d871ae -
Branch / Tag:
refs/heads/main - Owner: https://github.com/Joey-Kot
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish-pypi.yml@e6cfbf07922a5b45801c218f6d0261f928d871ae -
Trigger Event:
push
-
Statement type:
File details
Details for the file firecrawl_toolkit-0.0.24-py3-none-any.whl.
File metadata
- Download URL: firecrawl_toolkit-0.0.24-py3-none-any.whl
- Upload date:
- Size: 26.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0759f2ca2efd18179e2f32ecd26dad5aa866b980fe9e745da975e2bfdf637fdf
|
|
| MD5 |
d1465de5aff488ee6420e0a36d638610
|
|
| BLAKE2b-256 |
22451cbf84f520685ecd822e939771d033a723f5ed620f7bc5bc36ddc5f3395a
|
Provenance
The following attestation bundles were made for firecrawl_toolkit-0.0.24-py3-none-any.whl:
Publisher:
publish-pypi.yml on Joey-Kot/firecrawl-toolkit
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
firecrawl_toolkit-0.0.24-py3-none-any.whl -
Subject digest:
0759f2ca2efd18179e2f32ecd26dad5aa866b980fe9e745da975e2bfdf637fdf - Sigstore transparency entry: 1629711088
- Sigstore integration time:
-
Permalink:
Joey-Kot/firecrawl-toolkit@e6cfbf07922a5b45801c218f6d0261f928d871ae -
Branch / Tag:
refs/heads/main - Owner: https://github.com/Joey-Kot
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish-pypi.yml@e6cfbf07922a5b45801c218f6d0261f928d871ae -
Trigger Event:
push
-
Statement type: