websift
A lightweight, free, self-hosted MCP (Model Context Protocol) server that gives AI agents real-time web access — DuckDuckGo search + web page fetching (HTML → Markdown, PDF → text) — with built-in SSRF protection and DNS pinning. No API key required.
Table of Contents
- What It Does
- Why Use This
- Comparison with Alternatives
- Architecture
- Installation
- Usage
- Tools
- Configuration
- Connecting to AI Clients
- Use Cases
- Security
- Development
- FAQ
What It Does
This MCP server exposes two tools to any AI agent or LLM client:
| Tool | Input | Output |
|---|---|---|
web_search |
query (string) |
Title, URL, and snippet for each DuckDuckGo result |
web_fetch |
url (string) |
Readable text content from any webpage (HTML → Markdown, PDF → text) |
That's it — simple, focused, and reliable.
Why Use This
Core Strengths
- 🆓 Completely Free — No API keys, no subscriptions, no rate limits from a third-party provider. DuckDuckGo is free, and this server is free.
- 🪶 Lightweight — Single Python process, ~4 dependencies, runs in a tiny Docker container (
python:3.12-slim≈ 150 MB). - 🔒 Secure by Default — SSRF protection with private-IP blocking, DNS resolution pinning, SNI validation, redirect limits, and content-type validation.
- 🌐 Universal MCP Compatibility — Works with any MCP client (VS Code, Claude, Cursor, Windsurf, JetBrains, custom agents, etc.).
- 📄 Smart Content Extraction — HTML → clean Markdown via BeautifulSoup, PDF → text via pypdf/pdfminer, binary detection, charset auto-detection.
- 🐙 GitHub README Shortcut — Fetching a
github.com/owner/repoURL automatically uses the GitHub API to grab the raw README. - 🏠 Self-Hosted — Full control over your data. No traffic routed through third-party services.
Ideal For
- 🐍 Python Scripts & Apps — Import
WebSearchClientdirectly in your code. No server, no Docker, no MCP overhead. - AI Agents & Agentic Workflows — Give any autonomous agent the ability to search the web and read pages on demand.
- Development Assistants — Let Copilot, Claude, or Cursor look up documentation, error messages, or package info in real time.
- Research & Analysis — Fetch and summarize articles, papers, or documentation pages.
- Cost-Sensitive Deployments — Replace paid web-search APIs (Tavily, Firecrawl, Exa, etc.) with a free self-hosted alternative.
- Air-Gapped / Private Networks — Run entirely offline (search requires internet, but fetch can work with internal URLs if you adjust security rules).
Comparison with Alternatives
| Feature | websift | Tavily MCP | Firecrawl MCP | Exa MCP | Brave Search MCP |
|---|---|---|---|---|---|
| Price | ✅ Free | 💰 Paid (free tier limited) | 💰 Paid | 💰 Paid | 💰 Paid (free tier limited) |
| API Key Required | ✅ No | ❌ Yes | ❌ Yes | ❌ Yes | ❌ Yes |
| Self-Hosted | ✅ Yes | ❌ No | ⚠️ Partial | ❌ No | ❌ No |
| Web Search | ✅ DuckDuckGo | ✅ Proprietary | ❌ (scrape only) | ✅ Proprietary | ✅ Brave |
| Web Fetch | ✅ HTML + PDF | ✅ Yes | ✅ Yes (deep) | ✅ Yes | ❌ No |
| SSRF Protection | ✅ Built-in | ⚠️ Managed | ⚠️ Managed | ⚠️ Managed | ⚠️ Managed |
| Container Size | ~150 MB | N/A (SaaS) | ~500 MB+ | N/A (SaaS) | N/A (SaaS) |
| Dependencies | 4 packages | N/A | Many | N/A | N/A |
| Rate Limits | DuckDuckGo only | Provider limits | Provider limits | Provider limits | Provider limits |
| Privacy | ✅ Full control | ⚠️ Data to provider | ⚠️ Data to provider | ⚠️ Data to provider | ⚠️ Data to provider |
When to Choose What
| Scenario | Recommended |
|---|---|
| You wantfree, no-signup web access for AI | websift ✅ |
| You needdeep scraping (JS-rendered pages, sitemaps) | Firecrawl |
| You needsemantic search (AI-powered relevance) | Exa |
You wantagentic-optimized search (Tavily's extract mode) |
Tavily |
| You needmaximum privacy (self-hosted, no external calls) | websift ✅ |
| You're building acustom AI agent with minimal infra | websift ✅ |
Architecture
┌─────────────┐ MCP Protocol ┌──────────────────┐
│ AI Client │ ◄── (streamable-HTTP) ┤ MCP Server │
│ (Copilot, │ │ (FastMCP) │
│ Claude, │ └─────────┬────────┘
│ Cursor…) │ │
└─────────────┘ ┌────────────┴─────────┐
│ WebSearchClient │
┌────────────┐ │ │
│ search() │──► DuckDuckGo (ddgs)│ │
│ fetch() │──► urllib + SSRF │ │
│ │ ├── html.py (BS4)│ │
│ │ ├── http.py │ │
│ │ ├── security.py │ │
│ │ └── content.py │ │
└────────────┘ └──────────────────────┘
Module Structure
web_search/
├── __init__.py # exports WebSearchClient + __version__
├── config.py # constants (size limits, user-agents, MIME types, ...)
├── security.py # SSRF protection: private-IP check, DNS resolve + pin
├── content.py # content-type detection (PDF, binary, HTML heuristics)
├── http.py # raw HTTP fetch: redirect following, SNI pinning, charset decode
├── html.py # HTML → Markdown conversion, text truncation
├── client.py # WebSearchClient: search / fetch / GitHub README shortcut
└── server.py # MCP server module (FastMCP, importable or standalone)
server.py # thin entry point → delegates to web_search.server
Installation
Option 1: PyPI (Recommended)
pip install websift
That's it — you can now use it as a Python library (direct import) or as an MCP server (for AI clients).
Option 2: Docker Compose
# Clone and start
git clone <repo-url>
cd websift
docker compose up -d --build
The MCP server will be available at http://localhost:8787/mcp.
Option 3: Docker (Manual)
docker build -t websift .
docker run -d --name websift -p 8787:8787 websift
Option 4: Local Python (No Docker)
# Install dependencies
pip install -r requirements.txt
# Run the server
python server.py
Usage
As a Python Library (Direct Import)
Use WebSearchClient directly in your Python code — no server needed:
from web_search import WebSearchClient
client = WebSearchClient()
# Search the web (DuckDuckGo)
results = client.search("Python 3.12 features")
print(results)
# Fetch a web page (HTML → Markdown, PDF → text)
content = client.fetch("https://docs.python.org/3/")
print(content)
Perfect for:
- Custom scripts & automation
- Embedding in your own applications
- Data pipelines & ETL workflows
- Testing & prototyping
As an MCP Server (For AI Clients)
Run the server to expose tools to any MCP-compatible AI client:
# Start the server (default: port 8787)
websift
# Custom port & transport
MCP_PORT=9000 MCP_TRANSPORT=sse websift
Or via Python:
from web_search.server import main
main()
Perfect for:
- VS Code (GitHub Copilot)
- Claude Desktop / Claude Code
- Cursor, Windsurf, JetBrains IDEs
- Any MCP-compatible agent
Tools
web_search(query: str) -> str
Searches DuckDuckGo and returns formatted results with title, URL, and snippet.
Example:
Agent: web_search("latest Python 3.12 features")
Server:
Title: What's New in Python 3.12
URL: https://docs.python.org/3/whatsnew/3.12.html
Snippet: Python 3.12 introduces several performance improvements...
---
Title: Python 3.12 Release Notes
URL: https://www.python.org/downloads/release/python-3120/
Snippet: The Python 3.12 release includes bug fixes and...
web_fetch(url: str) -> str
Fetches a URL and returns readable text content. Handles:
- HTML pages → converted to clean Markdown (BeautifulSoup, main-content extraction)
- PDF files → text extracted via pypdf / pdfminer
- Plain text / JSON / XML → returned as-is
- GitHub repos → automatically fetches README via GitHub API
- Binary files → detected and blocked (images, executables, archives)
Example:
Agent: web_fetch("https://github.com/python/cpython")
Server:
README of https://github.com/python/cpython (via GitHub API):
# Python
The Python programming language...
Configuration
Environment Variables
| Variable | Default | Description |
|---|---|---|
MCP_HOST |
0.0.0.0 |
Bind address |
MCP_PORT |
8787 |
Listen port |
MCP_TRANSPORT |
streamable-http |
Transport:streamable-http, sse, or stdio |
SEARCH_MAX_RESULTS |
5 |
Max search results returned |
SEARCH_TIMEOUT |
30 |
Request timeout in seconds |
Internal Limits
| Setting | Value | Description |
|---|---|---|
| Max page size | 2 MB | Normal page fetch limit |
| Max PDF size | 20 MB | PDF fetch limit |
| Max output chars | 32,000 | Characters sent to LLM |
| Max redirects | 5 | HTTP redirect chain limit |
Connecting to AI Clients
1. VS Code (GitHub Copilot)
📖 Official docs: Add and manage MCP servers in VS Code | MCP configuration reference
Via Extensions View (Easiest)
- Open Extensions view (
Ctrl+Shift+X) - Search
@mcpin the search field - Install any MCP server from the gallery
Via mcp.json (Custom Server)
Create .vscode/mcp.json in your workspace:
{
"mcpServers": {
"web-search": {
"type": "http",
"url": "http://localhost:8787/mcp"
}
}
}
Or for global (user-level) configuration, run MCP: Open User Configuration from the Command Palette and add the same entry.
Verify
Open Chat (Ctrl+Cmd+I / Ctrl+Ctrl+I) and ask: "Search for the latest Python release notes"
2. Claude Desktop
📖 Official docs: Getting Started with Local MCP Servers on Claude Desktop | Desktop Extensions
Via Settings UI
- Open Claude Desktop → Settings → Extensions
- Click "Advanced settings" → "Install Extension…"
- Or manually add via the configuration file
Via Configuration File
Edit ~/.config/claude/claude_desktop_config.json (Linux) or ~/Library/Application Support/Claude/claude_desktop_config.json (macOS):
{
"mcpServers": {
"web-search": {
"url": "http://localhost:8787/mcp"
}
}
}
Restart Claude Desktop after editing.
3. Claude Code
📖 Official docs: Connect Claude Code to tools via MCP
# Add the MCP server (HTTP transport)
claude mcp add --transport http web-search http://localhost:8787/mcp
# Verify it's connected
claude mcp list
# Use it in conversation
claude "Search for the latest Rust release and summarize the key changes"
Scopes:
# Project scope (default, stored in .mcp.json)
claude mcp add --transport http web-search http://localhost:8787/mcp
# User scope (available across all projects)
claude mcp add --transport http web-search --scope user http://localhost:8787/mcp
4. Copilot CLI
📖 Official docs: Adding MCP servers for GitHub Copilot CLI
Interactive Mode
/mcp add
# Server Name: web-search
# Server Type: HTTP
# URL: http://localhost:8787/mcp
# Press Ctrl+S to save
Command Line
copilot mcp add web-search --transport http --url http://localhost:8787/mcp
Config File
Edit ~/.github/copilot/mcp-config.json:
{
"mcpServers": {
"web-search": {
"type": "http",
"url": "http://localhost:8787/mcp"
}
}
}
5. Cursor
📖 Official docs: Model Context Protocol (MCP) | Cursor Docs
Create ~/.cursor/mcp.json (global) or .cursor/mcp.json (project):
{
"mcpServers": {
"web-search": {
"type": "http",
"url": "http://localhost:8787/mcp"
}
}
}
Restart Cursor, then use Chat or Agent mode to invoke the tools.
6. Windsurf (Codeium)
📖 Official docs: Cascade MCP Integration
Create ~/.codeium/windsurf/mcp_config.json:
{
"mcpServers": {
"web-search": {
"type": "http",
"url": "http://localhost:8787/mcp"
}
}
}
Alternatively, use the built-in UI: Settings → Cascade → MCP Servers.
7. JetBrains IDEs
📖 Official docs: MCP Server | IntelliJ IDEA Documentation
- Install the "MCP Client" plugin from the JetBrains Marketplace
- Go to Settings → Tools → MCP
- Add a new server:
- Name:
web-search - Type:
HTTP - URL:
http://localhost:8787/mcp
- Name:
- Apply and restart
8. Any MCP Client (Generic)
For any MCP-compatible client, the server is accessible at:
- Streamable HTTP (recommended):
http://localhost:8787/mcp - SSE:
http://localhost:8787/mcp/sse(setMCP_TRANSPORT=sse) - STDIO: Run
python server.pywithMCP_TRANSPORT=stdio
Generic HTTP configuration:
{
"mcpServers": {
"web-search": {
"type": "http",
"url": "http://localhost:8787/mcp"
}
}
}
Use Cases
AI Agent Web Research
User: "Find the latest benchmarks for LLM inference optimization"
Agent: web_search("LLM inference optimization benchmarks 2025")
Agent: web_fetch("https://example.com/benchmark-article")
Agent: [Summarizes findings from fetched content]
Documentation Lookup
User: "How do I configure CORS in FastAPI?"
Agent: web_search("FastAPI CORS configuration")
Agent: web_fetch("https://fastapi.tiangolo.com/tutorial/cors/")
Agent: [Provides code example from documentation]
Error Debugging
User: "I'm getting 'ModuleNotFoundError: no module named '_sqlite3'"
Agent: web_search("ModuleNotFoundError _sqlite3 Python Docker")
Agent: [Finds solution: install python3-dev packages]
Competitive Analysis
User: "What are the latest features in React 19?"
Agent: web_search("React 19 new features")
Agent: web_fetch("https://react.dev/blog/2024")
Agent: [Summarizes new features]
Agentic AI Workflows
This server is particularly well-suited for agentic AI because:
- Deterministic tools —
web_searchandweb_fetchhave clear, predictable inputs and outputs. - No authentication overhead — agents don't need to manage API keys.
- Self-contained — single container, no external dependencies beyond DuckDuckGo.
- SSRF-safe — agents can safely fetch URLs without risking internal network exposure.
- Markdown output — clean, structured text that LLMs can process efficiently.
Security
Built-in Protections
| Protection | How It Works |
|---|---|
| SSRF Prevention | All resolved IPs are checked against private/loopback/link-local ranges |
| DNS Pinning | DNS resolution is pinned to the first resolved IP; SNI validation ensures certificate matches |
| Redirect Limits | Maximum 5 redirects to prevent redirect loops and SSRF bypass |
| Scheme Validation | Onlyhttp:// and https:// schemes are allowed |
| Binary Detection | Images, executables, archives, and other binary content are detected and blocked |
| Size Limits | 2 MB for normal pages, 20 MB for PDFs, 32,000 chars output limit |
| Charset Detection | BOM detection (UTF-8/16/32), Content-Type header parsing, meta tag fallback |
Network Considerations
- The server binds to
0.0.0.0by default — restrict withMCP_HOST=127.0.0.1for local-only access. - No authentication is built in — place behind a reverse proxy (nginx, Caddy) if exposing externally.
- Docker Compose isolates the server in its own container network.
Development
Project Structure
websift/
├── pyproject.toml # Package metadata, dependencies, console script
├── docker-compose.yml # Docker Compose setup
├── Dockerfile # Python 3.12-slim container
├── requirements.txt # Python dependencies
├── server.py # MCP server entry point (delegates to web_search.server)
├── .env.example # Environment variable template
├── .mcp.json # VS Code MCP configuration
├── README.md # This file
├── docs/
│ └── README.vi.md # Vietnamese documentation
└── web_search/
├── __init__.py # Package exports (WebSearchClient, __version__)
├── config.py # Constants and configuration
├── security.py # SSRF protection and DNS pinning
├── content.py # Content-type detection
├── http.py # HTTP fetching with SNI pinning
├── html.py # HTML to Markdown conversion
├── client.py # WebSearchClient (search + fetch)
└── server.py # MCP server module (importable or standalone)
Running Locally
# Install from PyPI
pip install websift
# Or install in editable mode (for development)
pip install -e .
# Run as MCP server
websift
# Run with custom settings
MCP_PORT=9000 MCP_TRANSPORT=sse websift
# Or use as a library (no server needed)
python -c "from web_search import WebSearchClient; print(WebSearchClient().search('test'))"
Running with Docker
# Build and start
docker compose up -d --build
# View logs
docker compose logs -f
# Stop
docker compose down
FAQ
Q: Does this require an API key?
No. DuckDuckGo search is free and doesn't require authentication. The entire server runs without any API keys.
Q: Can I use this behind a firewall?
Yes. The server only needs outbound HTTPS access to reach DuckDuckGo and target websites. Inbound access is only needed for the MCP endpoint (port 8787).
Q: How does this compare to Tavily or Firecrawl?
This is simpler and free, but doesn't offer JS rendering, deep scraping, or semantic search. For basic web search + page fetching, it's a solid free alternative. See the Comparison table for details.
Q: Can I add authentication?
The server itself doesn't include auth, but you can place it behind nginx/Caddy with basic auth or API key validation:
location /mcp {
auth_basic "MCP Server";
auth_basic_user_file /etc/nginx/.htpasswd;
proxy_pass http://localhost:8787/mcp;
}
Q: What transport protocols are supported?
- streamable-http (recommended, default) — modern MCP standard
- sse — legacy Server-Sent Events (still supported)
- stdio — for local process communication
Q: Why is the output limited to 32,000 characters?
This keeps responses within typical LLM context windows while providing substantial content. You can adjust MAX_PAGE_CHARS in web_search/config.py if needed.
Q: Can I use this for internal/private websites?
By default, SSRF protection blocks private IP ranges. To allow internal sites, modify web_search/security.py to whitelist specific domains or IP ranges.
Q: Can I use this as a Python library (without MCP)?
Yes! Just pip install websift and import directly:
from web_search import WebSearchClient
client = WebSearchClient()
client.search("your query")
client.fetch("https://example.com")
No server, no Docker, no MCP overhead — just pure Python.
Q: Can I publish my own version to PyPI?
Yes. After making changes:
# Install build tools
pip install build twine
# Build the package
python -m build
# Test upload (TestPyPI)
twine upload --repository testpypi dist/*
# Real upload (PyPI)
twine upload dist/*
License
MIT — see LICENSE for details.
Acknowledgments
- DuckDuckGo Search (ddgs) — search backend
- BeautifulSoup — HTML parsing
- pypdf — PDF text extraction
- FastMCP — MCP server framework
- Model Context Protocol — open protocol for AI tool integration
- VS Code MCP Documentation — official VS Code MCP guide
- Claude Code MCP Documentation — official Claude Code MCP guide
- Cursor MCP Documentation — official Cursor MCP guide
- Windsurf MCP Documentation — official Windsurf MCP guide
- Copilot CLI MCP Documentation — official GitHub Copilot CLI MCP guide
- Claude Desktop MCP Documentation — official Claude Desktop MCP guide
- JetBrains MCP Documentation — official JetBrains MCP guide
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file websift-0.1.0.tar.gz.
File metadata
- Download URL: websift-0.1.0.tar.gz
- Upload date:
- Size: 26.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.10.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
183963ea99bb10e1fa377b69a621fafdedd14987cccf9621a51c515bc5d0edc5
|
|
| MD5 |
4178a640154927030ed67dd19d414051
|
|
| BLAKE2b-256 |
8d049dbc2924fa0435f921bf0a09b7ee470ce2a6a59c558d8806fa33a15dd80b
|
File details
Details for the file websift-0.1.0-py3-none-any.whl.
File metadata
- Download URL: websift-0.1.0-py3-none-any.whl
- Upload date:
- Size: 19.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.10.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ac7cbcc5c41476a8ba1034734fca4e4eb4c5f952aaf088d1532ef058653de082
|
|
| MD5 |
ccbd59f61379efa3e79f0042e984001b
|
|
| BLAKE2b-256 |
ebc22a83fe51319b1cc0803c22ee2dcec9977fa7d51d596f5bdadda79de364f4
|