Skip to main content

 

websift

A lightweight, free, self-hosted MCP (Model Context Protocol) server that gives AI agents real-time web access — DuckDuckGo search + web page fetching (HTML → Markdown, PDF → text) — with built-in SSRF protection and DNS pinning. No API key required.

Python 3.10+ PyPI MCP License: MIT


Table of Contents


What It Does

This MCP server exposes two tools to any AI agent or LLM client:

Tool Input Output
web_search query (string) Title, URL, and snippet for each DuckDuckGo result
web_fetch url (string) Readable text content from any webpage (HTML → Markdown, PDF → text)

That's it — simple, focused, and reliable.


Why Use This

Core Strengths

  • 🆓 Completely Free — No API keys, no subscriptions, no rate limits from a third-party provider. DuckDuckGo is free, and this server is free.
  • 🪶 Lightweight — Single Python process, ~4 dependencies, runs in a tiny Docker container (python:3.12-slim ≈ 150 MB).
  • 🔒 Secure by Default — SSRF protection with private-IP blocking, DNS resolution pinning, SNI validation, redirect limits, and content-type validation.
  • 🌐 Universal MCP Compatibility — Works with any MCP client (VS Code, Claude, Cursor, Windsurf, JetBrains, custom agents, etc.).
  • 📄 Smart Content Extraction — HTML → clean Markdown via BeautifulSoup, PDF → text via pypdf/pdfminer, binary detection, charset auto-detection.
  • 🐙 GitHub README Shortcut — Fetching a github.com/owner/repo URL automatically uses the GitHub API to grab the raw README.
  • 🏠 Self-Hosted — Full control over your data. No traffic routed through third-party services.

Ideal For

  • 🐍 Python Scripts & Apps — Import WebSearchClient directly in your code. No server, no Docker, no MCP overhead.
  • AI Agents & Agentic Workflows — Give any autonomous agent the ability to search the web and read pages on demand.
  • Development Assistants — Let Copilot, Claude, or Cursor look up documentation, error messages, or package info in real time.
  • Research & Analysis — Fetch and summarize articles, papers, or documentation pages.
  • Cost-Sensitive Deployments — Replace paid web-search APIs (Tavily, Firecrawl, Exa, etc.) with a free self-hosted alternative.
  • Air-Gapped / Private Networks — Run entirely offline (search requires internet, but fetch can work with internal URLs if you adjust security rules).

Comparison with Alternatives

Feature websift Tavily MCP Firecrawl MCP Exa MCP Brave Search MCP
Price ✅ Free 💰 Paid (free tier limited) 💰 Paid 💰 Paid 💰 Paid (free tier limited)
API Key Required ✅ No ❌ Yes ❌ Yes ❌ Yes ❌ Yes
Self-Hosted ✅ Yes ❌ No ⚠️ Partial ❌ No ❌ No
Web Search ✅ DuckDuckGo ✅ Proprietary ❌ (scrape only) ✅ Proprietary ✅ Brave
Web Fetch ✅ HTML + PDF ✅ Yes ✅ Yes (deep) ✅ Yes ❌ No
SSRF Protection ✅ Built-in ⚠️ Managed ⚠️ Managed ⚠️ Managed ⚠️ Managed
Container Size ~150 MB N/A (SaaS) ~500 MB+ N/A (SaaS) N/A (SaaS)
Dependencies 4 packages N/A Many N/A N/A
Rate Limits DuckDuckGo only Provider limits Provider limits Provider limits Provider limits
Privacy ✅ Full control ⚠️ Data to provider ⚠️ Data to provider ⚠️ Data to provider ⚠️ Data to provider

When to Choose What

Scenario Recommended
You wantfree, no-signup web access for AI websift
You needdeep scraping (JS-rendered pages, sitemaps) Firecrawl
You needsemantic search (AI-powered relevance) Exa
You wantagentic-optimized search (Tavily's extract mode) Tavily
You needmaximum privacy (self-hosted, no external calls) websift
You're building acustom AI agent with minimal infra websift

Architecture

┌─────────────┐     MCP Protocol      ┌──────────────────┐
│  AI Client  │ ◄── (streamable-HTTP) ┤   MCP Server     │
│  (Copilot,  │                       │   (FastMCP)      │
│   Claude,   │                       └─────────┬────────┘
│   Cursor…)  │                                 │
└─────────────┘                    ┌────────────┴─────────┐
                                   │  WebSearchClient     │
┌────────────┐                     │                      │
│  search()  │──► DuckDuckGo (ddgs)│                      │   
│  fetch()   │──► urllib + SSRF    │                      │
│            │    ├── html.py (BS4)│                      │
│            │    ├── http.py      │                      │
│            │    ├── security.py  │                      │
│            │    └── content.py   │                      │
└────────────┘                     └──────────────────────┘

Module Structure

web_search/
├── __init__.py    # exports WebSearchClient + __version__
├── config.py      # constants (size limits, user-agents, MIME types, ...)
├── security.py    # SSRF protection: private-IP check, DNS resolve + pin
├── content.py     # content-type detection (PDF, binary, HTML heuristics)
├── http.py        # raw HTTP fetch: redirect following, SNI pinning, charset decode
├── html.py        # HTML → Markdown conversion, text truncation
├── client.py      # WebSearchClient: search / fetch / GitHub README shortcut
└── server.py      # MCP server module (FastMCP, importable or standalone)
server.py          # thin entry point → delegates to web_search.server

Installation

Option 1: PyPI (Recommended)

pip install websift

That's it — you can now use it as a Python library (direct import) or as an MCP server (for AI clients).

Option 2: Docker Compose

# Clone and start
git clone <repo-url>
cd websift
docker compose up -d --build

The MCP server will be available at http://localhost:8787/mcp.

Option 3: Docker (Manual)

docker build -t websift .
docker run -d --name websift -p 8787:8787 websift

Option 4: Local Python (No Docker)

# Install dependencies
pip install -r requirements.txt

# Run the server
python server.py

Usage

As a Python Library (Direct Import)

Use WebSearchClient directly in your Python code — no server needed:

from web_search import WebSearchClient

client = WebSearchClient()

# Search the web (DuckDuckGo)
results = client.search("Python 3.12 features")
print(results)

# Fetch a web page (HTML → Markdown, PDF → text)
content = client.fetch("https://docs.python.org/3/")
print(content)

Perfect for:

  • Custom scripts & automation
  • Embedding in your own applications
  • Data pipelines & ETL workflows
  • Testing & prototyping

As an MCP Server (For AI Clients)

Run the server to expose tools to any MCP-compatible AI client:

# Start the server (default: port 8787)
websift

# Custom port & transport
MCP_PORT=9000 MCP_TRANSPORT=sse websift

Or via Python:

from web_search.server import main
main()

Perfect for:

  • VS Code (GitHub Copilot)
  • Claude Desktop / Claude Code
  • Cursor, Windsurf, JetBrains IDEs
  • Any MCP-compatible agent

Tools

web_search(query: str) -> str

Searches DuckDuckGo and returns formatted results with title, URL, and snippet.

Example:

Agent: web_search("latest Python 3.12 features")
Server:
Title: What's New in Python 3.12
URL: https://docs.python.org/3/whatsnew/3.12.html
Snippet: Python 3.12 introduces several performance improvements...

---

Title: Python 3.12 Release Notes
URL: https://www.python.org/downloads/release/python-3120/
Snippet: The Python 3.12 release includes bug fixes and...

web_fetch(url: str) -> str

Fetches a URL and returns readable text content. Handles:

  • HTML pages → converted to clean Markdown (BeautifulSoup, main-content extraction)
  • PDF files → text extracted via pypdf / pdfminer
  • Plain text / JSON / XML → returned as-is
  • GitHub repos → automatically fetches README via GitHub API
  • Binary files → detected and blocked (images, executables, archives)

Example:

Agent: web_fetch("https://github.com/python/cpython")
Server:
README of https://github.com/python/cpython (via GitHub API):

# Python
The Python programming language...

Configuration

Environment Variables

Variable Default Description
MCP_HOST 0.0.0.0 Bind address
MCP_PORT 8787 Listen port
MCP_TRANSPORT streamable-http Transport:streamable-http, sse, or stdio
SEARCH_MAX_RESULTS 5 Max search results returned
SEARCH_TIMEOUT 30 Request timeout in seconds

Internal Limits

Setting Value Description
Max page size 2 MB Normal page fetch limit
Max PDF size 20 MB PDF fetch limit
Max output chars 32,000 Characters sent to LLM
Max redirects 5 HTTP redirect chain limit

Connecting to AI Clients

1. VS Code (GitHub Copilot)

📖 Official docs: Add and manage MCP servers in VS Code | MCP configuration reference

Via Extensions View (Easiest)

  1. Open Extensions view (Ctrl+Shift+X)
  2. Search @mcp in the search field
  3. Install any MCP server from the gallery

Via mcp.json (Custom Server)

Create .vscode/mcp.json in your workspace:

{
  "mcpServers": {
    "web-search": {
      "type": "http",
      "url": "http://localhost:8787/mcp"
    }
  }
}

Or for global (user-level) configuration, run MCP: Open User Configuration from the Command Palette and add the same entry.

Verify

Open Chat (Ctrl+Cmd+I / Ctrl+Ctrl+I) and ask: "Search for the latest Python release notes"

2. Claude Desktop

📖 Official docs: Getting Started with Local MCP Servers on Claude Desktop | Desktop Extensions

Via Settings UI

  1. Open Claude Desktop → Settings → Extensions
  2. Click "Advanced settings" → "Install Extension…"
  3. Or manually add via the configuration file

Via Configuration File

Edit ~/.config/claude/claude_desktop_config.json (Linux) or ~/Library/Application Support/Claude/claude_desktop_config.json (macOS):

{
  "mcpServers": {
    "web-search": {
      "url": "http://localhost:8787/mcp"
    }
  }
}

Restart Claude Desktop after editing.

3. Claude Code

📖 Official docs: Connect Claude Code to tools via MCP

# Add the MCP server (HTTP transport)
claude mcp add --transport http web-search http://localhost:8787/mcp

# Verify it's connected
claude mcp list

# Use it in conversation
claude "Search for the latest Rust release and summarize the key changes"

Scopes:

# Project scope (default, stored in .mcp.json)
claude mcp add --transport http web-search http://localhost:8787/mcp

# User scope (available across all projects)
claude mcp add --transport http web-search --scope user http://localhost:8787/mcp

4. Copilot CLI

📖 Official docs: Adding MCP servers for GitHub Copilot CLI

Interactive Mode

/mcp add
# Server Name: web-search
# Server Type: HTTP
# URL: http://localhost:8787/mcp
# Press Ctrl+S to save

Command Line

copilot mcp add web-search --transport http --url http://localhost:8787/mcp

Config File

Edit ~/.github/copilot/mcp-config.json:

{
  "mcpServers": {
    "web-search": {
      "type": "http",
      "url": "http://localhost:8787/mcp"
    }
  }
}

5. Cursor

📖 Official docs: Model Context Protocol (MCP) | Cursor Docs

Create ~/.cursor/mcp.json (global) or .cursor/mcp.json (project):

{
  "mcpServers": {
    "web-search": {
      "type": "http",
      "url": "http://localhost:8787/mcp"
    }
  }
}

Restart Cursor, then use Chat or Agent mode to invoke the tools.

6. Windsurf (Codeium)

📖 Official docs: Cascade MCP Integration

Create ~/.codeium/windsurf/mcp_config.json:

{
  "mcpServers": {
    "web-search": {
      "type": "http",
      "url": "http://localhost:8787/mcp"
    }
  }
}

Alternatively, use the built-in UI: Settings → Cascade → MCP Servers.

7. JetBrains IDEs

📖 Official docs: MCP Server | IntelliJ IDEA Documentation

  1. Install the "MCP Client" plugin from the JetBrains Marketplace
  2. Go to Settings → Tools → MCP
  3. Add a new server:
    • Name: web-search
    • Type: HTTP
    • URL: http://localhost:8787/mcp
  4. Apply and restart

8. Any MCP Client (Generic)

For any MCP-compatible client, the server is accessible at:

  • Streamable HTTP (recommended): http://localhost:8787/mcp
  • SSE: http://localhost:8787/mcp/sse (set MCP_TRANSPORT=sse)
  • STDIO: Run python server.py with MCP_TRANSPORT=stdio

Generic HTTP configuration:

{
  "mcpServers": {
    "web-search": {
      "type": "http",
      "url": "http://localhost:8787/mcp"
    }
  }
}

Use Cases

AI Agent Web Research

User: "Find the latest benchmarks for LLM inference optimization"
Agent: web_search("LLM inference optimization benchmarks 2025")
Agent: web_fetch("https://example.com/benchmark-article")
Agent: [Summarizes findings from fetched content]

Documentation Lookup

User: "How do I configure CORS in FastAPI?"
Agent: web_search("FastAPI CORS configuration")
Agent: web_fetch("https://fastapi.tiangolo.com/tutorial/cors/")
Agent: [Provides code example from documentation]

Error Debugging

User: "I'm getting 'ModuleNotFoundError: no module named '_sqlite3'"
Agent: web_search("ModuleNotFoundError _sqlite3 Python Docker")
Agent: [Finds solution: install python3-dev packages]

Competitive Analysis

User: "What are the latest features in React 19?"
Agent: web_search("React 19 new features")
Agent: web_fetch("https://react.dev/blog/2024")
Agent: [Summarizes new features]

Agentic AI Workflows

This server is particularly well-suited for agentic AI because:

  • Deterministic toolsweb_search and web_fetch have clear, predictable inputs and outputs.
  • No authentication overhead — agents don't need to manage API keys.
  • Self-contained — single container, no external dependencies beyond DuckDuckGo.
  • SSRF-safe — agents can safely fetch URLs without risking internal network exposure.
  • Markdown output — clean, structured text that LLMs can process efficiently.

Security

Built-in Protections

Protection How It Works
SSRF Prevention All resolved IPs are checked against private/loopback/link-local ranges
DNS Pinning DNS resolution is pinned to the first resolved IP; SNI validation ensures certificate matches
Redirect Limits Maximum 5 redirects to prevent redirect loops and SSRF bypass
Scheme Validation Onlyhttp:// and https:// schemes are allowed
Binary Detection Images, executables, archives, and other binary content are detected and blocked
Size Limits 2 MB for normal pages, 20 MB for PDFs, 32,000 chars output limit
Charset Detection BOM detection (UTF-8/16/32), Content-Type header parsing, meta tag fallback

Network Considerations

  • The server binds to 0.0.0.0 by default — restrict with MCP_HOST=127.0.0.1 for local-only access.
  • No authentication is built in — place behind a reverse proxy (nginx, Caddy) if exposing externally.
  • Docker Compose isolates the server in its own container network.

Development

Project Structure

websift/
├── pyproject.toml          # Package metadata, dependencies, console script
├── docker-compose.yml      # Docker Compose setup
├── Dockerfile              # Python 3.12-slim container
├── requirements.txt        # Python dependencies
├── server.py               # MCP server entry point (delegates to web_search.server)
├── .env.example            # Environment variable template
├── .mcp.json               # VS Code MCP configuration
├── README.md               # This file
├── docs/
│   └── README.vi.md        # Vietnamese documentation
└── web_search/
    ├── __init__.py         # Package exports (WebSearchClient, __version__)
    ├── config.py           # Constants and configuration
    ├── security.py         # SSRF protection and DNS pinning
    ├── content.py          # Content-type detection
    ├── http.py             # HTTP fetching with SNI pinning
    ├── html.py             # HTML to Markdown conversion
    ├── client.py           # WebSearchClient (search + fetch)
    └── server.py           # MCP server module (importable or standalone)

Running Locally

# Install from PyPI
pip install websift

# Or install in editable mode (for development)
pip install -e .

# Run as MCP server
websift

# Run with custom settings
MCP_PORT=9000 MCP_TRANSPORT=sse websift

# Or use as a library (no server needed)
python -c "from web_search import WebSearchClient; print(WebSearchClient().search('test'))"

Running with Docker

# Build and start
docker compose up -d --build

# View logs
docker compose logs -f

# Stop
docker compose down

FAQ

Q: Does this require an API key?

No. DuckDuckGo search is free and doesn't require authentication. The entire server runs without any API keys.

Q: Can I use this behind a firewall?

Yes. The server only needs outbound HTTPS access to reach DuckDuckGo and target websites. Inbound access is only needed for the MCP endpoint (port 8787).

Q: How does this compare to Tavily or Firecrawl?

This is simpler and free, but doesn't offer JS rendering, deep scraping, or semantic search. For basic web search + page fetching, it's a solid free alternative. See the Comparison table for details.

Q: Can I add authentication?

The server itself doesn't include auth, but you can place it behind nginx/Caddy with basic auth or API key validation:

location /mcp {
    auth_basic "MCP Server";
    auth_basic_user_file /etc/nginx/.htpasswd;
    proxy_pass http://localhost:8787/mcp;
}

Q: What transport protocols are supported?

  • streamable-http (recommended, default) — modern MCP standard
  • sse — legacy Server-Sent Events (still supported)
  • stdio — for local process communication

Q: Why is the output limited to 32,000 characters?

This keeps responses within typical LLM context windows while providing substantial content. You can adjust MAX_PAGE_CHARS in web_search/config.py if needed.

Q: Can I use this for internal/private websites?

By default, SSRF protection blocks private IP ranges. To allow internal sites, modify web_search/security.py to whitelist specific domains or IP ranges.

Q: Can I use this as a Python library (without MCP)?

Yes! Just pip install websift and import directly:

from web_search import WebSearchClient
client = WebSearchClient()
client.search("your query")
client.fetch("https://example.com")

No server, no Docker, no MCP overhead — just pure Python.

Q: Can I publish my own version to PyPI?

Yes. After making changes:

# Install build tools
pip install build twine

# Build the package
python -m build

# Test upload (TestPyPI)
twine upload --repository testpypi dist/*

# Real upload (PyPI)
twine upload dist/*

License

MIT — see LICENSE for details.

Acknowledgments

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

websift-0.1.0.tar.gz (26.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

websift-0.1.0-py3-none-any.whl (19.7 kB view details)

Uploaded Python 3

File details

Details for the file websift-0.1.0.tar.gz.

File metadata

  • Download URL: websift-0.1.0.tar.gz
  • Upload date:
  • Size: 26.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.10.12

File hashes

Hashes for websift-0.1.0.tar.gz
Algorithm Hash digest
SHA256 183963ea99bb10e1fa377b69a621fafdedd14987cccf9621a51c515bc5d0edc5
MD5 4178a640154927030ed67dd19d414051
BLAKE2b-256 8d049dbc2924fa0435f921bf0a09b7ee470ce2a6a59c558d8806fa33a15dd80b

See more details on using hashes here.

File details

Details for the file websift-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: websift-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 19.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.10.12

File hashes

Hashes for websift-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 ac7cbcc5c41476a8ba1034734fca4e4eb4c5f952aaf088d1532ef058653de082
MD5 ccbd59f61379efa3e79f0042e984001b
BLAKE2b-256 ebc22a83fe51319b1cc0803c22ee2dcec9977fa7d51d596f5bdadda79de364f4

See more details on using hashes here.

Release history Release notifications | RSS feed

2.0.0

2 files

1.3.0

2 files

1.2.0

2 files

1.1.0

2 files

1.0.1

2 files

1.0.0

2 files

0.3.1

2 files

0.3.0

2 files

0.2.0

2 files

This release

0.1.0 This release

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page