Skip to main content

webscout-mcp

PyPI version Python versions Tests Docker Pulls License

AI-powered web intelligence platform for AI agents. Search, fetch, crawl, extract, understand, and monitor the web — with built-in AI, vector search, browser automation, and alerting. Everything stays on your machine.

✨ Features

🔍 Core Web Tools

  • Multi-backend search — Bing, DuckDuckGo, Google, Brave HTML with automatic failover and result merging
  • Smart content extraction — trafilatura + readability-lxml + html2text fallback, clean article content
  • Concurrent crawler — BFS crawl with depth/page limits, robots.txt compliance, retry on failures
  • Structured data extraction — CSS selectors, attributes, regex extraction
  • Metadata extraction — JSON-LD, OpenGraph, Twitter Cards, article metadata, images, links
  • RSS/Atom support — parse feeds and feed indexes

🤖 AI Content Understanding

  • Text summarization — automatic article and page summarization
  • Question answering — ask questions about fetched content
  • Key points extraction — extract main ideas and takeaways
  • Content classification — categorize content into custom categories
  • Tag generation — auto-generate relevant tags
  • Sentiment analysis — analyze text sentiment
  • Document comparison — compare two documents side by side
  • Entity extraction — extract people, places, organizations, dates
  • Multiple LLM backends — Ollama (local/free), OpenAI, Doubao, custom OpenAI-compatible

🧠 Vector Search & RAG

  • Semantic search — search by meaning, not just keywords
  • RAG (Retrieval-Augmented Generation) — answer questions based on your crawled content
  • Local vector database — ChromaDB persistent storage
  • Multiple embedding backends — local sentence-transformers (free), OpenAI, custom
  • Document chunking — automatic text splitting with overlap
  • Similarity threshold — configurable relevance filtering

🌐 Headless Browser Automation

  • JavaScript rendering — fetch modern SPAs and dynamic content
  • User interaction simulation — scroll, click, fill forms
  • Screenshot capture — full-page screenshots
  • PDF export — convert web pages to PDF
  • Login state management — cookie persistence across sessions
  • Anti-detection stealth mode — navigator.webdriver, plugins, languages spoofing
  • Resource blocking — block images, media, CSS, fonts for faster loading
  • Proxy support — HTTP/HTTPS proxy configuration
  • Multiple browsers — Chromium, Firefox, WebKit

📡 Web Monitoring & Alerting

  • Scheduled monitoring — configurable check intervals
  • Content change detection — text, HTML, specific element changes
  • Keyword monitoring — appearance, disappearance, count changes
  • Price monitoring — track price changes with threshold alerts
  • Change history — persistent history with diff generation
  • Multi-channel alerts — Webhook, Email (SMTP), DingTalk, WeCom
  • Configurable thresholds — minimum change size, similarity thresholds

⚡ Performance & Security

  • Smart caching — SQLite cache with TTL, size limits, automatic eviction
  • Rate limiting — per-domain token-bucket rate limiting
  • SSRF protection — blocks localhost, sensitive ports, invalid schemes
  • Browser fingerprint rotation — random User-Agents + realistic headers
  • TLS fingerprint simulation — realistic TLS ClientHello fingerprints
  • Connection pooling — persistent HTTP connections
  • Cookie management — automatic cookie handling and persistence

🚀 Easy Setup & Deployment

  • One-click setupwebscout-mcp setup auto-installs all dependencies
  • System detection — auto-detects OS, CPU, memory, GPU
  • Smart recommendations — suggests optimal configuration based on hardware
  • Docker support — pre-built images for amd64 and arm64
  • Docker Compose — one-command deployment
  • systemd service — Linux service file for production
  • Kubernetes — deployment manifests for container orchestration
  • Configuration hot-reload — reload config without restart

📊 Export & Integration

  • Multiple export formats — JSON, CSV, Markdown
  • MCP server — native Model Context Protocol support
  • CLI interface — command-line tools for search, fetch, crawl
  • Python API — full programmatic access to all features
  • Sitemap support — parse sitemap.xml and sitemap indexes
  • Incremental crawling — only re-fetch changed pages via ETag/Last-Modified

📦 Installation

Quick Install

pip install webscout-mcp

Requires Python 3.10+.

One-Click Full Setup (Recommended)

# Install core package
pip install webscout-mcp

# Run setup to install all optional dependencies
webscout-mcp setup --playwright --ollama --vector-store

The setup command will:

  • Detect your system configuration (OS, CPU, memory, GPU)
  • Install Playwright and Chromium browser
  • Install Ollama and download a local LLM (optional)
  • Install ChromaDB and sentence-transformers for vector search (optional)
  • Generate a configuration file
  • Run a health check to verify everything works

Optional Dependencies

# Browser automation (Playwright)
pip install webscout-mcp[browser]
playwright install chromium

# Vector search & RAG (ChromaDB + sentence-transformers)
pip install webscout-mcp[vector]

# AI content understanding (OpenAI client)
pip install webscout-mcp[ai]

# All features
pip install webscout-mcp[all]

Docker

docker pull wxslang/webscout-mcp:latest
docker run -p 8000:8000 wxslang/webscout-mcp:latest

🚀 Quick Start

MCP Client Configuration

Add to your MCP client config:

Claude Desktop (~/Library/Application Support/Claude/claude_desktop_config.json on macOS, %APPDATA%\Claude\claude_desktop_config.json on Windows):

{
  "mcpServers": {
    "webscout": {
      "command": "webscout-mcp",
      "args": []
    }
  }
}

Cursor (Settings → MCP → Add new MCP server):

{
  "mcpServers": {
    "webscout": {
      "command": "webscout-mcp",
      "args": []
    }
  }
}

CLI Usage

# Search the web
webscout-mcp search "python web scraping" --max-results 10

# Fetch a page
webscout-mcp fetch https://example.com --extract --format markdown

# Crawl a site
webscout-mcp crawl https://example.com --depth 2 --pages 50

# Run setup
webscout-mcp setup --playwright

# Start MCP server
webscout-mcp serve

Python API

from webscout_mcp import WebScout

# Initialize
scout = WebScout()

# Search
results = scout.search("AI agents", max_results=5)
for result in results:
    print(result.title, result.url)

# Fetch and extract
page = scout.fetch("https://example.com/article")
print(page.title)
print(page.content)  # Clean article text

# Crawl
pages = scout.crawl("https://example.com", depth=2, max_pages=20)

🤖 AI Content Understanding

from webscout_mcp.ai_processor import AIProcessor, AIConfig

# Use local Ollama (free, no API key needed)
config = AIConfig(backend="ollama", model="qwen2.5:7b")
ai = AIProcessor(config=config)

# Summarize
summary = ai.summarize(page.content, max_length=500)
print(summary.content)

# Ask questions
answer = ai.answer_question(page.content, "What are the main points?")
print(answer.content)

# Extract key points
points = ai.extract_key_points(page.content, num_points=5)
print(points.content)

# Analyze sentiment
sentiment = ai.analyze_sentiment(page.content)
print(sentiment.content)

Using OpenAI API:

config = AIConfig(
    backend="openai",
    model="gpt-4o",
    api_key="your-api-key",
)

Using Doubao (豆包):

config = AIConfig(
    backend="doubao",
    model="ep-20240101",
    api_key="your-api-key",
)

🧠 Vector Search & RAG

from webscout_mcp.vector_store import VectorStore, RAGEngine, Document

# Initialize vector store (local, free)
store = VectorStore()

# Add documents
doc = Document.from_text(page.content, source=page.url)
store.add_document(doc)

# Semantic search
results = store.search("how to build AI agents", n_results=5)
for result in results:
    print(f"[{result.score:.2f}] {result.document.content[:100]}")

# RAG: Ask questions based on your documents
rag = RAGEngine(vector_store=store)
answer = rag.query("What is the best approach for web scraping?")
print(answer["answer"])
print("Sources:", answer["sources"])

🌐 Headless Browser Automation

from webscout_mcp.browser_fetcher import BrowserFetcher, BrowserConfig

# Initialize
config = BrowserConfig(headless=True, block_media=True)
browser = BrowserFetcher(config=config)

# Fetch JS-rendered page
result = browser.fetch(
    "https://example.com/spa",
    wait_for_selector=".content",
    scroll_to_bottom=True,
)
print(result.title)
print(result.content)

# Take screenshot
browser.take_screenshot("https://example.com", "screenshot.png", full_page=True)

# Export to PDF
browser.export_pdf("https://example.com", "page.pdf")

# Click element
result = browser.click_element("https://example.com", "button.load-more")

# Fill form
result = browser.fill_form(
    "https://example.com/login",
    {"#username": "user", "#password": "pass"},
    submit_selector="button[type=submit]",
)

browser.close()

📡 Web Monitoring & Alerting

from webscout_mcp.monitor import WebMonitor, MonitorConfig, WebhookAlert, EmailAlert

# Initialize
config = MonitorConfig(check_interval=300, min_change_size=10)
monitor = WebMonitor(config=config)

# Add alert channels
monitor.add_alert_channel(WebhookAlert("https://hooks.example.com/alert"))
monitor.add_alert_channel(EmailAlert(
    smtp_server="smtp.gmail.com",
    smtp_port=587,
    username="you@gmail.com",
    password="app-password",
    from_addr="you@gmail.com",
    to_addrs=["recipient@example.com"],
))

# Check for changes
changes = monitor.check_url("https://example.com/pricing")
for change in changes:
    print(f"{change.change_type}: {change.old_value} -> {change.new_value}")

# Get history
history = monitor.get_history("https://example.com/pricing")

⚙️ Configuration

Environment Variables

# Core
WEBSCOUT_CACHE_ENABLED=true
WEBSCOUT_CACHE_TTL=3600
WEBSCOUT_RATE_LIMIT_ENABLED=true

# Search
WEBSCOUT_SEARCH_DEFAULT_BACKEND=bing
WEBSCOUT_SEARCH_MAX_RESULTS=10

# AI
WEBSCOUT_AI_BACKEND=ollama
WEBSCOUT_AI_MODEL=qwen2.5:7b
WEBSCOUT_AI_API_KEY=your-key

# Vector Store
WEBSCOUT_VECTOR_DB=chroma
WEBSCOUT_EMBEDDING_BACKEND=local
WEBSCOUT_EMBEDDING_MODEL=BAAI/bge-small-zh-v1.5

# Browser
WEBSCOUT_BROWSER_TYPE=chromium
WEBSCOUT_BROWSER_HEADLESS=true
WEBSCOUT_BROWSER_BLOCK_MEDIA=true

# Monitor
WEBSCOUT_MONITOR_INTERVAL=300
WEBSCOUT_MONITOR_MIN_CHANGE=10

# Setup
WEBSCOUT_SETUP_PLAYWRIGHT=true
WEBSCOUT_SETUP_OLLAMA=false
WEBSCOUT_SETUP_CHROMADB=false

Config File

Create ~/.config/webscout/config.toml:

[server]
host = "127.0.0.1"
port = 8000

[cache]
enabled = true
ttl = 3600

[search]
default_backend = "bing"
max_results = 10

[ai]
backend = "ollama"
model = "qwen2.5:7b"

[vector_store]
enabled = true
vector_db = "chroma"
embedding_backend = "local"

[browser]
enabled = true
headless = true

[monitor]
check_interval = 300

📚 Documentation

🧪 Testing

# Install dev dependencies
pip install webscout-mcp[dev]

# Run all tests
pytest tests/

# Run with coverage
pytest tests/ --cov=webscout_mcp --cov-report=html

Test coverage: 279+ tests covering all modules.

🤝 Contributing

Contributions are welcome! Please see CONTRIBUTING.md for guidelines.

  1. Fork the repository
  2. Create a feature branch
  3. Make your changes
  4. Add tests
  5. Submit a pull request

📄 License

MIT License — see LICENSE for details.

🙏 Acknowledgments


Built with ❤️ for the AI agent community.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

webscout_mcp-0.5.0.tar.gz (96.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

webscout_mcp-0.5.0-py3-none-any.whl (80.6 kB view details)

Uploaded Python 3

File details

Details for the file webscout_mcp-0.5.0.tar.gz.

File metadata

  • Download URL: webscout_mcp-0.5.0.tar.gz
  • Upload date:
  • Size: 96.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.10.12

File hashes

Hashes for webscout_mcp-0.5.0.tar.gz
Algorithm Hash digest
SHA256 1c9fbe3b8ae616ff3f05f95145f8f33a6cf5ea481efddf1d455cc4898a9baf5b
MD5 3b617e9aa59a1122aead005960b45f62
BLAKE2b-256 000e636a5e87b8b0bee4eab39aa02fe60c26b584829734505e86d4e2676481b8

See more details on using hashes here.

File details

Details for the file webscout_mcp-0.5.0-py3-none-any.whl.

File metadata

  • Download URL: webscout_mcp-0.5.0-py3-none-any.whl
  • Upload date:
  • Size: 80.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.10.12

File hashes

Hashes for webscout_mcp-0.5.0-py3-none-any.whl
Algorithm Hash digest
SHA256 d9480c7ace1d4ce455525972692049a02977820b3157c3ee20605ef5c7ca7f13
MD5 9b6ed30c51d32ede6f30744d413509ef
BLAKE2b-256 f588d22730554f3915af05afaf0ceeb38d458d81cab2ea1026c6bf8b4b54f954

See more details on using hashes here.

Release history Release notifications | RSS feed

0.6.4

2 files

0.6.2

2 files

0.6.1

2 files

0.6.0

2 files

This release

0.5.0 This release

2 files

0.4.0

2 files

0.3.0

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page