Skip to main content

wesearch🕸️

PyPI version CI Python 3.12+ License: Apache-2.0 Discord

A wascally web research toolkit for agents; web search, web fetch, and publication research.

Quick Start

# Mac:
#   # Required for quick install.
#   brew install uv

# Ubuntu/Debian:
#   # Required for quick install.
#   sudo apt-get install -y curl
#   curl -LsSf https://astral.sh/uv/install.sh | sh

uv add wesearch

# Alternatively: python -m pip install wesearch

wesearch is a synchronous, batteries-included library for programmatic web access: run a search, fetch a page through a real-browser fingerprint, extract content, and look up scholarly papers across multiple providers. It is the web layer factored out of a larger agent stack, so it is built to survive bot-detection, rate limits, and flaky endpoints without a running browser in the common case.

Example

from wesearch.search import search
from wesearch.fetch import RequestParams, fetch
from wesearch.scrape import get_element_content

# Web search (DuckDuckGo by default)
hits = search("denoising recursion models", num_results=10)
for r in hits:
    print(r.title, r.url)

# Fetch + scrape
body, _session = fetch("https://example.com", request=RequestParams(timeout_sec=10))
title = get_element_content(body.decode("utf-8"), "h1")

# Scholarly papers: Semantic Scholar + OpenAlex, reciprocal-rank-fused by default
from wesearch.paper.search import search as paper_search
from wesearch.paper.ids import normalize_id
from wesearch.paper.details import metadata

result = paper_search("attention is all you need", limit=5)
for rec in result.records:
    print(rec.title, rec.year)

meta = metadata(*normalize_id("arXiv:1706.03762"))

Each name is imported from the submodule that defines it; the top-level __init__ re-exports nothing.

What's inside

wesearch/
├── search.py        search(...) over DuckDuckGo (+ configured backends);
│                      SearchResult / PaperResult / ImageResult
├── fetch/           the sole HTTP egress
│   ├── fetch.py     fetch(url, request=RequestParams(...)) -> (body, session)
│   ├── curl.py      curl-cffi transport (TLS/JA3 browser impersonation)
│   ├── stdlib.py    dependency-free urllib transport
│   ├── zendriver.py opt-in real-Chrome backend for JS-gated pages
│   ├── transport_routing.py  per-domain transport selection
│   └── challenge.py bot-challenge detection and classification
├── chrome/          real-browser fingerprints
│   ├── headers.py   Chrome request headers (incl. x-browser-validation)
│   └── useragents.py  vendored, refreshable User-Agent pools
├── profile.py       cross-process per-(egress_ip, domain) cookie + UA jar
├── ratelimit.py     cross-process, per-domain rate limiting
├── scrape.py        get_element_content(html, selector)
├── errors.py        FetchError, BotDetectionError + subclasses
└── paper/           scholarly-paper lookup
    ├── search.py    search(...) across Semantic Scholar, OpenAlex, SearXNG
    ├── details.py   metadata / references / citations
    ├── authors.py   author search and publication lists
    ├── fetch.py     PDF download cascade
    └── providers/   per-source backends (openalex, s2, searxng)

Transports

fetch picks a transport per domain:

  • curl-cffi (default): TLS/JA3 browser impersonation, no browser process.
  • stdlib: dependency-free urllib fallback.
  • zendriver: an opt-in headless-Chrome backend for JavaScript-gated pages, used only for domains that require it.

A persistent per-(egress_ip, domain) profile (cookies + User-Agent) is loaded and saved transparently, and cross-process rate limiting paces requests so concurrent workers stay under each site's threshold.

API keys & configuration

Everything works keyless out of the box; the environment variables below raise your rate limits or unlock a backend, and are all optional unless noted.

  • SEMANTIC_SCHOLAR_API_KEY -- optional. Without it, Semantic Scholar lookups (paper.search, paper.details) share a low-rate public tier, and paper.search(..., source="fused") may return complete=False when S2 throttles the request. Set it for a higher-throughput tier.
  • OPENALEX_EMAIL -- optional. Identifies you to OpenAlex's "polite pool" for more headroom than anonymous requests.
  • OPENALEX_API_KEY -- optional. A higher OpenAlex request budget.
  • SEARXNG_URL -- required only to use a "searxng" backend (search(..., backend="searxng") or paper.search(..., source="searxng")); the base URL of a SearXNG instance you control or trust.
  • WESEARCH_BROWSER_CONNECTION_TIMEOUT_SEC -- optional; only affects the "zendriver" browser transport. Seconds to wait for Chrome's DevTools channel per launch attempt (default 1.0, which fails fast when no usable browser exists so the curl-then-zendriver cascade stays snappy). Wrapper-launched browsers need much longer -- see the snap note below.

State (the cookie/User-Agent profile jar, cross-process rate-limit lockfiles, the browser transport's persistent Chrome profile) is written under the OS's standard per-user data directory -- XDG_DATA_HOME (or ~/.local/share) on Linux, ~/Library/Application Support on macOS, %LOCALAPPDATA% on Windows -- namespaced per component (see wesearch/lib/userdirs.py). No configuration file is required or read.

Snap-packaged Chromium (stock Ubuntu) needs two accommodations for the "zendriver" transport, or it fails every launch with BrowserUnavailableError:

  1. WESEARCH_BROWSER_CONNECTION_TIMEOUT_SEC=30 -- the snap wrapper takes several seconds to expose DevTools, far beyond the 1-second default, and each retry restarts the browser.
  2. XDG_DATA_HOME pointing at a non-hidden directory (e.g. ~/wesearch-data) -- snap's AppArmor confinement silently blocks writes under hidden home paths like ~/.local/share, so Chrome dies on its profile lock at every launch when the profile jar lives in the default location.

A non-snap Chrome or Chromium (e.g. Google Chrome's .deb on x86_64) needs neither.

Development

uv sync --all-groups

Tests are tiered with pytest markers; the default run (uv run pytest) executes only the fast unit tier:

addopts = -m 'not ci_smoke and not cuda and not integration and not performance and not cluster and not slow'
  • ci_smoke -- slower package smoke tests, run explicitly in CI.
  • cuda -- requires a real CUDA device.
  • integration -- requires networking or external CLIs.
  • performance -- timing-sensitive.
  • cluster -- requires live cluster access.
  • slow -- expensive local correctness tests (JIT, full fixtures, git, bash, a fresh interpreter).
  • real_llm -- spawns a live LLM CLI; skipped unless RUN_REAL_LLM=1.

Run a specific tier with uv run pytest -m integration, or everything with uv run pytest -m ''.

The zendriver backend tests need a Chrome or Chromium binary on PATH; without one they are skipped automatically. No other system dependency is required to run the fast unit tier.

MCP server

The toolkit is directly callable by coding agents over MCP:

pip install 'wesearch[mcp]'
wesearch-mcp                  # serves stdio; register it with your MCP client

For Claude Code: claude mcp add --scope user wesearch -- wesearch-mcp.

Tools exposed: paper_search (fused Semantic Scholar + OpenAlex), paper_details, paper_references, paper_citations, paper_pdf (downloads into the user cache and returns the path), author_search, author_papers, web_search, and web_fetch (extracted page text, with an opt-in headless-browser fallback for bot-walled sites). Outputs are deliberately compact for model consumption; abstracts are truncated and empty fields dropped. The server is synchronous and per-client — state that must be shared (rate limits, cookie/UA profiles) is already cross-process safe on disk, so no daemon is needed.

See also

Sibling libraries in the rekursiv-ai family:

  • sagent — The self-mutating multi-provider coding-agent CLI and typed Python library.
  • trackinizer — Centralized agent database for tracking inquiries, work, and the evidence behind conclusions.
  • madcatter — Rich-based Markdown renderer for the terminal; ships the mdcat CLI.
  • priml — Composable PyTorch building blocks: models, optimizers, losses, and a step-based training loop.
  • configgle — Hierarchical experiment configuration in typed pure-Python dataclasses instead of YAML.

Citing

If you find our work useful, please consider citing:

@misc{rekursivai2026wesearch,
      title={Wesearch - A wascally web research toolkit for agents.},
      author={Joshua V. Dillon},
      year={2026},
      howpublished={Github},
      url={https://github.com/rekursiv-ai/wesearch},
}

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

wesearch-0.1.8.tar.gz (345.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

wesearch-0.1.8-py3-none-any.whl (186.0 kB view details)

Uploaded Python 3

File details

Details for the file wesearch-0.1.8.tar.gz.

File metadata

  • Download URL: wesearch-0.1.8.tar.gz
  • Upload date:
  • Size: 345.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for wesearch-0.1.8.tar.gz
Algorithm Hash digest
SHA256 f8ae8b3e3cb191e215b67c4b69dab6418b5d0754482d8f852571299456c6b126
MD5 d6d4e23b105557c89e1c32ba0210eef4
BLAKE2b-256 5095cbd4f1a4ed44b830a18a90eb36dcfdcad47fd9c9f64dc8e5b14af361e041

See more details on using hashes here.

Provenance

The following attestation bundles were made for wesearch-0.1.8.tar.gz:

Publisher: publish-pypi.yml on rekursiv-ai/wesearch

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file wesearch-0.1.8-py3-none-any.whl.

File metadata

  • Download URL: wesearch-0.1.8-py3-none-any.whl
  • Upload date:
  • Size: 186.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for wesearch-0.1.8-py3-none-any.whl
Algorithm Hash digest
SHA256 00a3c18c0ea1627a1d8893ba619e6132ef93c83ec7a6780e2ffa0beb9b05648a
MD5 abe4561a2701c764294c27f83868ea52
BLAKE2b-256 9d96397504a325c54ee392872c799458b5663372269bb0b3aa00973d92fa83d2

See more details on using hashes here.

Provenance

The following attestation bundles were made for wesearch-0.1.8-py3-none-any.whl:

Publisher: publish-pypi.yml on rekursiv-ai/wesearch

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.1.11

2 files

0.1.10

2 files

0.1.9

2 files

This release

0.1.8 This release

2 files

0.1.7

2 files

0.1.6

2 files

0.1.5

2 files

0.1.4

2 files

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page