wesearch🕸️
A wascally web research toolkit for agents; web search, web fetch, and publication research.
Quick Start
# Mac:
# # Required for quick install.
# brew install uv
# Ubuntu/Debian:
# # Required for quick install.
# sudo apt-get install -y curl
# curl -LsSf https://astral.sh/uv/install.sh | sh
uv add wesearch
# Alternatively: python -m pip install wesearch
wesearch is a synchronous, batteries-included library for programmatic web access: run a
search, fetch a page through a real-browser fingerprint, extract content, and look up scholarly
papers across multiple providers. It is the web layer factored out of a larger agent stack, so
it is built to survive bot-detection, rate limits, and flaky endpoints without a running browser
in the common case.
Example
from wesearch.search.search import search
from wesearch.fetch import RequestParams, fetch
from wesearch.scrape import get_element_content
# Web search (DuckDuckGo by default)
hits = search("denoising recursion models", num_results=10)
for r in hits:
print(r.title, r.url)
# Fetch + scrape
body, _session = fetch("https://example.com", request=RequestParams(timeout_sec=10))
title = get_element_content(body.decode("utf-8"), "h1")
# Scholarly papers: Semantic Scholar + OpenAlex, reciprocal-rank-fused by default
from wesearch.paper.search import search as paper_search
from wesearch.paper.ids import normalize_id
from wesearch.paper.details import metadata
result = paper_search("attention is all you need", limit=5)
for rec in result.records:
print(rec.title, rec.year)
meta = metadata(*normalize_id("arXiv:1706.03762"))
Each name is imported from the submodule that defines it; the top-level __init__ re-exports
nothing.
What's inside
wesearch/
├── types/ the vocabulary every layer shares; imports nothing internal
│ ├── params.py RequestParams + Content/Retry/Observe/Policy; Transport,
│ │ Extractor, Trust
│ ├── extractor.py the Extract protocol each extractor satisfies
│ └── errors.py FetchError, BotDetectionError + subclasses
├── fetch/ the sole HTTP egress
│ ├── fetch.py fetch(url, request=RequestParams(...)) -> (body, session)
│ ├── transport/ how bytes are retrieved
│ │ ├── curl.py curl-cffi transport (TLS/JA3 browser impersonation)
│ │ ├── stdlib.py dependency-free urllib transport
│ │ ├── zendriver.py opt-in real-Chrome backend for JS-gated pages
│ │ └── transport_routing.py per-domain transport selection
│ ├── extractor/ how a fetched page becomes text
│ │ ├── html2text.py every text node as Markdown
│ │ ├── markdownify.py the document's elements as Markdown
│ │ ├── trafilatura.py the scored article body only (the default)
│ │ └── raw.py the source, untouched
│ ├── providers/ per-site fetch strategies (reddit, google_news, x, ...)
│ └── challenge.py bot-challenge detection and classification
├── search/ web search over pluggable backends
│ └── search.py search(...) over SearXNG / DuckDuckGo / Google;
│ SearchResult / PaperResult / ImageResult
├── paper/ scholarly-paper lookup
│ ├── search.py search(...) across Semantic Scholar, OpenAlex, SearXNG
│ ├── details.py metadata / references / citations
│ ├── authors.py author search and publication lists
│ ├── fetch.py PDF download cascade
│ └── providers/ per-source backends (openalex, s2, searxng)
├── mcp/ the MCP surface; the only place the mcp SDK is imported
│ └── server.py wesearch-mcp, one tool per public function
├── chrome/ real-browser fingerprints
│ ├── headers.py Chrome request headers (incl. x-browser-validation)
│ └── useragents.py vendored, refreshable User-Agent pools
├── web.py fetch_web(...): fetch + provider dispatch + extraction
├── profile.py cross-process per-(egress_ip, domain) cookie + UA jar
├── ratelimit.py cross-process, per-domain rate limiting
└── scrape.py get_element_content(html, selector)
Transports
fetch picks a transport per domain:
- curl-cffi (default): TLS/JA3 browser impersonation, no browser process.
- stdlib: dependency-free
urllibfallback. - zendriver: an opt-in headless-Chrome backend for JavaScript-gated pages, used only for domains that require it.
A persistent per-(egress_ip, domain) profile (cookies + User-Agent) is loaded and saved
transparently, and cross-process rate limiting paces requests so concurrent workers stay under
each site's threshold.
API keys & configuration
Everything works keyless out of the box; the environment variables below raise your rate limits or unlock a backend, and are all optional unless noted.
SEMANTIC_SCHOLAR_API_KEY-- optional. Without it, Semantic Scholar lookups (paper.search,paper.details) share a low-rate public tier, andpaper.search(..., source="fused")may returncomplete=Falsewhen S2 throttles the request. Set it for a higher-throughput tier.OPENALEX_EMAIL-- optional. Identifies you to OpenAlex's "polite pool" for more headroom than anonymous requests.OPENALEX_API_KEY-- optional. A higher OpenAlex request budget.SEARXNG_URL-- required only to use a"searxng"backend (search(..., backend="searxng")orpaper.search(..., source="searxng")); the base URL of a SearXNG instance you control or trust.
State (the cookie/User-Agent profile jar, cross-process rate-limit lockfiles, the browser
transport's persistent Chrome profile) is written under the OS's standard per-user data
directory -- XDG_DATA_HOME (or ~/.local/share) on Linux, ~/Library/Application Support on
macOS, %LOCALAPPDATA% on Windows -- namespaced per component (see wesearch/lib/userdirs.py).
No configuration file is required or read.
The "zendriver" transport needs a non-snap Chrome or Chromium (e.g. Google Chrome's
.deb on x86_64). Snap-packaged Chromium fails every launch with BrowserUnavailableError
for two reasons, neither of which the library can work around:
- The snap wrapper takes several seconds to expose DevTools -- far beyond the launch budget
(0.5s per attempt, 6 attempts), which is sized so the
curl-then-zendrivercascade fails fast on hosts with no usable browser rather than stalling every fetch. - Snap's AppArmor confinement silently blocks writes under hidden home paths like
~/.local/share, so Chrome dies on its profile lock wherever the profile jar lands by default.
Development
uv sync --all-groups
Tests are tiered with pytest markers; the default run (uv run pytest) executes only the fast
unit tier:
addopts = -m 'not ci_smoke and not cuda and not integration and not performance and not cluster and not slow'
ci_smoke-- slower package smoke tests, run explicitly in CI.cuda-- requires a real CUDA device.integration-- requires networking or external CLIs.performance-- timing-sensitive.cluster-- requires live cluster access.slow-- expensive local correctness tests (JIT, full fixtures, git, bash, a fresh interpreter).real_llm-- spawns a live LLM CLI; skipped unlessRUN_REAL_LLM=1.
Run a specific tier with uv run pytest -m integration, or everything with
uv run pytest -m ''.
The zendriver backend tests need a Chrome or Chromium binary on PATH; without one they are
skipped automatically. No other system dependency is required to run the fast unit tier.
MCP server
The toolkit is directly callable by coding agents over MCP:
pip install 'wesearch[mcp]'
wesearch-mcp # serves stdio; register it with your MCP client
For Claude Code: claude mcp add --scope user wesearch -- wesearch-mcp.
Tools exposed: paper_search (fused Semantic Scholar + OpenAlex),
paper_details, paper_references, paper_citations, paper_pdf
(downloads into the user cache and returns the path), author_search,
author_papers, web_search, and web_fetch (extracted page text, with an
opt-in headless-browser fallback for bot-walled sites). Outputs are
deliberately compact for model consumption; abstracts are truncated and
empty fields dropped. The server is synchronous and per-client — state that
must be shared (rate limits, cookie/UA profiles) is already cross-process
safe on disk, so no daemon is needed.
See also
Sibling projects in the rekursiv-ai family:
- sagent — The self-mutating multi-provider coding-agent CLI and typed Python library.
- trackinizer — Centralized agent database for tracking inquiries, work, and the evidence behind conclusions.
- madcatter — Rich-based Markdown renderer for the terminal; ships the
mdcatCLI. - priml — Composable PyTorch building blocks: models, optimizers, losses, and a step-based training loop.
- configgle — Hierarchical experiment configuration in typed pure-Python dataclasses instead of YAML.
- copybarista — Bidirectional source sync for publishing OSS-ready trees from a monorepo.
- sudoku — Sudoku-Extreme solved end to end with a 7M-parameter recursive transformer.
Citing
If you find our work useful, please consider citing:
@misc{rekursivai2026wesearch,
title={Wesearch - A wascally web research toolkit for agents.},
author={Joshua V. Dillon},
year={2026},
howpublished={Github},
url={https://github.com/rekursiv-ai/wesearch},
}
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file wesearch-0.1.10.tar.gz.
File metadata
- Download URL: wesearch-0.1.10.tar.gz
- Upload date:
- Size: 413.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8099a47f03292a732f8e354875b8fd0147747908347709964c405279e0838f7f
|
|
| MD5 |
95301e89252c854b415f9250b25025f7
|
|
| BLAKE2b-256 |
a9a82ae3bbafc6c5b24d37fc222ba7060eaeecaa63dd13ef5bf5e5cfe4e0ab1a
|
Provenance
The following attestation bundles were made for wesearch-0.1.10.tar.gz:
Publisher:
publish-pypi.yml on rekursiv-ai/wesearch
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
wesearch-0.1.10.tar.gz -
Subject digest:
8099a47f03292a732f8e354875b8fd0147747908347709964c405279e0838f7f - Sigstore transparency entry: 2417659138
- Sigstore integration time:
-
Permalink:
rekursiv-ai/wesearch@ac1328c30e12b3500fe8f680c266bd97273a6c19 -
Branch / Tag:
refs/tags/v0.1.10 - Owner: https://github.com/rekursiv-ai
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish-pypi.yml@ac1328c30e12b3500fe8f680c266bd97273a6c19 -
Trigger Event:
release
-
Statement type:
File details
Details for the file wesearch-0.1.10-py3-none-any.whl.
File metadata
- Download URL: wesearch-0.1.10-py3-none-any.whl
- Upload date:
- Size: 222.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
6acd1119c5213584e66825af5a5362965ddb42a3c7f817f94c675d2e77233472
|
|
| MD5 |
c813660cf68d83a4379cd11ddbc8bada
|
|
| BLAKE2b-256 |
bfacab874021fff0eb918cf59e68ad9d034f4849cdb999c500ace8c79fe073c1
|
Provenance
The following attestation bundles were made for wesearch-0.1.10-py3-none-any.whl:
Publisher:
publish-pypi.yml on rekursiv-ai/wesearch
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
wesearch-0.1.10-py3-none-any.whl -
Subject digest:
6acd1119c5213584e66825af5a5362965ddb42a3c7f817f94c675d2e77233472 - Sigstore transparency entry: 2417659634
- Sigstore integration time:
-
Permalink:
rekursiv-ai/wesearch@ac1328c30e12b3500fe8f680c266bd97273a6c19 -
Branch / Tag:
refs/tags/v0.1.10 - Owner: https://github.com/rekursiv-ai
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish-pypi.yml@ac1328c30e12b3500fe8f680c266bd97273a6c19 -
Trigger Event:
release
-
Statement type: