Skip to main content

docseek

Goal-driven document discovery. Give it a website and a goal in plain language, in any language:

"Find the English datasheet of every battery charger" · "Töltsd le a 2026-os havi alap jelentéseket" · "Finde alle Sitzungsunterlagen vom März"

and it crawls the site with a real browser and returns the documents the goal asks for, each with a relevance verdict (accepted, unsure, rejected) and facets such as the period and year it covers.

It is built for sites where the documents are not one link away: behind year filters and tabs, in "load more" listings, in single-page apps that load them from a JSON API, one or two clicks below a listing, or spread over hundreds of pages of look-alike documents.

How it decides

The crawl is driven by a relevance judge that answers four questions, and code does everything else:

  1. Is this link a document the goal asks for? Every candidate gets a relevance score and a verdict.
  2. What kind of page does this link lead to? Pages are visited in order of how likely they list documents.
  3. Does this page still hide documents? If so, the page is escalated to a browsing agent that clicks, fills in and scrolls until it finds them.
  4. Which filter value shows the documents? Code finds the page's filters, the judge picks the value, code sets it.

Two judges are built in:

Judge What it is
llm Any LLM: Anthropic, or any OpenAI-compatible endpoint (OpenAI, Azure OpenAI, OpenRouter, Gemini, Ollama, vLLM, LM Studio). Answers on five levels mapped into the verdict bands.
jev TypeSafe Jev, a decision model that returns calibrated probabilities. Needs a TypeSafe key. If it becomes unavailable mid-crawl, the LLM judge takes over.

The judge's wording comes from a domain profile (profile): generic works for any goal; fund-reports is an example of a profile tuned for one domain (periodic fund documents). A profile is a small JSON file - see docseek/profiles/ - so you can write one for your own domain.

The browsing agent (escalations, and the agent decision layer that crawls with the agent alone) runs on the same LLM configuration.

Quick start

pip install docseek
playwright install chromium

or from source:

git clone https://github.com/mohojojo/docseek && cd docseek
uv venv && source .venv/bin/activate      # or python -m venv .venv
uv pip install -e ".[dev]"
playwright install chromium

Then start the API:

export ANTHROPIC_API_KEY=...              # or configure any other model, see below
uvicorn docseek.server:app --port 8010
curl -s localhost:8010/v1/discover -H 'Content-Type: application/json' -d '{
  "url": "https://www.example.com/",
  "goal": "Find the 2025 annual reports (PDF)",
  "profile": "generic",
  "max_pages": 30
}'

Or with Docker: docker compose up --build.

Or from the command line (the same crawl, printed as JSON):

docseek https://www.example.com/ "Find the 2025 annual reports (PDF)" --max-pages 30

Any model

# OpenAI
export LLM_PROVIDER=openai-compatible LLM_MODEL=gpt-4.1-mini LLM_API_KEY=sk-...

# a local model through Ollama
export LLM_PROVIDER=openai-compatible LLM_BASE_URL=http://localhost:11434/v1 LLM_MODEL=llama3.1

# Anthropic (the default when ANTHROPIC_API_KEY is set)
export LLM_PROVIDER=anthropic LLM_MODEL=claude-haiku-4-5

The client adapts to what a server accepts (for example OpenAI's reasoning models, which take max_completion_tokens and no temperature). The browsing agent needs a model with tool calling; screenshots need a vision-capable model.

Generated programs

For a site you query again and again, docseek can write a discovery program: a coding agent explores the site once - raw HTML, rendered pages, the JSON calls a page makes - and writes a small Python function that finds the documents the goal asks for. Later requests run the program with no model at all, in seconds, and the relevance judge scores what it returns exactly as it scores a crawl.

export PROGRAMS_DIR=./programs CODEGEN_MODEL=claude-sonnet-5
docseek generate https://www.example.com/ "Find the 2025 annual reports (PDF)"   # a few minutes, once
docseek https://www.example.com/ "Find the 2025 annual reports (PDF)" --programs  # seconds, no model
docseek check                                                                      # has any site changed?

With "programs": true, /v1/discover answers from the site's program when it is healthy, and crawls otherwise: when the program errors, returns nothing where it used to find documents, or keeps fewer than half of what it kept last time, it is marked stale, the crawl answers, and a new program is written in the background. docseek check (or POST /v1/programs/check) replays every program against a snapshot of what it returned before - no model, no judge - and reports ok, grew, shrank or broken.

What a program can do is fenced in: it runs in a separate process with an empty environment, CPU and memory limits, an allowlist of standard-library modules, and no network except the parent's fetcher - the site and its subdomains, the data hosts and POST endpoints the site's own pages used while the program was written, robots.txt, and the same public-address rule as the crawl. This is a boundary, not a hardened jail: the code was written by a model that read untrusted page text, so the feature is off unless PROGRAMS_DIR is set, and a docseek that strangers can reach belongs in a container.

Health cannot tell when a program confidently returns the wrong slice of a site; a crawl or a person can.

Configuration

Variable Description
LLM_PROVIDER anthropic or openai-compatible. Default: anthropic when ANTHROPIC_API_KEY is set.
LLM_MODEL Model for the agent and the LLM judge. Default on anthropic: claude-haiku-4-5; required for openai-compatible.
LLM_BASE_URL Base URL of an OpenAI-compatible endpoint (default https://api.openai.com/v1).
LLM_API_KEY Key for the LLM (falls back to ANTHROPIC_API_KEY / OPENAI_API_KEY; a local server may need none).
ANTHROPIC_API_KEY Enough on its own to run everything on Claude. /v1/search-sites uses Anthropic's web search and always needs it.
TYPESAFE_API_KEY Enables the jev judge (the default judge when set).
JEV_PRICE_PER_MTOK Your TypeSafe price per million input tokens, used only for the cost the result reports (jev_cost_usd). Default 0.
CRAWLER_API_KEY When set, every request must send it as X-API-Key. Unset, the API is open - set it before exposing the service.
PROGRAMS_DIR Directory for generated discovery programs. Unset: the feature is off.
CODEGEN_MODEL Model that writes programs, on the configured provider (default: LLM_MODEL). Use a strong coding model.
CODEGEN_MAX_TURNS, CODEGEN_MAX_INPUT_TOKENS Budget for writing one program (default 45 turns, 3M input tokens including cached reads).
PATTERNS_DIR Directory where learned site knowledge (gate sequences, replayable escalation steps) is kept. Unset: nothing is learned.

HTTP API

Endpoint
POST /v1/discover Crawl and return the result as JSON.
POST /v1/discover-stream The same crawl as a server-sent event stream (every page, verdict and escalation as it happens).
GET /v1/profiles The bundled domain profiles.
POST /v1/search-sites Suggest websites for a goal (Anthropic web search).
GET/DELETE /v1/patterns[/{domain}] Learned site knowledge.
POST /v1/programs Write (or rewrite) a site's discovery program in the background.
GET/DELETE /v1/programs[/{key}] Generated programs: code, notes, health.
POST /v1/programs/check Replay programs against their snapshots (drift check).
GET /health Liveness.

Main request fields for /v1/discover:

Field Default
url, goal Where to start and what to find.
decision_layer judge-driven jev for the judge-driven crawl (despite the name, with any judge), agent for the browsing agent alone.
judge jev if its key is set, else llm Which relevance judge.
profile generic Domain profile name, or a path to a profile file.
max_pages, max_seconds, max_depth 10, 180, 3 Crawl budget.
same_domain_only, allowed_hosts true, [] Off-domain policy: with same_domain_only: false the crawl may cross to one host linked from the start site.
programs false Answer from the site's generated program when it is healthy (needs PROGRAMS_DIR).
include_rejected false Also return rejected candidates, to see what the judge threw away.
model LLM_MODEL Agent model override.

The result lists the documents with relevance, verdict, source (page, sitemap, api, agent), period and year, plus stop_reason, token counts and which model decided.

Crawling responsibly

  • robots.txt is honoured for every host the crawl touches.
  • SSRF protection: only http(s) URLs whose host resolves exclusively to public addresses are fetched or returned - a hostname pointing at 127.0.0.1 or a cloud metadata address is refused. The same rule applies to URLs a page or a model hands the crawler.
  • Pages are visited one at a time by default (max_concurrent); set max_pages and max_seconds to what the site can take.
  • Page text reaches the models as data. Instruction-like text in a page is flagged and never acted on.
  • Crawl only what you are allowed to, and respect each site's terms.

Evaluating

eval/ measures recall and precision against hand-checked ground truth: see eval/ground_truth/README.md for the format and the scripts (run_jev.py for the judge-driven crawl, run_baseline.py for the agent path, run_codegen.py for generated programs, judge_compare.py for scoring a judge offline on a frozen, labelled candidate set).

Development

uv pip install -e ".[dev]"
playwright install chromium
pytest

The tests run offline: DNS resolution and robots.txt are faked (tests/conftest.py) and every model call is mocked. Browser tests use a local Chromium and are skipped when it is missing.

Licence

Apache-2.0 - see LICENSE. TypeSafe Jev is a third-party service with its own terms; its API client is included, its model is not.

Release files for docseek 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for docseek 0.1.0
File Size Uploaded
docseek-0.1.0.tar.gz 214.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for docseek 0.1.0
File Interpreter ABI Platform
docseek-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 344.4 kB

Release files / docseek-0.1.0.tar.gz

Download URL docseek-0.1.0.tar.gz
Size 214.0 kB
Tags Source
SHA-256 checksum
How to use checksums
0ebd4473dfbe33827b69f0c23a7ca10619cfce9add7a482738f277d9fd94fc0c
BLAKE2b-256 checksum
How to use checksums
77af2537e285a83ed3b5516529393956c6796235e86292b928cc26ee1d82111a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.

Transparency log

Release files / docseek-0.1.0-py3-none-any.whl

Download URL docseek-0.1.0-py3-none-any.whl
Size 130.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
112eaf120e89c3458725b8d1de0d230560c87a69dffed8420a77ba02075acf1f
BLAKE2b-256 checksum
How to use checksums
6c90179285603717bb094565717e58e34f12e7d4a9811ab85aba8ce85be71b8d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.

Transparency log

Release history Release notifications | RSS feed

0.7.1

2 release files

0.7.0

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page