Skip to main content

docduty-search

The difference between a useful tool and just an API wrapper is the opinionated extraction — giving Claude clean, structured data specific to what each source actually contains, not generic markdown.

MCP server that unifies web search, academic papers, videos, news, patents, and GitHub issues/PRs behind a small set of composable tools — with optional Chrome browser integration for fetching JS-rendered pages, running JavaScript, parsing with Python, and taking screenshots. save() turns findings into durable, cited notes outside the ephemeral cache; blog() scaffolds an Astro-ready draft to write up what you found.

Install

uv tool install docduty-search            # search + details only
uv tool install "docduty-search[browser]" # + fetch, js, py, screenshot

Setup

docduty-search setup

One command does everything:

  1. Prompts for API keys → saves to ~/.config/docduty-search/.env
  2. Installs the research skill → ~/.claude/skills/docduty-search/SKILL.md
  3. Registers the MCP server with Claude Code (if claude CLI is found)

API keys:

  • SerpAPI key (required) — serpapi.com/manage-api-key
  • Perplexity key (optional) — console.perplexity.ai — enables domain-filtered search
  • GitHub token (optional) — setup auto-detects an authenticated gh CLI and uses gh auth token without prompting; only asks for a manual token if gh isn't available. Without either, falls back to unauthenticated GitHub API requests (60 req/h)

Tools

search

Search across 8 source types.

search("svelte 5 runes")                                      # web
search("machine learning", type="scholar")                     # papers with citations
search("react tutorial", type="video")                         # YouTube
search("AI regulation EU", type="news")                        # news
search("piezoelectric transducer", type="patents")             # patents
search("svelte logo", type="images")                           # images
search("electric vehicles", type="trends")                     # Google Trends
search("repo:sveltejs/svelte is:open label:bug", type="issues") # GitHub issues/PRs
Parameter Description
query Search query (required). For type="issues", raw GitHub search qualifiers (repo:, is:, label:, language:, ...)
type web · scholar · video · news · patents · images · trends · issues
domains Filter by domain. Prefix with - to exclude: ["-reddit.com"]
language ISO 639-1 code: "en", "no", "de"
recency hour · day · week · month · year
num Max results (default 10)

Patent-specific filters:

Parameter Description
before "priority:YYYYMMDD" or "publication:YYYYMMDD"
after Same format as before
inventor Filter by inventor name
assignee Filter by assignee/company
country Country codes: "US,WO,EP"
status "GRANT" or "APPLICATION"
sort "new" or "old" (by filing date)

Per-type extras in results:

Type Extra fields
scholar authors, cited_by, result_id, resources (PDF links), author_ids
video duration, channel, views, video_id
news source
patents patent_id, pdf, filing_date, inventor, assignee
issues state, is_pr, comments, repo, labels, created_at, updated_at

Every result also carries a ref — a short id valid for a few days, resolvable back to the exact cached response. Pass a ref to save()/blog() to embed that content permanently instead of just linking the URL. See Saving & writing up.

details

Get structured details for a result. Returns a summary inline; full data (claims, articles, transcript text) goes to the py() namespace.

details("patent/US11734097B1/en")          # patent claims, citations, similar
details("scholar/6497879044063343659")     # scholar article via Google Patents
details("t-NybWd6Sz0J")                   # citation formats (MLA, APA, BibTeX)
details("video/dQw4w9WgXcQ")              # video metadata + description
details("transcript/dQw4w9WgXcQ")          # full transcript text
details("author/nTJ7ihUAAAAJ")            # author profile, h-index, articles
ID format Inline response In py() namespace
patent/... title, abstract, dates, classifications, counts claims, patent_citations, cited_by, similar_documents
scholar/... same as patent same as patent
video/... title, channel, views, description_preview description
transcript/... chars, chapters full_text
author/... name, affiliations, h-index, articles_count articles, co_authors
result_id citation formats + export links

Also carries a ref (see above).

fetch (browser extra)

Navigate Chrome to a URL and extract content. Returns a summary with preview; full content available via py().

fetch("https://svelte.dev/docs/svelte/$state")
→ { ref, url, title, chars, content_preview }

py("content[:500]")    # full extracted markdown
py("soup.select('h2')") # BeautifulSoup on raw HTML

Tries the cheapest, cleanest extraction first, falling back to a full browser navigation only when nothing else matches:

  1. Patents — patent PDF URLs redirect to the Google Patents web page (always has text)
  2. Documents — PDF/DOCX/PPTX/XLSX via Content-Type detection and MarkItDown conversion
  3. GitHub — repos (README + metadata), file blobs (raw content), directory listings, and issue/PR threads (with comments) via the GitHub API — no browser needed
  4. Markdown mirrors — many docs sites (Mintlify, GitBook, Vercel docs, code.claude.com) serve raw markdown at <page>.md; Google's developer docs use <page>.md.txt. Both are tried before falling back to a full page render
  5. Full browser fallback — CDP navigate + trafilatura extraction, handles Cloudflare challenge pages automatically (waits for real Chrome to pass the challenge)

js (browser extra)

Run JavaScript on the current page.

js("document.title")
js("[...document.querySelectorAll('h2')].map(h => h.textContent)")

py (browser extra)

Run Python with access to page content and all previous results.

py("len(content)")
py("[h.text for h in soup.select('article h2')]")
py("cache.clear()")

Available in namespace:

  • content — clean extracted markdown from last fetch()
  • html — raw HTML from last fetch()
  • soup — BeautifulSoup parsed HTML
  • url — current page URL
  • _ / _1.._N — result history across all tools
  • js_result — last js() result
  • cache — TTL cache, call cache.clear() to reset
  • claims, cited_by, full_text, etc. — from last details() call

screenshot (browser extra)

Capture the current page as PNG.

screenshot()                    # saves to temp file
screenshot("/tmp/page.png")    # saves to specific path

Saving & writing up

Everything above is ephemeral — cached for a few days, then swept. Two tools turn a finding into something durable:

save

One call, one atomic fact — a durable, cited note outside the cache.

save(
    topic="safari no beforeinstallprompt",
    content="Safari never fires beforeinstallprompt...",
    sources=["53459432a9"],           # a ref from fetch()/search()/details() — solidifies the source
    tags=["pwa", "safari", "install"],
    questions=["Does Safari support beforeinstallprompt?"],
)
→ { path, solidified: [...], linked: [...] }

Citing a ref embeds that source's actual content permanently ("solidified"); a bare URL is just a link that can rot. Keep content scoped to one fact — every entry in questions should be answerable by it almost verbatim. code/language hold a verbatim snippet worth preserving separately from prose.

Notes land in ~/.local/share/docduty-search/memory/, one markdown file per topic, with a maintained INDEX.md. Check there (Glob/Grep, or read the memory://index resource for the path + full frontmatter schema) before researching something from scratch.

blog

Scaffolds a post — filename, Astro-shaped frontmatter, cited sources — then hands off to Write/Edit for the actual narrative, unlike save()'s one-shot fact.

blog(
    title="Teaching an agent to dig, remember, and write about it",
    intent="What this post should cover — seeds the Astro `description` field",
    sources=["910b6da9df"],
    tags=["docduty-search", "devlog"],
)
→ { path, solidified: [...], linked: [...] }

Frontmatter matches Astro's content-collections blog schema (title/pubDate/description/tags) so a finished draft needs no reformatting to be publish-ready — author/image are deliberately left out since they're project-identity details the tool has no basis to invent. Drafts land in ~/.local/share/docduty-search/blog/ (see blog://index for the schema).

Example: Patent Research

# Search with filters
search("MEMS transducer", type="patents", assignee="Murata", country="US,JP", status="GRANT")

# Get structured details — claims, citations, similar documents
details(_1[0]["patent_id"])
py("len(claims)")           # 55 claims
py("claims[0][:200]")       # first claim preview
py("len(cited_by)")         # 61 forward citations

# Find prior art (replicates Google Patents "Find Prior Art" button)
search(" ".join(_2["prior_art_keywords"]), type="patents",
       before=f"priority:{_2['prior_art_date'].replace('-','')}")

# Follow a citation chain
details(_2["cited_by"][0]["patent_id"])

Example: Video Learning

# Find tutorials
search("svelte 5 runes tutorial", type="video", num=5)

# Get video details
details(f"video/{_1[0]['video_id']}")

# Get the full transcript
details(f"transcript/{_1[0]['video_id']}")
py("full_text[:500]")    # first 500 chars of spoken content

Example: Scholar Research

# Search papers
search("piezoelectric micromachined ultrasonic transducer review", type="scholar")

# Get citation formats
details(_1[0]["result_id"])    # MLA, APA, Chicago, BibTeX

# Author profile
details(f"author/{_1[0]['author_ids'][0]}")
py("articles[:3]")             # top 3 papers
py("co_authors")               # collaborators

Backends

Backend Used when API key
SerpAPI Search types web/scholar/video/news/patents/images/trends + details SERPAPI_API_KEY (required)
Perplexity type="web" + domains set PERPLEXITY_API_KEY (optional)
GitHub API type="issues" search, and fetch() on github.com/raw.githubusercontent.com URLs GITHUB_TOKEN or gh auth token (optional — works unauthenticated at lower rate limits)
Chrome CDP fetch fallback, js, py, screenshot None (uses local Chrome)

Example: GitHub Research

# Search open bugs in a repo
search("repo:sveltejs/svelte is:open is:issue label:bug", type="issues", num=5)

# Read a repo's README + metadata
fetch("https://github.com/sveltejs/svelte")

# Read a specific source file (raw, no browser rendering)
fetch("https://github.com/sveltejs/svelte/blob/main/packages/svelte/src/index-client.js")

# Read a full issue thread, including comments
fetch(_1[0]["url"])

Development

git clone https://github.com/user/docduty-search
cd docduty-search
uv sync --extra browser
ruff format src/ && ruff check src/

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

docduty_search-0.1.6.tar.gz (34.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

docduty_search-0.1.6-py3-none-any.whl (42.5 kB view details)

Uploaded Python 3

File details

Details for the file docduty_search-0.1.6.tar.gz.

File metadata

  • Download URL: docduty_search-0.1.6.tar.gz
  • Upload date:
  • Size: 34.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.26 {"installer":{"name":"uv","version":"0.11.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"EndeavourOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for docduty_search-0.1.6.tar.gz
Algorithm Hash digest
SHA256 8a3969ed5cbddb74f2192f1b1c11a560a09f133390f073c09fe0494adf4aa914
MD5 92388d29cd61719d8f366f09b64913db
BLAKE2b-256 c4620ad77ca2860c2d6ea51f49b32b383c8d9a522b0107c0cd009aff5590fac2

See more details on using hashes here.

File details

Details for the file docduty_search-0.1.6-py3-none-any.whl.

File metadata

  • Download URL: docduty_search-0.1.6-py3-none-any.whl
  • Upload date:
  • Size: 42.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.26 {"installer":{"name":"uv","version":"0.11.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"EndeavourOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for docduty_search-0.1.6-py3-none-any.whl
Algorithm Hash digest
SHA256 b7c2c488b55b296a8b2dc49d4172eda051af45f9dc76d316fc5c1c2caab5bcf4
MD5 4cf27e67d82ada9e58d2cdf40971e96a
BLAKE2b-256 7d327f672b2cc6962fd456b7fd9cbff6b0018554aa58e5b5d4d471561a370bf0

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page