Skip to main content

Scientific knowledge base — papers, packages, codebases → queryable markdown

Project description

PaperMind

Scientific knowledge base: papers, packages, and codebases → queryable markdown.

PaperMind ingests heterogeneous scientific sources — PDFs, PyPI packages, and source trees — into a portable, plain-text knowledge base. A CLI manages ingestion, search, and discovery. An MCP server exposes the KB as tools to any AI assistant that speaks the Model Context Protocol.

Install

Minimum (no PDF or browser support):

pip install papermind

With PDF ingestion (GLM-OCR — requires GPU):

pip install "papermind[ocr]"

Note: GLM-OCR requires a recent transformers build with GLM-OCR support. If pip install "papermind[ocr]" gives a model loading error, install the dev branch: pip install "transformers @ git+https://github.com/huggingface/transformers.git"

With semantic search (qmd):

npm install -g @tobilu/qmd

With browser-based package docs:

pip install "papermind[browser]"
playwright install chromium

Requirements: Python 3.11+, GPU recommended for PDF ingestion

Quick Start

# 1. Create a knowledge base
papermind --kb ~/kb init

# 2. Fetch papers (search + download + OCR + ingest in one step)
papermind --kb ~/kb fetch "SWAT+ calibration machine learning" -n 10 -t swat_ml

# 3. Ingest a local PDF
papermind --kb ~/kb ingest paper path/to/paper.pdf --topic hydrology

# 4. Ingest a Python package's API docs
papermind --kb ~/kb ingest package numpy

# 5. Ingest a codebase (Python, Fortran, C, Rust)
papermind --kb ~/kb ingest codebase ~/src/myproject --name myproject

# 6. Search
papermind --kb ~/kb search "evapotranspiration calibration"
papermind --kb ~/kb search "SWAT" --topic swat_ml

# 7. Check what's in the KB
papermind --kb ~/kb catalog show

CLI Reference

All commands take --kb <path> as a global option. Pass --offline to disable all network access.

Command Description
init Initialize a new knowledge base directory
fetch <query> Search + download + OCR + ingest papers in one step
ingest paper <path> Add a paper (PDF) via GLM-OCR
ingest package <name> Extract a PyPI package's API and docs
ingest codebase <path> Walk a source tree (Python, Fortran, C)
search <query> Search the KB (semantic via qmd, or grep fallback)
catalog show List all KB entries (--json for machine-readable)
catalog stats Summary statistics by type and topic
remove <id> Remove an entry from the KB
discover <query> Find papers via OpenAlex / Semantic Scholar / Exa
download <url|doi> Download a paper PDF
export-bibtex Export paper citations as BibTeX
doctor Check installed dependencies and tool availability
reindex Rebuild catalog.json and catalog.md from filesystem
serve Start the MCP server (stdio transport)
version Print version

Examples

# Fetch 10 papers on a topic, auto-download and ingest
papermind --kb ~/kb fetch "differentiable hydrology neural ODE" -n 10 -t diff_hydro

# Preview what fetch would do (no download/ingest)
papermind --kb ~/kb fetch "SWAT calibration" -n 5 --dry-run

# Ingest multiple papers from a directory
papermind --kb ~/kb ingest paper papers/ --topic swat

# Export citations for reference managers
papermind --kb ~/kb export-bibtex > references.bib

# Machine-readable catalog
papermind --kb ~/kb catalog show --json

# Search with topic filter
papermind --kb ~/kb search "calibration" --topic swat_ml

# Run fully offline (no network calls at all)
papermind --kb ~/kb --offline search "groundwater recharge"

# Check tool health
papermind --kb ~/kb doctor

Paper Discovery

PaperMind searches three academic APIs in parallel:

  • OpenAlex — free, no API key, direct PDF URLs for open-access papers
  • Semantic Scholar — structured metadata, citation counts (optional API key for higher rate limits)
  • Exa — broad web search (requires API key)
  • Unpaywall — DOI→PDF resolver fallback (free, no key)

PDF OCR

PaperMind uses GLM-OCR (MIT, 0.9B params, #1 OmniDocBench) for PDF→markdown conversion. Features:

  • Runs locally on GPU (RTX 3060+ recommended, ~2GB VRAM)
  • Outputs structured markdown with LaTeX equations
  • Auto-detects section headings (numbered sections, ALL-CAPS)
  • Extracts embedded figures as PNG files alongside the markdown
  • Source PDF copied next to markdown for easy comparison

Install with pip install "papermind[ocr]". Model downloaded from HuggingFace on first use (~2GB, cached).

MCP Server

PaperMind exposes your KB to AI assistants via the Model Context Protocol.

Claude Code (.claude/mcp.json):

{
  "mcpServers": {
    "papermind": {
      "command": "papermind",
      "args": ["--kb", "/path/to/kb", "serve"]
    }
  }
}

Available MCP tools:

Tool Description
query Search the KB; optional scope, topic, limit
get Read a single document by relative path
multi_get Read multiple documents in one call
catalog_stats KB statistics (counts by type and topic)
list_topics All topics in the KB
discover_papers Search academic APIs

Search

Two search backends:

  • qmd — hybrid search (BM25 + vector embeddings + LLM reranking). Install: npm install -g @tobilu/qmd, then qmd collection add ~/kb --name my-kb
  • Built-in fallback — grep-based term matching (zero dependencies)

Configuration

Each KB has a .papermind/config.toml. All keys are optional.

[search]
qmd_path = "qmd"
fallback_search = true

[apis]
semantic_scholar_key = ""
exa_key = ""

[ingestion]
ocr_model = "zai-org/GLM-OCR"
ocr_dpi = 150
default_paper_topic = "uncategorized"

[firecrawl]
api_key = ""

[privacy]
offline_only = false

Environment variables override config file values:

Variable Purpose
PAPERMIND_EXA_KEY Exa search API key
PAPERMIND_SEMANTIC_SCHOLAR_KEY Semantic Scholar API key
PAPERMIND_FIRECRAWL_KEY Firecrawl API key
HF_TOKEN HuggingFace token (faster model downloads)

Contributing

git clone https://github.com/dmbrmv/papermind
cd papermind
pip install -e ".[dev]"
uv run pytest tests/ -v
uv run ruff check src/

The test suite is fully offline — no network calls, no external tools required.

License

MIT — see LICENSE.

Third-party dependency licenses: LICENSE_THIRD_PARTY.md.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

papermind-1.5.0.tar.gz (212.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

papermind-1.5.0-py3-none-any.whl (90.8 kB view details)

Uploaded Python 3

File details

Details for the file papermind-1.5.0.tar.gz.

File metadata

  • Download URL: papermind-1.5.0.tar.gz
  • Upload date:
  • Size: 212.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for papermind-1.5.0.tar.gz
Algorithm Hash digest
SHA256 7e86ad5315ddb3b7adf0a3824e541f2fb01a83043b8ceeba83b6add4fb1368ea
MD5 3adfc3a29267ac958cc486043f8a9122
BLAKE2b-256 f60ec340dafd8cf212ad2ec1456157332440feefae7b57d5c497ce4e994aa621

See more details on using hashes here.

Provenance

The following attestation bundles were made for papermind-1.5.0.tar.gz:

Publisher: publish.yml on dmbrmv/papermind

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file papermind-1.5.0-py3-none-any.whl.

File metadata

  • Download URL: papermind-1.5.0-py3-none-any.whl
  • Upload date:
  • Size: 90.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for papermind-1.5.0-py3-none-any.whl
Algorithm Hash digest
SHA256 f7cc6c32560d78cf9b79d57106053782069029966a2a88acdf877dfb36155646
MD5 8323808bf6a65cde9fa0570bc2022a4b
BLAKE2b-256 b102a7fad697390d2a8f691109f33bd495d3bc199090a191d1841dc30bf4c637

See more details on using hashes here.

Provenance

The following attestation bundles were made for papermind-1.5.0-py3-none-any.whl:

Publisher: publish.yml on dmbrmv/papermind

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page