Skip to main content

arXiv MCP Server

MCP (Model Context Protocol) server for searching, downloading, and analyzing arXiv papers. Supports keyword search, Semantic Scholar-backed semantic search, PDF text extraction, and SQLite caching.

Features

  • Keyword Search — full arXiv API query syntax (ti:, au:, cat:, all:)
  • Semantic Search — server-side semantic search via the Semantic Scholar API
  • Hybrid Search — arXiv keyword results enriched with Semantic Scholar citation counts
  • PDF Download — download and extract full text via PyMuPDF
  • Smart Caching — SQLite cache with configurable TTL (24h metadata, 30-day PDFs)
  • Categories — built-in taxonomy of ~150 arXiv categories
  • Graceful Degradation — semantic/hybrid fall back to keyword when Semantic Scholar is unavailable
  • Lightweight — no local ML model; depends only on mcp and pymupdf

Install

pip install arxiv-mcp-server

No model download — the server starts instantly.

Claude Code Integration

Add to ~/.claude/settings.json or .claude/settings.local.json:

{
  "mcpServers": {
    "arxiv": {
      "command": "arxiv-mcp-server"
    }
  }
}

Or from a virtual environment:

{
  "mcpServers": {
    "arxiv": {
      "command": "python",
      "args": ["-m", "arxiv_mcp_server.server"]
    }
  }
}

Configuration

All via environment variables:

Variable Default Description
ARXIV_CACHE_DIR ~/.arxiv-mcp-server Cache root (DB + PDFs)
ARXIV_CACHE_TTL_HOURS 24 Metadata cache lifetime
ARXIV_PDF_TTL_DAYS 30 PDF retention period
ARXIV_RATE_LIMIT_SEC 1.0 Min seconds between API calls
ARXIV_REQUEST_TIMEOUT 30 API request timeout (seconds)
ARXIV_MAX_RESULTS 50 Max results per search
ARXIV_LOG_LEVEL INFO Python log level

Tools

search_papers

Search arXiv papers.

query: string (required) — search query
max_results: int (default 10, max 50)
start: int (default 0) — pagination offset
search_mode: "keyword" | "semantic" | "hybrid" (default "hybrid")

get_paper

Get a single paper by arXiv ID.

arxiv_id: string (required) — e.g. "1706.03762" or "1706.03762v7"

get_recent

Get recent papers in a category.

category: string (default "cs.AI")
max_results: int (default 20)

list_categories

List arXiv categories.

parent: string (optional) — filter by parent, e.g. "cs"

download_pdf

Download a paper PDF and extract text.

arxiv_id: string (required)
output_dir: string (optional) — defaults to ARXIV_CACHE_DIR/pdfs

get_cached_papers

List cached papers (diagnostic). No arguments.

Development

git clone <repo-url>
cd arxiv-mcp-server
pip install -e ".[dev]"
pytest

License

MIT

Metadata

Release files for arxiv-mcp-search 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for arxiv-mcp-search 0.1.0
File Size Uploaded
arxiv_mcp_search-0.1.0.tar.gz 17.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for arxiv-mcp-search 0.1.0
File Interpreter ABI Platform
arxiv_mcp_search-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 39.5 kB

Release files / arxiv_mcp_search-0.1.0.tar.gz

Download URL arxiv_mcp_search-0.1.0.tar.gz
Size 17.1 kB
Tags Source
SHA-256 checksum
How to use checksums
06a2e943769f118099d22cc531cab19d356266f19de1da0a05537bce4e5bd66b
BLAKE2b-256 checksum
How to use checksums
ee7d1fee692d42035215abf90f6bccdb5f53a670fccc1f5a6911b5db0d4396e1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.9

Release files / arxiv_mcp_search-0.1.0-py3-none-any.whl

Download URL arxiv_mcp_search-0.1.0-py3-none-any.whl
Size 22.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c9a22f18c8185d8c61c111c0a3db4fcc9acd45b949832a85586eff6f8e6b50a6
BLAKE2b-256 checksum
How to use checksums
4cb0b0110ddfad403c5941b39fe8275d5ef1b4a03c40229ac9d5d492c4763650
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.9

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page