Skip to main content

asta-papers

Legal-only paper full-text retrieval and conversion. Identifier (DOI / PMID / PMCID / arXiv / Semantic Scholar Corpus ID) — or BYO PDF / JATS — to markdown with explicit license classification.

Lifts paper recovery on biomedical literature using only publisher-blessed legal channels (NCBI E-utilities, EuropePMC, bioRxiv, arXiv). It does not query Unpaywall or scrape institutional repositories.

It also installs the asta-papers command: Semantic Scholar search, paper lookup and citation walks, whose JSON output is what downstream tools such as theorize papers fetch read.

export S2_API_KEY=...        # optional; 1 request/s without, 100 with
asta-papers search "astrocyte lactate shuttle" --limit 20 --date 2018-
asta-papers get ARXIV:2005.14165 --format text
asta-papers citations CorpusId:218487638 --limit 50
asta-papers search "..." --limit 30 | asta-papers store --into store.json    # fetch the text, build the store

store fetches the full text of what a search found and writes the paper store Theorizer reads, resumably. PDFs go through Mistral OCR when MISTRAL_API_KEY is set, otherwise through pymupdf4llm (pip install 'asta-papers[pymupdf]').

Install

pip install asta-papers                 # core (JATS conversion only)
pip install 'asta-papers[mistral]'      # + Mistral OCR for PDFs
pip install 'asta-papers[olmocr]'       # + local olmOCR for PDFs (offline)
pip install 'asta-papers[s3]'           # + s3:// BYO support
pip install 'asta-papers[all]'          # everything

Quickstart

import os
from asta_papers import Client
from asta_papers.converters.mistral import MistralConverter

c = Client(
    email="me@allenai.org",
    ncbi_api_key=os.environ.get("NCBI_API_KEY"),       # optional, lifts NCBI 3→10 rps
    converters=[MistralConverter()],                    # for PDF→markdown
)

# By identifier
r = c.fetch(doi="10.1186/s12943-024-02093-w")
print(r.success, r.license_class, r.markdown[:200])

# Storage-tier policy
if r.may_redistribute:                                   # CC BY / CC0 / CC BY-SA
    save_artifact(r.bytes)
elif r.may_use_for_tdm:                                  # TDM-permissive licenses; not BRONZE/UNKNOWN/CLOSED
    extract_inline(r.markdown)

# BYO PDF — bytes, local path, or URI
r = c.fetch(pdf=b"%PDF-...")
r = c.fetch(pdf="paper.pdf")
r = c.fetch(pdf="s3://my-bucket/paper.pdf")              # requires [s3]

# Batch with bounded concurrency + per-host rate limits
results = c.fetch_many([
    {"doi": "10.1038/foo"},
    {"pmcid": "PMC123"},
    {"pdf": "paper.pdf", "doi": "10.99/local"},
])

How it works

A strategy ladder runs against legal aggregator APIs in order, returning the first successful result:

  1. PMC E-utilities efetch — JATS XML for PMC OA Subset articles
  2. NCBI elink — PMID → PMC self-link when S2 didn't surface it
  3. Published-version handoff — when input is a preprint DOI, route to the published version (bioRxiv API or Crossref relation) so callers get the most-recent public version of the paper
  4. arXiv — direct PDF for arXiv papers
  5. bioRxiv / medRxiv API — JATS XML or PDF for preprints
  6. EuropePMC PDF render — text-mining-licensed PDFs for free-to-read papers NCBI's OA Subset doesn't include

Per-host token-bucket rate limiting honors every publisher's published quota exactly. arXiv at 0.33 rps (their explicit rule). NCBI at 3 rps (10 with key) shared across eutils.*, pmc.*, www.* hostnames. Independent services run fully in parallel.

License classification

Every successful fetch carries a LicenseClass:

cc-by  cc-by-sa  cc-by-nd  cc-by-nc  cc-by-nc-sa  cc-by-nc-nd  cc0
text-mining-only   bronze   arxiv-default   closed   unknown

Plus helper booleans on FetchResult: may_redistribute, may_redistribute_nc, may_make_derivatives, may_train_models, may_use_for_tdm, plus a source_type field (publisher / repository / other) for callers who want publisher-vs-repository policy without parsing strings. Storage-tier policy is a one-line check.

Successful results also carry an attribution blockquote at the top of the markdown by default (source URL + DOI/PMCID + license + retrieval strategy), so the markdown is self-attributing when it travels to end users. Disable with Client(include_attribution=False).

What's NOT here

  • No Sci-Hub, no archive scraping, no UA spoofing past WAFs.
  • No Unpaywall and no institutional-repository scraping (removed in 0.1.0). A paper whose only open copy is one of those comes back with success=False.
  • No title-only paper fetch: use asta-papers search (or PaperFinder) first to get an identifier.
  • No multi-tenant Credentials per-call object (one Client per credential set).
  • No async API (sync only in v0.1).

Configuration

See Client.__init__ and docs/concepts.md for the full list. Required: email (Crossref polite-pool identifier; kwarg or ASTA_PAPERS_EMAIL env). Recommended: NCBI_API_KEY env (free, 5-minute registration, 3.3× throughput).

Tests

pytest tests/unit -q                         # offline
pytest tests/integration -v                  # real-API tests
python tools/check_test_legitimacy.py --strict  # asserts mock-ratio < 30%

The full integration suite hits live upstream APIs — no mocks. Tests run in ~90 seconds. A per-paper snapshot recovery benchmark (53 biomedical DOIs that fail Mistral-OCR-only retrieval) gates regressions. Its recorded baseline predates the removal of the Unpaywall route and will need re-recording: papers that were only reachable through Unpaywall now count as not recovered.

Design

Full design at docs/DESIGN.md.

Metadata

Release files for asta-papers 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for asta-papers 0.1.0
File Size Uploaded
asta_papers-0.1.0.tar.gz 103.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for asta-papers 0.1.0
File Interpreter ABI Platform
asta_papers-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 167.8 kB

Release files / asta_papers-0.1.0.tar.gz

Download URL asta_papers-0.1.0.tar.gz
Size 103.2 kB
Tags Source
SHA-256 checksum
How to use checksums
1595869633d6957196496b7bb14b8fe6b8f372a50f1c3cb5f05e834784d790ba
BLAKE2b-256 checksum
How to use checksums
6f6e48505b7982f3e2671bab4887a907fb613942d6666ce7dea0d1d8f135dca3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.7.13

Release files / asta_papers-0.1.0-py3-none-any.whl

Download URL asta_papers-0.1.0-py3-none-any.whl
Size 64.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
52a601ac8dbf7c137a659e4f1ee39d6d979b8bbf1a200fdcd11984baa25ef559
BLAKE2b-256 checksum
How to use checksums
6a612ff8b02525318a989ba477f270b423b768563a043c8208aa37b6e9949f84
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.7.13

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page