asta-papers
Legal-only paper full-text retrieval and conversion. Identifier (DOI / PMID / PMCID / arXiv / Semantic Scholar Corpus ID) — or BYO PDF / JATS — to markdown with explicit license classification.
Lifts paper recovery on biomedical literature using only publisher-blessed legal channels (NCBI E-utilities, EuropePMC, bioRxiv, arXiv). It does not query Unpaywall or scrape institutional repositories.
It also installs the asta-papers command: Semantic Scholar search, paper
lookup and citation walks, whose JSON output is what downstream tools such as
theorize papers fetch read.
export S2_API_KEY=... # optional; 1 request/s without, 100 with
asta-papers search "astrocyte lactate shuttle" --limit 20 --date 2018-
asta-papers get ARXIV:2005.14165 --format text
asta-papers citations CorpusId:218487638 --limit 50
asta-papers search "..." --limit 30 | asta-papers store --into store.json # fetch the text, build the store
store fetches the full text of what a search found and writes the paper store Theorizer reads,
resumably. PDFs go through Mistral OCR when MISTRAL_API_KEY is set, otherwise through pymupdf4llm
(pip install 'asta-papers[pymupdf]').
Install
pip install asta-papers # core (JATS conversion only)
pip install 'asta-papers[mistral]' # + Mistral OCR for PDFs
pip install 'asta-papers[olmocr]' # + local olmOCR for PDFs (offline)
pip install 'asta-papers[s3]' # + s3:// BYO support
pip install 'asta-papers[all]' # everything
Quickstart
import os
from asta_papers import Client
from asta_papers.converters.mistral import MistralConverter
c = Client(
email="me@allenai.org",
ncbi_api_key=os.environ.get("NCBI_API_KEY"), # optional, lifts NCBI 3→10 rps
converters=[MistralConverter()], # for PDF→markdown
)
# By identifier
r = c.fetch(doi="10.1186/s12943-024-02093-w")
print(r.success, r.license_class, r.markdown[:200])
# Storage-tier policy
if r.may_redistribute: # CC BY / CC0 / CC BY-SA
save_artifact(r.bytes)
elif r.may_use_for_tdm: # TDM-permissive licenses; not BRONZE/UNKNOWN/CLOSED
extract_inline(r.markdown)
# BYO PDF — bytes, local path, or URI
r = c.fetch(pdf=b"%PDF-...")
r = c.fetch(pdf="paper.pdf")
r = c.fetch(pdf="s3://my-bucket/paper.pdf") # requires [s3]
# Batch with bounded concurrency + per-host rate limits
results = c.fetch_many([
{"doi": "10.1038/foo"},
{"pmcid": "PMC123"},
{"pdf": "paper.pdf", "doi": "10.99/local"},
])
How it works
A strategy ladder runs against legal aggregator APIs in order, returning the first successful result:
- PMC E-utilities efetch — JATS XML for PMC OA Subset articles
- NCBI elink — PMID → PMC self-link when S2 didn't surface it
- Published-version handoff — when input is a preprint DOI, route to
the published version (bioRxiv API or Crossref
relation) so callers get the most-recent public version of the paper - arXiv — direct PDF for arXiv papers
- bioRxiv / medRxiv API — JATS XML or PDF for preprints
- EuropePMC PDF render — text-mining-licensed PDFs for free-to-read papers NCBI's OA Subset doesn't include
Per-host token-bucket rate limiting honors every publisher's published quota
exactly. arXiv at 0.33 rps (their explicit rule). NCBI at 3 rps (10 with key)
shared across eutils.*, pmc.*, www.* hostnames. Independent services
run fully in parallel.
License classification
Every successful fetch carries a LicenseClass:
cc-by cc-by-sa cc-by-nd cc-by-nc cc-by-nc-sa cc-by-nc-nd cc0
text-mining-only bronze arxiv-default closed unknown
Plus helper booleans on FetchResult: may_redistribute,
may_redistribute_nc, may_make_derivatives, may_train_models,
may_use_for_tdm, plus a source_type field (publisher /
repository / other) for callers who want publisher-vs-repository
policy without parsing strings. Storage-tier policy is a one-line check.
Successful results also carry an attribution blockquote at the top of
the markdown by default (source URL + DOI/PMCID + license + retrieval
strategy), so the markdown is self-attributing when it travels to end
users. Disable with Client(include_attribution=False).
What's NOT here
- No Sci-Hub, no archive scraping, no UA spoofing past WAFs.
- No Unpaywall and no institutional-repository scraping (removed in 0.1.0). A
paper whose only open copy is one of those comes back with
success=False. - No title-only paper fetch: use
asta-papers search(or PaperFinder) first to get an identifier. - No multi-tenant
Credentialsper-call object (one Client per credential set). - No async API (sync only in v0.1).
Configuration
See Client.__init__ and docs/concepts.md for the full list. Required:
email (Crossref polite-pool identifier; kwarg or ASTA_PAPERS_EMAIL env).
Recommended: NCBI_API_KEY env (free, 5-minute registration, 3.3× throughput).
Tests
pytest tests/unit -q # offline
pytest tests/integration -v # real-API tests
python tools/check_test_legitimacy.py --strict # asserts mock-ratio < 30%
The full integration suite hits live upstream APIs — no mocks. Tests run in ~90 seconds. A per-paper snapshot recovery benchmark (53 biomedical DOIs that fail Mistral-OCR-only retrieval) gates regressions. Its recorded baseline predates the removal of the Unpaywall route and will need re-recording: papers that were only reachable through Unpaywall now count as not recovered.
Design
Full design at docs/DESIGN.md.
Metadata
Release files for asta-papers 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| asta_papers-0.1.0.tar.gz | 103.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| asta_papers-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 167.8 kB
Release files / asta_papers-0.1.0.tar.gz
| Download URL | asta_papers-0.1.0.tar.gz |
|---|---|
| Size | 103.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1595869633d6957196496b7bb14b8fe6b8f372a50f1c3cb5f05e834784d790ba
|
|
BLAKE2b-256 checksum How to use checksums |
6f6e48505b7982f3e2671bab4887a907fb613942d6666ce7dea0d1d8f135dca3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.7.13
|
Release files / asta_papers-0.1.0-py3-none-any.whl
| Download URL | asta_papers-0.1.0-py3-none-any.whl |
|---|---|
| Size | 64.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
52a601ac8dbf7c137a659e4f1ee39d6d979b8bbf1a200fdcd11984baa25ef559
|
|
BLAKE2b-256 checksum How to use checksums |
6a612ff8b02525318a989ba477f270b423b768563a043c8208aa37b6e9949f84
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.7.13
|