Skip to main content

PyEuropePMC

PyPI version PyPI - Downloads Python 3.10+ License: MIT Tests Coverage Documentation MCP Server

🔄 Build Status

CI/CD Pipeline Python Compatibility Documentation CodeQL codecov

PyEuropePMC is a robust Python toolkit for automated search, extraction, and analysis of scientific literature from Europe PMC.

✨ Key Features

  • 🔍 Comprehensive Search API - Query Europe PMC with advanced search options
  • Advanced Query Builder - Fluent API for building complex search queries with type safety
  • �📄 Full-Text Retrieval - Download PDFs, XML, and HTML content from open access articles
  • 🔬 XML Parsing & Conversion - Parse full text XML and convert to plaintext, markdown, extract tables and metadata
  • 🏷️ Text-Mining Annotations - Retrieve and parse entity annotations, sentences, and relationships (genes, diseases, chemicals)
  • 📊 Multiple Output Formats - JSON, XML, Dublin Core (DC)
  • 📦 Bulk FTP Downloads - Efficient bulk PDF downloads from Europe PMC FTP servers
  • 🔄 Smart Pagination - Automatic handling of large result sets
  • 🛡️ Robust Error Handling - Built-in retry logic and connection management
  • 🧑‍💻 Type Safety - Extensive use of type annotations and validation
  • Rate Limiting - Respectful API usage with configurable delays
  • 🧪 Extensively Tested - 200+ tests with 90%+ code coverage
  • 📋 Systematic Review Tracking - PRISMA-compliant search logging and audit trails
  • 📈 Advanced Analytics - Publication trends, citation analysis, quality metrics, and duplicate detection
  • 📉 Rich Visualizations - Interactive plots and dashboards using matplotlib and seaborn
  • 🔗 External API Enrichment - Enhance metadata with CrossRef, Unpaywall, Semantic Scholar, and OpenAlex
  • 🤖 MCP Server Support - Model Context Protocol integration for LLM tool usage

📁 Project Structure

The repository is organized as follows:

  • src/pyeuropepmc/ - Main package source code
  • tests/ - Unit and integration tests
  • docs/ - Documentation and guides
  • examples/ - Example scripts and usage demonstrations
  • benchmarks/ - Performance benchmarking scripts and results
  • data/ - Downloads, outputs, and generated data files
  • conf/ - Configuration files for RDF mapping and other settings

🚀 Quick Start

Installation

pip install pyeuropepmc                 # light core
pip install "pyeuropepmc[all]"          # everything (1.x-equivalent)
pip install "pyeuropepmc[analytics,agentic]"   # pick what you need

Upgrading from 1.x? See docs/migration/v1-to-v2.md.

Basic Usage

from pyeuropepmc import SearchClient

# Search for papers
with SearchClient() as client:
    results = client.search("CRISPR gene editing", pageSize=10)

    for paper in results["resultList"]["result"]:
        print(f"Title: {paper['title']}")
        print(f"Authors: {paper.get('authorString', 'N/A')}")
        print("---")

Advanced Search with QueryBuilder

from pyeuropepmc import QueryBuilder

# Build complex queries with fluent API
qb = QueryBuilder()
query = (qb
    .keyword("cancer", field="title")
    .and_()
    .keyword("immunotherapy")
    .and_()
    .date_range(start_year=2020, end_year=2023)
    .and_()
    .citation_count(min_count=10)
    .build())

print(f"Generated query: {query}")
# Output: (TITLE:cancer) AND immunotherapy AND (PUB_YEAR:[2020 TO 2023]) AND (CITED:[10 TO *])

Advanced Search with Parsing

# Search and automatically parse results
papers = client.search_and_parse(
    query="COVID-19 AND vaccine",
    pageSize=50,
    sort="CITED desc"
)

for paper in papers:
    print(f"Citations: {paper.get('citedByCount', 0)}")
    print(f"Title: {paper.get('title', 'N/A')}")

Full-Text Content Retrieval

from pyeuropepmc import FullTextClient

# Initialize full-text client
fulltext_client = FullTextClient()

# Download PDF
pdf_path = fulltext_client.download_pdf_by_pmcid("PMC1234567", output_dir="./downloads")

# Download XML
xml_content = fulltext_client.download_xml_by_pmcid("PMC1234567")

# Bulk FTP downloads
from pyeuropepmc import FTPDownloader

ftp_downloader = FTPDownloader()
results = ftp_downloader.bulk_download_and_extract(
    pmcids=["1234567", "2345678"],
    output_dir="./bulk_downloads"
)

Full-Text XML Parsing

Parse full text XML files and extract structured information:

from pyeuropepmc import FullTextClient, FullTextXMLParser

# Download and parse XML
with FullTextClient() as client:
    xml_path = client.download_xml_by_pmcid("PMC3258128")

# Parse the XML
with open(xml_path, 'r') as f:
    parser = FullTextXMLParser(f.read())

# Extract metadata
metadata = parser.extract_metadata()
print(f"Title: {metadata['title']}")
print(f"Authors: {', '.join(metadata['authors'])}")

# Convert to different formats
plaintext = parser.to_plaintext()  # Plain text
markdown = parser.to_markdown()     # Markdown format

# Extract tables
tables = parser.extract_tables()
for table in tables:
    print(f"Table: {table['label']} - {len(table['rows'])} rows")

# Extract references
references = parser.extract_references()
print(f"Found {len(references)} references")

Text-Mining Annotations

Retrieve and parse entity annotations, sentences, and relationships from scientific literature:

from pyeuropepmc import AnnotationsClient, parse_annotations

# Initialize annotations client
with AnnotationsClient() as client:
    # Get annotations for specific articles
    annotations = client.get_annotations_by_article_ids(
        article_ids=["PMC3359311"],
        section="abstract"  # or "fulltext", "all"
    )

    # Parse annotations to extract structured data
    parsed = parse_annotations(annotations)

    print(f"Found {len(parsed['entities'])} entities")
    print(f"Found {len(parsed['relationships'])} relationships")

    # Display entities by type
    for entity in parsed['entities'][:5]:
        print(f"{entity['name']} ({entity['type']})")

    # Search for specific entities (e.g., chemicals)
    entity_annotations = client.get_annotations_by_entity(
        entity_id="CHEBI:16236",  # Ethanol
        entity_type="CHEBI",
        page_size=20
    )

    # Filter by annotation provider
    provider_annotations = client.get_annotations_by_provider(
        provider="Europe PMC",
        annotation_type="Disease"
    )

Supported Entity Types:

  • 🧬 Genes and proteins
  • 🦠 Diseases and conditions
  • 🧪 Chemicals and drugs (CHEBI)
  • 🔬 Gene Ontology terms
  • 🌱 Organisms and species
  • 🔗 Entity relationships

See examples/10-annotations for detailed examples.

Advanced Analytics and Visualization

Analyze search results with built-in analytics and create visualizations:

from pyeuropepmc import (
    SearchClient,
    to_dataframe,
    citation_statistics,
    quality_metrics,
    remove_duplicates,
    plot_publication_years,
    create_summary_dashboard,
)

# Search and convert to DataFrame
with SearchClient() as client:
    response = client.search("machine learning", pageSize=100)
    papers = response.get("resultList", {}).get("result", [])

# Convert to pandas DataFrame for analysis
df = to_dataframe(papers)

# Remove duplicates
df = remove_duplicates(df, method="title", keep="most_cited")

# Get citation statistics
stats = citation_statistics(df)
print(f"Mean citations: {stats['mean_citations']:.2f}")
print(f"Highly cited (top 10%): {stats['citation_distribution']['90th_percentile']:.0f}")

# Assess quality metrics
metrics = quality_metrics(df)
print(f"Open access: {metrics['open_access_percentage']:.1f}%")
print(f"With PDF: {metrics['with_pdf_percentage']:.1f}%")

# Create visualizations
plot_publication_years(df, save_path="publications_by_year.png")
create_summary_dashboard(df, save_path="analysis_dashboard.png")

External API Enrichment

Enhance paper metadata with data from CrossRef, Unpaywall, Semantic Scholar, and OpenAlex:

Professional Semantic Scholar Integration (v0.12.0)

PyEuropePMC now uses the danielnsilva/semanticscholar professional library for robust Semantic Scholar API integration:

Usage with API Key:

from pyeuropepmc.features.enrich.sources.semantic_scholar import SemanticScholarClient

# With API key (recommended for higher rate limits)
client = SemanticScholarClient(api_key="your_api_key_here")

# Get enriched paper data
result = client.enrich(semantic_scholar_id="649def34f8be52c8b66281af98ae884c09aef38b")
print(f"Citations: {result['citation_count']}")  # 439
print(f"Influential: {result['influential_citation_count']}")

# Get recommendations
recommendations = client.get_recommendations_for_paper("649def34f8be52c8b66281af98ae884c09aef38b")

Usage for Bulk Search:

For search operations that may be rate-limited, use bulk=True for faster results:

# Search with bulk retrieval (no relevance ranking, faster)
results = client.search_papers(query="machine learning cancer", bulk=True)

# Or with relevance ranking (default, may be rate-limited)
results = client.search_papers(query="machine learning cancer", bulk=False)

Note: Search operations may be rate-limited depending on API usage. For reliable results with specific papers, use client.enrich() with paper IDs (DOI, S2PaperId, etc.). The search_papers() method is best used with bulk=True for faster, non-ranked results, or with specific filters to reduce the result set.

Benefits of Professional Library:

  • ✅ Typed response objects (Paper, Author, Venue)
  • ✅ Automatic retries and rate limiting
  • ✅ Full API coverage (Graph, Recommendations, Datasets)
  • ✅ Async support for concurrent requests
  • ✅ Built-in pagination handling
  • ✅ Production-ready (461 stars on GitHub)

Usage: examples/09-enrichment/ (basic_enrichment.py, advanced_enrichment.py)

from pyeuropepmc import PaperEnricher, EnrichmentConfig

# Configure enrichment with multiple APIs
config = EnrichmentConfig(
    enable_crossref=True,
    enable_semantic_scholar=True,
    enable_openalex=True,
    enable_unpaywall=True,
    unpaywall_email="your@email.com"  # Required for Unpaywall
)

# Enrich paper metadata
with PaperEnricher(config) as enricher:
    result = enricher.enrich_paper(doi="10.1371/journal.pone.0308090")

    # Access merged data from all sources
    merged = result["merged"]
    print(f"Title: {merged['title']}")
    print(f"Citations: {merged['citation_count']}")
    print(f"Open Access: {merged['is_oa']}")

    # Access individual source data
    if "crossref" in result["sources"]:
        print(f"Funders: {result['crossref']['funders']}")

    if "semantic_scholar" in result["sources"]:
        print(f"Influential Citations: {result['semantic_scholar']['influential_citation_count']}")

Features:

  • 🔄 Automatic data merging from multiple sources
  • 📊 Citation metrics from multiple databases
  • 🔓 Open access status and full-text URLs
  • 💰 Funding information
  • 🏷️ Topic classifications and fields of study
  • ⚡ Optional caching for performance
  • 📚 Professional Semantic Scholar client (typed responses, async support)

See examples/09-enrichment for more details.

Knowledge Graph Structure Options 🕸️

PyEuropePMC supports flexible knowledge graph structures for different use cases:

from pyeuropepmc.mappers import RDFMapper

mapper = RDFMapper()

# Metadata-only KG (for citation networks and bibliometrics)
metadata_graphs = mapper.save_metadata_rdf(
    entities_data,
    output_dir="rdf_output"
)  # Papers + authors + institutions

# Content-only KG (for text analysis and document processing)
content_graphs = mapper.save_content_rdf(
    entities_data,
    output_dir="rdf_output"
)  # Papers + sections + references + tables

# Complete KG (for comprehensive analysis)
complete_graphs = mapper.save_complete_rdf(
    entities_data,
    output_dir="rdf_output"
)  # All entities and relationships

# Use configured default from conf/rdf_map.yml
graphs = mapper.save_rdf(entities_data, output_dir="rdf_output")

Use Cases:

  • 📊 Citation Networks: Use metadata-only KGs for bibliometric analysis
  • 📝 Text Mining: Use content-only KGs for NLP and information extraction
  • 🔬 Full Analysis: Use complete KGs for comprehensive research workflows

See examples/kg_structure_demo.py for a complete working example.

Unified Processing Pipeline 🏗️

The new unified pipeline dramatically simplifies the complex workflow of XML parsing → enrichment → RDF conversion:

from pyeuropepmc import PaperProcessingPipeline, PipelineConfig

# Simple configuration
config = PipelineConfig(
    enable_enrichment=True,      # Enable metadata enrichment
    enable_crossref=True,        # CrossRef API
    enable_semantic_scholar=True, # Semantic Scholar API
    enable_openalex=True,        # OpenAlex API
    enable_ror=True,             # ROR institution data
    crossref_email="your@email.com",  # Required for higher CrossRef rate limits
    output_format="turtle",      # RDF output format
    output_dir="output"          # Where to save RDF files
)

# Create unified pipeline
pipeline = PaperProcessingPipeline(config)

# Process single paper - replaces 8+ separate steps!
result = pipeline.process_paper(
    xml_content=xml_string,
    doi="10.1038/nature11476",
    save_rdf=True
)

print(f"Generated {result['triple_count']} RDF triples")
print(f"Output saved to: {result['output_file']}")

# Process multiple papers in batch
xml_contents = {
    "10.1038/nature11476": xml_content_1,
    "10.1038/nature11477": xml_content_2,
}

batch_results = pipeline.process_papers(xml_contents)
for doi, result in batch_results.items():
    print(f"{doi}: {result['triple_count']} triples")

What it does automatically:

  • ✅ Parses XML and extracts entities (paper, authors, sections, tables, figures, references)
  • ✅ Enriches metadata from external APIs (citations, fields of study, etc.)
  • ✅ Converts everything to RDF with proper relationships
  • ✅ Saves structured output files
  • ✅ Handles errors gracefully

Before vs After:

# OLD: Complex multi-step workflow (8+ steps)
parser = FullTextXMLParser()
parser.parse(xml_content)
paper, authors, sections, tables, figures, references = build_paper_entities(parser)
enricher = PaperEnricher(config)
enrichment_data = enricher.enrich_paper(doi)
rdf_mapper = RDFMapper()
paper.to_rdf(graph, related_entities=...)
rdf_mapper.serialize_graph(graph, format='turtle')

# NEW: Single pipeline call (3 steps)
config = PipelineConfig(...)
pipeline = PaperProcessingPipeline(config)
result = pipeline.process_paper(xml_content, doi=doi)

See examples/pipeline_demo.py for a complete working example.

📚 Documentation

📖 Read the Full Documentation ← Start Here!

Quick Links:

Note: Enable GitHub Pages first! See Setup Guide for instructions.

📊 Parser Quality Benchmark

The XML full-text parser is continuously evaluated against a curated benchmark of 55 open-access JATS articles from Europe PMC. Results demonstrate high-fidelity extraction across all quality dimensions:

Metric Mean Min Max Std Dev
Composite Score 0.9871 0.9643 0.9992 0.0086
Metadata Accuracy 1.0000 1.0000 1.0000 0.0000
Text Fidelity 1.0000 1.0000 1.0000 0.0000
Element Coverage 0.9925 0.9655 1.0000 0.0087
Section Accuracy 0.9431 0.8333 1.0000 0.0445
Inline Recall 1.0000 1.0000 1.0000 0.0000

Parse speed: 55.0 articles in 2.05s (26.8 articles/s)

PLOS XML Support

The parser now handles PLOS articles that use bare <p> elements directly under <body> (without <sec> wrappers). This structure was previously ignored, causing near-zero scores on PLOS-only benchmarks.

Metric Before Fix After Fix
Composite Score 0.4734 (min) 0.9778 ± 0.0393
Metadata Accuracy 0.6000 ± 0.0000 1.0000 ± 0.0000
Inline Recall 0.0000 (min) 1.0000 ± 0.0000
Text Fidelity 0.3026 (min) 1.0000 ± 0.0000
Section Accuracy 0.5745 ± 0.1812 0.9272 ± 0.0870
Element Coverage 0.9617 ± 0.0118 0.9617 ± 0.0118

Key fixes:

  • Section parser: extract text from bare <p> elements directly under <body>
  • Content blocks: collect bare <p> paragraphs as a synthetic body section
  • Plaintext converter: include bare <p> elements in body text output
  • Metadata matching: empty-empty fields (e.g., no PMID/PMCID) count as matches
  • Section accuracy: detect bare <p> sections for correct section path tracking

Run the benchmark yourself:

pyeuropepmc benchmark run local --local-path benchmark_xmls/xml --limit 55
pyeuropepmc benchmark run local --local-path benchmark_xmls/xml --dataset plos1000

See the Benchmarking Guide for full methodology and profiling tools.

🤝 Contributing

We welcome contributions! See the development docs and open an issue or PR to get started.

📄 License

Distributed under the MIT License. See LICENSE for more information.

🌐 Links

🤖 MCP Server

PyEuropePMC includes a Model Context Protocol (MCP) server for use with LLMs and AI assistants.

Installation

pip install pyeuropepmc

Usage

The MCP server provides four tools:

  • search_papers - Search for papers in Europe PMC
  • get_paper_details - Get detailed information about a paper (by PMID, PMCID, or DOI)
  • search_authors - Search for authors in Europe PMC
  • get_paper_citations - Get citations for a paper

Running the MCP Server

# As a standalone server
pyeuropepmc-mcp

# Or using Python directly
python -m pyeuropepmc.mcp.server

Using with LLMs

The server implements the MCP protocol and can be configured in your LLM application:

{
  "mcpServers": {
    "pyeuropepmc": {
      "command": "python",
      "args": ["/path/to/pyeuropepmc-mcp"]
    }
  }
}

The server exposes these tools over the MCP protocol (JSON-RPC on stdio): unified_search, get_paper_details, search_authors, get_paper_citations, citation_snowball, clinical_trial_search, fulltext_index_query, paper_figures, plus optional LLM and bibliography tools. See src/pyeuropepmc/mcp/server.py for the full registry and input schemas.

For direct Python use (no MCP client), call the same underlying APIs:

from pyeuropepmc import SearchClient
from pyeuropepmc.features.search import UnifiedSearch

# Europe PMC only
papers = SearchClient().search_and_parse("CRISPR gene editing", pageSize=10)

# Europe PMC + other sources, deduplicated
merged, report = UnifiedSearch(sources=["europepmc", "pubmed", "arxiv"]).search(
    "CRISPR gene editing", limit=10
)

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pyeuropepmc-2.0.0.tar.gz (604.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pyeuropepmc-2.0.0-py3-none-any.whl (723.1 kB view details)

Uploaded Python 3

File details

Details for the file pyeuropepmc-2.0.0.tar.gz.

File metadata

  • Download URL: pyeuropepmc-2.0.0.tar.gz
  • Upload date:
  • Size: 604.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for pyeuropepmc-2.0.0.tar.gz
Algorithm Hash digest
SHA256 26fa5125f982a67a1415f1facfe798826297cc456665753d74f345d55cdb76c0
MD5 cf22b90849465eef517c95c5e9aec251
BLAKE2b-256 6726cbbe0abdfcba1b818dcb8f94a0801a3ee708b547ae9731b39d1f27af3bfe

See more details on using hashes here.

Provenance

The following attestation bundles were made for pyeuropepmc-2.0.0.tar.gz:

Publisher: release.yml on JonasHeinickeBio/pyEuropePMC

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pyeuropepmc-2.0.0-py3-none-any.whl.

File metadata

  • Download URL: pyeuropepmc-2.0.0-py3-none-any.whl
  • Upload date:
  • Size: 723.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for pyeuropepmc-2.0.0-py3-none-any.whl
Algorithm Hash digest
SHA256 953c7bbb0b1e82e2a5e54422ef2b1a05ae338750dba2d449787d42f6383bd22d
MD5 abca4d592a97265aea9f88eb35326aab
BLAKE2b-256 105751d00948ec5743dfc1613507103c5474c6aaaa35d5e768bb4b68c8ad46d7

See more details on using hashes here.

Provenance

The following attestation bundles were made for pyeuropepmc-2.0.0-py3-none-any.whl:

Publisher: release.yml on JonasHeinickeBio/pyEuropePMC

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

2.2.1

2 files

2.2.0

2 files

2.1.1

2 files

2.1.0

2 files

This release

2.0.0 This release

2 files

1.17.0

2 files

1.16.0

2 files

1.14.0

2 files

1.13.0

2 files

1.12.0

2 files

1.11.3

2 files

1.11.2

2 files

1.11.1

2 files

1.11.0

2 files

1.10.1

2 files

1.10.0

2 files

1.9.1

2 files

1.9.0

2 files

1.8.1

2 files

1.8.0

2 files

1.7.0

2 files

1.6.0

2 files

1.5.0

2 files

1.4.0

2 files

1.3.0

2 files

1.2.0

2 files

1.1.0

2 files

1.0.0

2 files

0.0.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page