Skip to main content

Shared library for biomedical literature tools — LLM abstraction, quality assessment, transparency analysis, and database utilities

Project description

bmlib

CI

Shared Python library for biomedical literature tools — LLM abstraction, quality assessment, transparency analysis, full-text retrieval, publication ingestion, and database utilities.

Version: 0.5.0 | License: AGPL-3.0-or-later | Python: >=3.11

Installation

# Core (only jinja2 dependency)
pip install bmlib

# Editable install with all extras
uv pip install -e ".[all,dev]"

Optional dependency groups

Group Install command Provides
anthropic pip install bmlib[anthropic] Anthropic Claude LLM provider
ollama pip install bmlib[ollama] Ollama local LLM provider
openai pip install bmlib[openai] OpenAI, DeepSeek, Mistral, Gemini, and OpenAI-compatible providers
postgresql pip install bmlib[postgresql] PostgreSQL database backend
transparency pip install bmlib[transparency] Transparency analysis (httpx)
publications pip install bmlib[publications] Publication ingestion and sync (httpx)
pdf pip install bmlib[pdf] PDF → text conversion (pymupdf)
dev pip install bmlib[dev] pytest, pytest-cov, ruff
all pip install bmlib[all] Every runtime extra above (not dev)

Modules

Module Description
bmlib.db Thin database abstraction (SQLite + PostgreSQL) with pure functions over DB-API connections
bmlib.llm Unified LLM client with pluggable providers (Anthropic, OpenAI, Ollama, DeepSeek, Mistral, Gemini) — chat, tool calling, embeddings, JSON repair, and text chunking
bmlib.templates Jinja2-based prompt template engine with user-override directory fallback
bmlib.agents Base agent class for LLM-driven tasks with template rendering and JSON parsing
bmlib.quality 3-tier quality assessment pipeline for biomedical publications (metadata → LLM classifier → deep assessment), plus Cochrane risk-of-bias models and rule-based extractors
bmlib.transparency Multi-API transparency and bias analysis (CrossRef, Europe PMC, OpenAlex, ClinicalTrials.gov)
bmlib.publications Publication ingestion from PubMed, bioRxiv, medRxiv, and OpenAlex with deduplication and sync
bmlib.fulltext Full-text retrieval (Europe PMC → Unpaywall → DOI), JATS XML parsing, PDF → text conversion, and disk-based caching

Quick Start

Database

from bmlib.db import connect_sqlite, execute, fetch_all, transaction

conn = connect_sqlite("~/.myapp/data.db")
with transaction(conn):
    execute(conn, "INSERT INTO papers (doi, title) VALUES (?, ?)", ("10.1101/x", "A paper"))
rows = fetch_all(conn, "SELECT * FROM papers")

LLM

from bmlib.llm import LLMClient, LLMMessage

client = LLMClient(default_provider="ollama")
response = client.chat(
    messages=[LLMMessage(role="user", content="Summarise this paper.")],
    model="ollama:medgemma4B_it_q8",
)
print(response.content)

Model strings use the format "provider:model_name":

"anthropic:claude-sonnet-4-20250514"
"openai:gpt-4o"
"ollama:medgemma4B_it_q8"
"deepseek:deepseek-chat"
"mistral:mistral-large-latest"
"gemini:gemini-2.0-flash"

Tool Calling

from bmlib.llm import LLMClient, LLMMessage, LLMToolDefinition

search = LLMToolDefinition(
    name="search_pubmed",
    description="Search PubMed for articles matching a query.",
    parameters={
        "type": "object",
        "properties": {"query": {"type": "string"}},
        "required": ["query"],
    },
)

client = LLMClient()
response = client.chat(
    messages=[LLMMessage(role="user", content="Find recent trials on statins.")],
    model="anthropic:claude-sonnet-4-20250514",
    tools=[search],
)

for call in response.tool_calls or []:
    print(call.name, call.arguments)  # arguments is already a parsed dict

To continue the conversation, append the assistant message (carrying tool_calls) and one role="tool" message per call, each with the matching tool_call_id, then send the whole list again.

Long Documents

from bmlib.llm import chunk_text, process_with_map_reduce

for chunk in chunk_text(paper_text, chunk_size=8000, overlap=200):
    print(chunk.chunk_index, chunk.size)

summary = process_with_map_reduce(
    paper_text,
    map_fn=lambda part: summarise(part),
    reduce_fn=lambda parts: summarise("\n".join(parts)),
)

Publication Sync

from datetime import date
from bmlib.db import connect_sqlite
from bmlib.publications import sync

conn = connect_sqlite("publications.db")
report = sync(
    conn,
    sources=["pubmed", "biorxiv"],
    date_from=date(2025, 1, 1),
    date_to=date(2025, 1, 7),
    email="researcher@example.com",
)
print(f"Added: {report.records_added}, Merged: {report.records_merged}")

Full-Text Retrieval

from bmlib.fulltext import FullTextService

service = FullTextService(email="researcher@example.com")

# Passing identifier= enables the built-in disk cache (platform default dir).
result = service.fetch_fulltext(
    pmc_id="PMC7614751", doi="10.1234/example", identifier="PMC7614751"
)

if result.html:
    print(result.html[:200])

Quality Assessment

from bmlib.llm import LLMClient
from bmlib.quality import QualityManager

llm = LLMClient()
manager = QualityManager(
    llm=llm,
    classifier_model="anthropic:claude-3-haiku-20240307",
    assessor_model="anthropic:claude-sonnet-4-20250514",
)

assessment = manager.assess(
    title="A Randomized Controlled Trial of ...",
    abstract="We conducted a double-blind RCT ...",
    publication_types=["Randomized Controlled Trial"],
)
print(assessment.study_design, assessment.quality_tier)

Transparency Analysis

from bmlib.transparency import TransparencyAnalyzer

analyzer = TransparencyAnalyzer(email="researcher@example.com")
result = analyzer.analyze("doc-001", doi="10.1038/s41586-024-00001-0")
print(result.transparency_score, result.risk_level)

Development

# Install with dev dependencies
uv pip install -e ".[all,dev]"

# Run tests
uv run pytest tests/ -v

# Lint and format
uv run ruff check .
uv run ruff format --check .

Documentation

Full API documentation is available in docs/manual/.

License

AGPL-3.0-or-later

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

bmlib-0.5.0.tar.gz (193.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

bmlib-0.5.0-py3-none-any.whl (166.5 kB view details)

Uploaded Python 3

File details

Details for the file bmlib-0.5.0.tar.gz.

File metadata

  • Download URL: bmlib-0.5.0.tar.gz
  • Upload date:
  • Size: 193.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.6

File hashes

Hashes for bmlib-0.5.0.tar.gz
Algorithm Hash digest
SHA256 b8c53e15402e9056afb83918d6783c8236b286df25202b4bc4f6f0187d97050c
MD5 5ea669749be848e54351b39b239e2a98
BLAKE2b-256 4f2fd71b477d59c10982503157279a767ec151f8bedc5c90041d018bdb55aef6

See more details on using hashes here.

File details

Details for the file bmlib-0.5.0-py3-none-any.whl.

File metadata

  • Download URL: bmlib-0.5.0-py3-none-any.whl
  • Upload date:
  • Size: 166.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.6

File hashes

Hashes for bmlib-0.5.0-py3-none-any.whl
Algorithm Hash digest
SHA256 6d1418fcc26b6bb838216e98fac504c25aa8e933d0c2ac526f1a46ad119011b4
MD5 fb48e68d041dc22695ebd8091aed247b
BLAKE2b-256 43c365db18d2db77d7a5a454f88f76e49e9e5cc6849b5dca32dafbe0e32f4075

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page