Skip to main content

bmlib

CI

Shared Python library for biomedical literature tools — LLM abstraction, quality assessment, transparency analysis, full-text retrieval, publication ingestion, and database utilities.

Version: 0.6.0 | License: AGPL-3.0-or-later | Python: >=3.11

Installation

# Core (only jinja2 dependency)
pip install bmlib

# Editable install with all extras
uv pip install -e ".[all,dev]"

Optional dependency groups

Group Install command Provides
anthropic pip install bmlib[anthropic] Anthropic Claude LLM provider
ollama pip install bmlib[ollama] Ollama local LLM provider
openai pip install bmlib[openai] OpenAI, DeepSeek, Mistral, Gemini, and OpenAI-compatible providers
postgresql pip install bmlib[postgresql] PostgreSQL database backend
transparency pip install bmlib[transparency] Transparency analysis (httpx)
publications pip install bmlib[publications] Publication ingestion and sync (httpx)
pdf pip install bmlib[pdf] PDF → text conversion (pymupdf)
dev pip install bmlib[dev] pytest, pytest-cov, ruff
all pip install bmlib[all] Every runtime extra above (not dev)

Modules

Module Description
bmlib.db Thin database abstraction (SQLite + PostgreSQL) with pure functions over DB-API connections
bmlib.llm Unified LLM client with pluggable providers (Anthropic, OpenAI, Ollama, DeepSeek, Mistral, Gemini) — chat, tool calling, embeddings, JSON repair, and text chunking
bmlib.templates Jinja2-based prompt template engine with user-override directory fallback
bmlib.agents Base agent class for LLM-driven tasks with template rendering and JSON parsing
bmlib.quality 3-tier quality assessment pipeline for biomedical publications (metadata → LLM classifier → deep assessment), plus Cochrane risk-of-bias models and rule-based extractors
bmlib.transparency Multi-API transparency and bias analysis (CrossRef, Europe PMC, OpenAlex, ClinicalTrials.gov)
bmlib.publications Publication ingestion from PubMed, bioRxiv, medRxiv, and OpenAlex with deduplication and sync
bmlib.fulltext Full-text retrieval (Europe PMC → Unpaywall → DOI), JATS XML parsing, PDF → text conversion, and disk-based caching

Quick Start

Database

from bmlib.db import connect_sqlite, execute, fetch_all, transaction

conn = connect_sqlite("~/.myapp/data.db")
with transaction(conn):
    execute(conn, "INSERT INTO papers (doi, title) VALUES (?, ?)", ("10.1101/x", "A paper"))
rows = fetch_all(conn, "SELECT * FROM papers")

LLM

from bmlib.llm import LLMClient, LLMMessage

client = LLMClient(default_provider="ollama")
response = client.chat(
    messages=[LLMMessage(role="user", content="Summarise this paper.")],
    model="ollama:medgemma4B_it_q8",
)
print(response.content)

Model strings use the format "provider:model_name":

"anthropic:claude-sonnet-4-20250514"
"openai:gpt-4o"
"ollama:medgemma4B_it_q8"
"deepseek:deepseek-chat"
"mistral:mistral-large-latest"
"gemini:gemini-2.0-flash"

Tool Calling

from bmlib.llm import LLMClient, LLMMessage, LLMToolDefinition

search = LLMToolDefinition(
    name="search_pubmed",
    description="Search PubMed for articles matching a query.",
    parameters={
        "type": "object",
        "properties": {"query": {"type": "string"}},
        "required": ["query"],
    },
)

client = LLMClient()
response = client.chat(
    messages=[LLMMessage(role="user", content="Find recent trials on statins.")],
    model="anthropic:claude-sonnet-4-20250514",
    tools=[search],
)

for call in response.tool_calls or []:
    print(call.name, call.arguments)  # arguments is already a parsed dict

To continue the conversation, append the assistant message (carrying tool_calls) and one role="tool" message per call, each with the matching tool_call_id, then send the whole list again.

Long Documents

from bmlib.llm import chunk_text, process_with_map_reduce

for chunk in chunk_text(paper_text, chunk_size=8000, overlap=200):
    print(chunk.chunk_index, chunk.size)

summary = process_with_map_reduce(
    paper_text,
    map_fn=lambda part: summarise(part),
    reduce_fn=lambda parts: summarise("\n".join(parts)),
)

Publication Sync

from datetime import date
from bmlib.db import connect_sqlite
from bmlib.publications import sync

conn = connect_sqlite("publications.db")
report = sync(
    conn,
    sources=["pubmed", "biorxiv"],
    date_from=date(2025, 1, 1),
    date_to=date(2025, 1, 7),
    email="researcher@example.com",
)
print(f"Added: {report.records_added}, Merged: {report.records_merged}")

Full-Text Retrieval

from bmlib.fulltext import FullTextService

service = FullTextService(email="researcher@example.com")

# Passing identifier= enables the built-in disk cache (platform default dir).
result = service.fetch_fulltext(
    pmc_id="PMC7614751", doi="10.1234/example", identifier="PMC7614751"
)

if result.html:
    print(result.html[:200])

Quality Assessment

from bmlib.llm import LLMClient
from bmlib.quality import QualityManager

llm = LLMClient()
manager = QualityManager(
    llm=llm,
    classifier_model="anthropic:claude-3-haiku-20240307",
    assessor_model="anthropic:claude-sonnet-4-20250514",
)

assessment = manager.assess(
    title="A Randomized Controlled Trial of ...",
    abstract="We conducted a double-blind RCT ...",
    publication_types=["Randomized Controlled Trial"],
)
print(assessment.study_design, assessment.quality_tier)

Transparency Analysis

from bmlib.transparency import TransparencyAnalyzer

analyzer = TransparencyAnalyzer(email="researcher@example.com")
result = analyzer.analyze("doc-001", doi="10.1038/s41586-024-00001-0")
print(result.transparency_score, result.risk_level)

Development

# Install with dev dependencies
uv pip install -e ".[all,dev]"

# Run tests
uv run pytest tests/ -v

# Lint and format
uv run ruff check .
uv run ruff format --check .

Documentation

Full API documentation is available in docs/manual/.

License

AGPL-3.0-or-later

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

bmlib-0.6.0.tar.gz (293.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

bmlib-0.6.0-py3-none-any.whl (210.4 kB view details)

Uploaded Python 3

File details

Details for the file bmlib-0.6.0.tar.gz.

File metadata

  • Download URL: bmlib-0.6.0.tar.gz
  • Upload date:
  • Size: 293.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.6

File hashes

Hashes for bmlib-0.6.0.tar.gz
Algorithm Hash digest
SHA256 08935ef21cd4f8ca4bf77f67bb369871ebc1941d4c909db2db6f81eb4186404e
MD5 289fff42996a56d4c451c7f42cb02506
BLAKE2b-256 5b6c90a6cb6878a6553e604fcbe5b2e389f727755286702e8871a29019b9c0a6

See more details on using hashes here.

File details

Details for the file bmlib-0.6.0-py3-none-any.whl.

File metadata

  • Download URL: bmlib-0.6.0-py3-none-any.whl
  • Upload date:
  • Size: 210.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.6

File hashes

Hashes for bmlib-0.6.0-py3-none-any.whl
Algorithm Hash digest
SHA256 971d4efbfc67a20f30b756365a897a1197940b7e9b8b70507db890ee405ae2a9
MD5 551122e35279ee87d65a6c9ff818454b
BLAKE2b-256 60003dd4da904e9f4cd68da25f3d51da39aeb047f1d7d32d27c4f2ea4d670264

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page