Skip to main content

BioHarbor

Run real bioinformatics from your AI agent. Reliable, reproducible, on your own GPUs.

CI PyPI License

🧪 Alpha (v0.1). Sequence tools, homology search (MMseqs2) and structure prediction (ESMFold, validated on RTX 5090) work. Feedback welcome — see the roadmap.

BioHarbor is an MCP server that lets AI agents such as Claude execute bioinformatics tools — not just look things up. Agents ask for an analysis; BioHarbor validates the input, schedules it on a GPU with room, records exactly how it ran, and hands back a compact, agent-readable summary.

Why another bio MCP server?

Most bio MCP servers wrap databases (UniProt, PDB, PubMed…). Use them — BioHarbor complements them by running the compute:

Database MCP servers BioHarbor
Runs real analyses (search, fold, cluster) ❌ ✅
Validates inputs before burning GPU time ❌ ✅
GPU-aware queue, polite on shared GPUs ❌ ✅
Long jobs return a job_id instead of timing out ❌ ✅
Compact summaries + files on disk (saves tokens) ❌ ✅
Provenance for every run, export to a pipeline ❌ ✅ (export: planned)

Quick start

pip install bioharbor
bioharbor doctor             # checks Python, GPUs, workspace, tools
bioharbor setup-db swissprot # reference database for search_homologs (needs MMseqs2)
bioharbor install-claude     # prints the config for Claude Desktop / Claude Code

For structure prediction on a GPU: pip install "bioharbor[esmfold]" — see docs/gpu-setup.md (RTX 50xx needs a CUDA 12.8+ PyTorch).

Claude Code:

claude mcp add bioharbor -- bioharbor serve

Claude Desktop: bioharbor install-claude --write, then restart the app.

Then ask your agent something like:

Find the longest ORF in this contig, translate it, search Swiss-Prot for homologs and predict its structure. Which regions are low confidence?

Use it without an agent

Every tool is also a CLI command, with identical behaviour:

bioharbor tools list
bioharbor run find_orfs sequence=@contig.fa min_aa=100
bioharbor run seq_stats sequence=MKTAYIAKQRQISFVKSHFSRQ
bioharbor jobs

Shared GPU server

Run one BioHarbor on the lab GPU box and point everyone's agent at it:

bioharbor serve --http --host 0.0.0.0 --port 8765   # MCP endpoint: http://<host>:8765/mcp

⚠️ Authentication for HTTP mode is on the roadmap; until then expose it only on a trusted network or behind an SSH tunnel.

Tools

Tool What it does Runs
seq_stats Validate sequences; type, length, GC%, molecular weight inline
translate_sequence DNA/RNA → protein, one or all six frames inline
find_orfs Longest ORFs on both strands, with coordinates inline
search_homologs MMseqs2 search (protein, or translated DNA) vs local DBs job
predict_structure ESMFold structure, pLDDT bands, low-confidence regions, pTM job (GPU)
scrna_pipeline scanpy QC → clustering → markers planned

Runtime tools: get_job, list_jobs, cancel_job, describe_tool, list_databases, gpu_status, read_file.

How it works

Agent ──MCP──▶ validate input ─▶ inline? ──yes──▶ run ─┐
                                   │ no                  ├─▶ provenance + summary ─▶ Agent
                                   ▼                     │
                     job queue (SQLite) ─▶ GPU placement ┘
                     (waits politely for a GPU with free memory)
  • Every call is a job recorded in SQLite with params, versions, timings and GPU used, plus a provenance.json next to its outputs.
  • GPU placement reads live free memory and utilisation (NVML or nvidia-smi), keeps headroom, and reserves memory for jobs it has started so two jobs never grab the same space. Other users' processes are respected.
  • Fail fast: input, binaries and databases are checked before a job is queued, so a bad request never waits behind a busy GPU.
  • Results are agent-shaped: summary, message, files, suggestions. Errors carry a hint and a retryable flag.

Details: docs/design.md.

Writing a tool

from pydantic import BaseModel, Field
from bioharbor.registry import Resources, RunContext, tool
from bioharbor.results import ToolResult


class FoldParams(BaseModel):
    sequence: str = Field(..., description="Protein sequence")


@tool(
    version="1",
    slow=True,
    resources=Resources(gpu=True, gpu_mem_gb=lambda p: 4 + len(p.sequence) / 100),
)
def predict_structure(params: FoldParams, ctx: RunContext) -> ToolResult:
    """Predict a protein structure with ESMFold."""
    ...
    return ToolResult(summary={"mean_plddt": 87.1}, files=["model.pdb"])

Plugins can ship tools in their own package via the bioharbor.tools entry-point group. See CONTRIBUTING.md.

License

Apache-2.0

Metadata

Release files for bioharbor 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for bioharbor 0.1.0
File Size Uploaded
bioharbor-0.1.0.tar.gz 45.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for bioharbor 0.1.0
File Interpreter ABI Platform
bioharbor-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 86.4 kB

Release files / bioharbor-0.1.0.tar.gz

Download URL bioharbor-0.1.0.tar.gz
Size 45.1 kB
Tags Source
SHA-256 checksum
How to use checksums
edad8de24fe8efbdf0b07cc8b4583256a756635205be21a3687b3416f74ecca3
BLAKE2b-256 checksum
How to use checksums
168875dd09860b655b6e9362b7b44a8be4279daa93cd7805f7c1271c51c3145b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release files / bioharbor-0.1.0-py3-none-any.whl

Download URL bioharbor-0.1.0-py3-none-any.whl
Size 41.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f6a0e6cfa145099c4f205ba11c7b7bc4f603a07d56f0450c07cc742b191a56b4
BLAKE2b-256 checksum
How to use checksums
9f4db898ffd58d8f45259735f06788e7c7957295f02f7c88920e76a4de0a020f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page