Skip to main content

mhcseqs

Self-contained pipeline for downloading, curating, and extracting binding grooves from MHC (Major Histocompatibility Complex) protein sequences.

Install

pip install mhcseqs

For development:

git clone https://github.com/pirl-unc/mhcseqs.git
cd mhcseqs
./develop.sh          # uv pip install -e ".[dev]"
./test.sh             # pytest
./lint.sh             # ruff

Quick start

CLI

# Build all output CSVs (writes to ~/.cache/mhcseqs/)
mhcseqs build

# Build to a specific directory instead
mhcseqs build --output-dir output/

# Look up a specific allele
mhcseqs lookup "HLA-A*02:01"

# Inspect cached downloads and built CSVs (path, size, age)
mhcseqs data list

# Re-download source FASTAs (IMGT/HLA + IPD-MHC publish from a rolling "Latest")
mhcseqs data refresh
mhcseqs build --force-download   # equivalently, rebuild with fresh sources

# Delete cached files
mhcseqs data clear               # source FASTAs
mhcseqs data clear --built       # also remove built CSVs/reports
mhcseqs data clear --built-only  # only built CSVs/reports

# Check version
mhcseqs --version

Python API

import mhcseqs

# Build the database (downloads to ~/.cache/mhcseqs/, only needed once)
paths = mhcseqs.build()  # BuildPaths dataclass

# Look up any allele → AlleleRecord with everything
r = mhcseqs.lookup("HLA-A*02:01")
r.sequence          # full protein (with signal peptide)
r.mature_sequence   # signal peptide removed (computed property)
r.mature_start      # signal peptide length (24 for HLA-A*02:01)
r.groove1           # α1 domain
r.groove2           # α2 domain
r.ig_domain         # α3 Ig-fold
r.tail              # TM + cytoplasmic
r.domains           # typed domain spans
r.domain_architecture
r.domain_spans
r.species_category  # "human"

# Apply mutations (IEDB-style, e.g. "K66A")
m = mhcseqs.lookup("HLA-A*02:01", mutations=["K66A", "D77S"])

Load as a DataFrame

import mhcseqs

# As a DataFrame (full sequence + groove decomposition + metadata)
df = mhcseqs.load_sequences_dataframe()

# Or as a list of dicts (no pandas dependency)
rows = mhcseqs.load_sequences_dict()

Current data summary

All sources (IMGT/HLA, IPD-MHC, UniProt curated references, and 15,860 diverse MHC sequences from UniProt) are merged into a single dataset:

Category Class I Class II Total
human 17,841 8,013 25,854
nhp 4,639 2,486 7,125
murine 980 565 1,545
ungulate 638 1,128 1,766
carnivore 166 318 484
other_mammal 1,015 740 1,755
bird 11,906 6,690 18,596
fish 2,439 7,106 9,545
other_vertebrate 940 1,332 2,272
total 40,564 28,378 68,942

Covering 558+ species. Groove parse success rate on IMGT/IPD-MHC entries: 99.3%.

Structural decomposition

The parser materializes an explicit domain grammar:

Chain Grammar
Class I alpha signal_peptide? -> g_alpha1 -> g_alpha2 -> c1_alpha3 -> transmembrane? -> cytoplasmic_tail?
Class II alpha signal_peptide? -> g_alpha1 -> c1_alpha2 -> transmembrane? -> cytoplasmic_tail?
Class II beta signal_peptide? -> g_beta1 -> c1_beta2 -> transmembrane? -> cytoplasmic_tail?

The exported contiguous sequence fields are:

Column Class I alpha Class II alpha Class II beta
groove1 α1 domain (~80-95 aa typical) α1 domain (~75-95 aa typical) —
groove2 α2 domain (~80-100 aa typical) — β1 domain (~70-100 aa typical)
ig_domain α3 C-like support domain α2 C-like support domain β2 C-like support domain
tail linker + TM + cytoplasmic tail linker + TM + cytoplasmic tail linker + TM + cytoplasmic tail

domain_architecture and domain_spans expose the typed domain grammar directly, for example:

  • class I: signal_peptide>g_alpha1>g_alpha2>c1_alpha3>tail_linker>transmembrane>cytoplasmic_tail
  • class II beta: signal_peptide>g_beta1>c1_beta2>tail_linker>transmembrane>cytoplasmic_tail

How Parsing Works

The parser is alignment-free and holistic. It does not rely on one absolute Cys position to define the mature start.

For each sequence it:

  1. Enumerates all plausible Cys-Cys pairs in the Ig/C-like separation range.
  2. Scores each pair as a candidate G-domain or C-like anchor using fold-topology evidence, especially the Trp41-like signal around c1+14.
  3. Enumerates candidate SP boundaries and whole domain parses, including partial parses when only fragment evidence is available.
  4. Chooses the best full parse using factored multiplicative scoring: three structural claims (SP grammar, domain architecture, completeness) each produce a [0,1] factor. Contradictory evidence in any factor gates the score down multiplicatively, while missing evidence is a softer penalty.

The strongest evidence types are:

  • SP cleavage grammar: hydrophobic h-region, short c-region, von Heijne -3/-1 compatibility, exclusion of impossible -3/-1 property pairs, and mild +1 mature-sequence penalties.
  • Domain-fold grammar: canonical G-domain versus C-like disulfide topology, including the IMGT-style Cys11-Cys74 G-domain signature and the Cys23/Trp41/Cys104 C-like grammar.
  • Class-specific groove boundaries:
    • class I α1/α2 junction motifs
    • class I α2 -> α3 boundary motifs
    • class II α1 -> α2 and β1 -> β2 boundary motifs
  • Soft priors on groove/support-domain lengths and TM support downstream.

The parser handles:

  • full-length proteins with or without signal peptides
  • SP-stripped deposits (mature_start = 0)
  • common fragments:
    • class I exon 2 only -> alpha1_only
    • class I exon 3 only -> alpha2_only
    • class II exon 2-like fragments -> fragment_fallback
  • low-evidence salvage:
    • class I from α3 C-like support only -> inferred_from_alpha3
    • class II beta from β1 groove pair only -> beta1_only_fallback
  • true groove absence / insufficient structural evidence -> missing_groove

Groove Status Values

These are the important parser-facing statuses:

Status Meaning
ok Full decomposition from the main structural grammar
alpha1_only Class I fragment consistent with α1 / exon 2 only
alpha2_only Class I fragment consistent with α2 / exon 3 only
fragment_fallback Short fragment retained as the observable groove half
inferred_from_alpha3 Class I salvage parse using a downstream α3 C-like anchor
beta1_only_fallback Class II beta salvage parse using only the β1 groove pair
missing_groove No recoverable groove architecture from the available evidence
non_classical Non-classical class-I lineage flagged post-parse
short Groove half too short to look functionally peptide-binding

Pipeline-only statuses can still appear in CSV outputs:

Status Meaning
not_applicable Row intentionally excluded from groove functionality, mainly B2M in build outputs

Literature Basis

The parser is built around conserved sequence grammar from the MHC literature:

  • MHC domain organization is more conserved than short local motifs across vertebrates: Primordial Linkage of β2-Microglobulin to the MHC
  • IMGT domain numbering and the G-domain versus C-like disulfide grammar: PMC3913909
  • Classical class-I domain layout and landmarks: PMC2434379
  • Salmonid class-II alpha/beta cysteine topology and lineage-specific extra cysteines: PMC2386828
  • Teleost class-II evolutionary divergence while retaining the same modular architecture: PMC4219347

Signal-peptide logic follows the standard SPase grammar:

Key columns

Column Description
two_field_allele Allele name at two-field resolution
gene MHC gene (e.g., A, DRB1, BF, UA)
mhc_class I or II
chain alpha, beta, or B2M
species Latin binomial from source
species_category One of 9 categories above
source imgt, ipd_mhc, or uniprot
source_id Database accession for provenance
groove_status See table above
is_functional True if groove parsed and not null/pseudogene

Dependencies

  • Python 3.10+
  • mhcgnomes >= 3.41.0 — allele parsing, species provenance, NHP taxonomy, and species-directed gene classification

mhcseqs emits full-binomial species aliases such as HomoSapiens-A*02:01 and MusMusculus-K*b. Its input registry accepts 476 current, historical, and external-database prefix assignments—including all 137 designation tokens in the current 125-organism IPD-MHC taxonomy register—and records an evidence URL for every one. Colliding short codes require explicit species context. mhcseqs does not invent abbreviated species prefixes, and mechanically generated 2+2/4+4/5+5 aliases are rejected unless that exact spelling has external evidence. See the prefix audit. The audit includes a versioned UniProtKB/UniSave provenance snapshot for all 239 records behind the 21 historical pairs that previously lacked raw inputs.

No alignment tools, BLAST, or structure databases are required.

License

Apache 2.0

Metadata

Release files for mhcseqs 2.6.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mhcseqs 2.6.2
File Size Uploaded
mhcseqs-2.6.2.tar.gz 900.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for mhcseqs 2.6.2
File Interpreter ABI Platform
mhcseqs-2.6.2-py3-none-any.whl Python 3 none any Details

Total release size: 1.8 MB

Release files / mhcseqs-2.6.2.tar.gz

Download URL mhcseqs-2.6.2.tar.gz
Size 900.9 kB
Tags Source
SHA-256 checksum
How to use checksums
ce833bff9e1fec1a5cbe18e05bc095306909b4ed5dffdfca8c2464d681af9dd3
BLAKE2b-256 checksum
How to use checksums
39bf3ec2e80bf54f1cff1c4eb27a50ad86c06732045fcab5fd6b63e28989c5f8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.6

Release files / mhcseqs-2.6.2-py3-none-any.whl

Download URL mhcseqs-2.6.2-py3-none-any.whl
Size 882.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2bcb4c8c9022ca91e5e72478dcc4367cbdd7981b09bb358d6b711592e2dae9b0
BLAKE2b-256 checksum
How to use checksums
8f870bf3a5f78876b55906b32bddcc5ebff5a52e3fac05cc08323cd8427311ac
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.6

Release history Release notifications | RSS feed

2.6.9

2 release files

2.6.6

2 release files

2.6.5

2 release files

2.6.4

2 release files

2.6.3

2 release files

This release

2.6.2 This release

2 release files

2.6.1

2 release files

2.6.0

2 release files

2.5.12

2 release files

2.5.11

2 release files

2.5.9

2 release files

2.5.8

2 release files

2.5.7

2 release files

2.5.6

2 release files

2.5.5

2 release files

2.5.4

2 release files

2.5.3

2 release files

2.5.2

2 release files

2.5.1

2 release files

2.5.0

2 release files

2.4.1

2 release files

2.4.0

2 release files

2.3.6

2 release files

2.3.5

2 release files

2.3.4

2 release files

2.3.3

2 release files

2.3.2

2 release files

2.3.1

2 release files

2.3.0

2 release files

2.2.2

2 release files

2.2.1

2 release files

2.2.0

2 release files

1.2.0

2 release files

1.1.0

2 release files

1.0.0

2 release files

0.11.0

2 release files

0.10.0

2 release files

0.9.0

2 release files

0.8.0

2 release files

0.6.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page