Skip to main content

Tests Coverage Status PyPI

mhcgnomes: Parsing MHC nomenclature in the wild

Documentation site: https://pirl-unc.github.io/mhcgnomes/

MHCgnomes is a parsing library for multi-species MHC nomenclature which aims to correctly parse every name in IEDB, IMGT/HLA, IPD/MHC, and the allele lists for both NetMHCpan and NetMHCIIpan predictors. This allows for standardization between immune databases and tools, which often use different naming conventions.

Usage example

In [1]: mhcgnomes.parse("HLA-A0201")
Out[1]: Allele(
    gene=Gene(
        species=Species(name="Homo sapiens", mhc_prefix="HLA"),
        name="A"),
    allele_fields=("02", "01"),
    annotations=(),
    mutations=())

In [2]: mhcgnomes.parse("HLA-A0201").to_string()
Out[2]: 'HLA-A*02:01'

In [3]: mhcgnomes.parse("HLA-A0201").compact_string()
Out[3]: 'A0201'

The problem: MHC nomenclature is nuts

Despite the valiant efforts of groups such as the Comparative MHC Nomenclature Committee, the names of MHC alleles you might encounter in different datasets (or accepted by immunoinformatics tools) are frustratingly ill-specified. It's not uncommon to see dozens of different forms for the same allele.

For example, these all refer to the same MHC protein sequence:

  • "HLA-A*02:01"
  • "HLA-A02:01"
  • "HLA-A:02:01"
  • "HLA-A0201"

Additionally, for human alleles, the species prefix is often omitted:

  • "A*02:01"
  • "A*0201"
  • "A02:01"
  • "A:02:01"
  • "A0201"

Annotations

Sometimes, alleles are bundled with modifier suffixes which specify the functionality or abundance of the MHC. Here's an example with an allele which is secreted instead of membrane-bound:

  • "HLA-A*02:01:01S"

These are collected in the annotations field of an Allele result.

Multi-letter annotations are also used in some non-human systems. In particular, Ps (pseudogene) and Sp (splice variant) appear as suffixes on allele fields, e.g. Mamu-B*074:03Sp or Caja-B5*01:01Ps, and are parsed into the annotations field as Sp or Ps respectively.

Note that Ps can also appear as part of a gene name (prefix or suffix) in non-human primates, such as Caja-G2Ps*01. In those cases Ps is treated as part of the gene name, not an allele annotation.

The suffix-specific annotation_pseudogene property remains separate from species/gene ontology status. Gene.pseudogene_status is True, False, or None when no status has been curated; Gene.is_pseudogene is true only for a curated pseudogene. Allele.is_pseudogene combines gene-level status with an explicit Ps suffix:

>>> mhcgnomes.parse("HLA-H*02:01").is_pseudogene
True
>>> mhcgnomes.parse("Mamu-G*01:01").is_pseudogene
True
>>> mhcgnomes.parse("HLA-G*01:01").is_pseudogene
False
>>> mhcgnomes.parse("Caja-B5*01:01Ps").annotation_pseudogene
True

Mutations

MHC proteins are sometimes described in terms of mutations to a known allele.

  • "HLA-B*08:01 N80I mutant"

These mutations are collected in the mutations field of an Allele result.

Beyond humans

To make things worse, several model organisms (like mice and rats) use archaic naming systems, where there is no notion of allele groups or four/six/eight digit alleles but every allele is simply given a name, such as:

  • "H2-Kk"
  • "RT1-9.5f"

In the above example "H2"/"RT1" correspond to species, "K"/"9.5" are the gene names and "k"/"f" are the allele names.

To make these even worse, the name of a species is subject to variation (e.g. "H2" vs. "H-2") as well as drift over time (e.g. ChLA -> MhcPatr -> Patr).

Serotypes, supertypes, haplotypes, and other named entities

Besides alleles there are also other named MHC related entities you'll encounter in immunological data. Closely related to alleles are serotypes, which effectively denote a grouping of alleles that are all recognized by the same antibody:

  • "HLA-A2"
  • "A2"

Supertypes are functional groupings based on shared peptide-binding specificity rather than serological reactivity (Sidney et al. 2008). These are parsed when the "supertype" keyword is present:

  • "A2 supertype"
  • "HLA-B44 supertype"

Class II heterodimers can be specified using dot notation, which is common in celiac disease literature:

  • "DQ2.5" (equivalent to DQA1*05:01/DQB1*02:01)
  • "DQ8.5"

In many datasets the exact allele is not known but an experiment might note the genetic background of a model animal, resulting in loose haplotype restrictions such as:

  • "H2-k class I"

Yes, good luck disambiguating "H2-k" the haplotype from "H2-K" the gene, especially since capitalization is not stable enough to be relied on for parsing.

In some cases immunological data comes only with a denoted species (e.g. "mouse"), a gene (e.g. "HLA-A"), or an MHC class ("human class I"). MHCgnomes has a structured representation for all of these cases and more.

CLI

After installation, a mhcgnomes CLI is available:

mhcgnomes "HLA-A*02:01" "DQ2.5"
# or:
python -m mhcgnomes "HLA-A*02:01" "DQ2.5"

This prints a table with:

  • input string
  • parsed result type
  • normalized and compact forms
  • species/gene/MHC class
  • parsed properties from to_record()

You can also use machine-friendly output:

mhcgnomes --format tsv "HLA-A*02:01" "HLA-A2"
mhcgnomes --format json "HLA-A*02:01" "not a real allele"

By default, unparseable values are shown as ParseError rows. Use strict mode to fail fast:

mhcgnomes --strict "not a real allele"

Parsing strategy

It is a fool's errand to curate all possible MHC allele names since that list grows daily as the MHC loci of more people (and non-human animals) are sequenced. Instead, MHCgnomes contains an ontology of curated species and genes and then attempts to parse any given string into multiple candidates of the following types:

The set of candidate interpretations for each string are then ranked according to heuristic rules. For example, a string will be preferentially interpreted as an Allele rather than a Serotype or Haplotype.

Parsing untrusted input

parse accepts anything that looks like MHC nomenclature, which is the right default for a known-good allele list and the wrong one for free text. Scanning curator notes or a spreadsheet column, "it parsed" is not the same as "this is an MHC molecule": a stray HLA class I is a valid MhcClass, and a bare d is a valid mouse Haplotype.

Say what you will accept, and let anything else come back as None:

from mhcgnomes import parse, Allele, Gene, Pair

result = parse(token, required_result_types=(Allele, Gene, Pair), raise_on_error=False)
if result is None:
    continue  # not an MHC molecule
HLA-A*02:01                -> Allele        n/a                 -> None
H2-K*b                     -> Allele        -                   -> None
Patr-AL                    -> Gene          HLA class I         -> None
HLA-DPB1*06:01/DPA1*01:03  -> Pair          d                   -> None

This is more robust than filtering on the returned type yourself, since it also prevents a lower-ranked interpretation from being chosen in the first place.

Which result type means what

Several result types are legitimate MHC designations without being molecules, which is why HLA-DR15 and BoLA-DR parse cleanly but yield no allele:

Type Means Example
Allele A specific allele HLA-A*02:01
Gene A locus, no allele given Patr-AL
Pair Class II alpha/beta pairing HLA-DPA1*01:03/DPB1*06:01
AlleleWithoutGene Allele name whose gene is unknown BoLA-D18.4
Class2Locus A class II locus family BoLA-DR
Serotype Serological group, may cover many alleles HLA-DR15
Haplotype Named haplotype H2-b
Supertype Functional grouping of alleles HLA-A02
MhcClass Just a class, with or without species HLA class I
Species Only a species was named HLA

Was the species explicit, or guessed?

Species inference is right most of the time and load-bearing when it is wrong. A bare gene symbol can carry a species you never supplied, and the default species means a deliberately generic string still comes back with one:

>>> parse("Gaga-BLB2*02").species_source
'explicit'
>>> parse("BLB2*02").species_source     # inferred from the gene name
'inferred'
>>> parse("MHC class II").species_source  # fell back to default_species
'default'

result.species_from_input is the boolean form, true only for explicit — it answers the question a caller actually has: did I supply this species, or did the parser? To reject anything you did not name outright — the right rule when validating curated data — ask for it at parse time:

>>> parse("BLB2*02", require_explicit_species=True, raise_on_error=False) is None
True

Provenance is not part of a result's identity: two alleles that differ only in how their species was determined still compare equal and hash the same. It is also computed only when asked, so it costs nothing if you never look.

How many digits per field?

Originally alleles for many genes were numbered with two digits:

  • "HLA-MICB*01"

But as the number of identified alleles increased, the number of fields specifying a distinct protein increased to two. This became conventionally called a "four digit" format, since each field has two digits. Yet, as the number of identified alleles continued to increase, the number of digits per field has often increased from two to three:

  • "MICB*002:01"
  • "HLA-A00201"
  • "A:002:01"
  • "A*00201"

MHCgnomes normalizes allele field widths by zero-padding to each gene's canonical minimum (e.g. 3 digits for MICA/MICB). Coverage of per-gene field widths is still incomplete for some non-human species.

However, if databases such as IPD-MHC or IMGT-HLA recorded an older form of an allele, then MHCgnomes can optionally map it onto the modern version (including capturing differences in numbers of digits per field).

Species-directed parsing

species= constrains parsing to a single species. The final parsed object must match that species exactly, or parsing fails. This is useful when you know the organism and want to reject cross-species mismatches:

>>> mhcgnomes.parse("BoLA-DRB3*01:01", species="Bos taurus").to_string()
'Bota-DRB3*01:01'
>>> mhcgnomes.parse("HLA-A*02:01", species="Bos taurus", raise_on_error=False) is None
True
>>> mhcgnomes.parse("A*02:01", species="Homo sapiens").species.name
'Homo sapiens'

When the input uses an ancestor prefix (like BoLA for genus-level Bos sp.), species= rewrites the result to the requested descendant species if valid.

default_species= is a less strict alternative — it provides a fallback species hint for inputs that don't contain a species prefix, but does not reject inputs that resolve to a different species:

>>> mhcgnomes.parse("A*02:01", default_species="Homo sapiens").species.name
'Homo sapiens'
>>> mhcgnomes.parse("DMA", default_species="Chelonia mydas").species.name
'Chelonia mydas'

Species and gene ontology

MHCgnomes maintains a curated ontology of species prefixes and MHC gene names in YAML data files under mhcgnomes/data/. The key files are:

File Purpose
species.yaml Canonical species entries with MHC prefixes, genes, classes, properties, and families
gene_aliases.yaml Alternative gene spellings that normalize to canonical genes
allele_aliases.yaml Retired or shorthand allele names that normalize to canonical alleles
known_alleles.yaml Curated known allele labels per species/gene

The species tree

Species entries form a tree through their parent links, and it is nomenclature-shaped rather than taxonomic. Two things follow that are easy to get wrong:

Gnathostomata sp. is a universal root. Every species descends from it, so "do these two share a common ancestor?" is true for any pair and useless as a compatibility test — worse, it fails open. Only a direct ancestor or descendant relation is meaningful. Species.compatible_with implements that rule:

>>> Species.get("Bos taurus").compatible_with("Bos sp.")
True                     # less specific, not contradictory
>>> Species.get("Macaca mulatta").compatible_with("Macaca fascicularis")
False                    # siblings

This is what you want when checking a declared species against the species mhcgnomes derives from an allele, since the two legitimately differ in specificity: BoLA belongs to the genus-level Bos sp., so a BoLA-N*013:01 allele on a sample curated as Bos taurus is compatible.

The tree is shallow where a species owns its own MHC system. Homo sapiens attaches directly to the root, so Primata sp. is not an ancestor of human, while Saimiri sciureus -> Primata sp. -> Gnathostomata sp. is. HLA is its own system rather than a generic primate one, and the tree reflects that. A caller reasoning purely taxonomically will guess wrong.

An X sp. node is a group entry: it holds the prefix and the genes shared by everything beneath it. Bos sp. owns BoLA and the cattle gene list; Bos taurus inherits both.

Species prefix conventions

Each species is identified by a short prefix (usually 2-4 characters) such as HLA (human), H2 (mouse), Gaga (chicken), or Dare (zebrafish). The parser uses these prefixes to identify species before parsing gene names and allele fields.

Prefixes are matched case-insensitively after stripping punctuation. A leading Mhc prefix (common in bird MHC literature, e.g. MhcTyal-DAB1*01:01) is automatically stripped as a fallback when normal prefix matching fails.

Some historically important prefixes are not single-species codes. Prefixes such as DLA, SLA, OLA, BoLA, and CELA are curated as umbrella taxon nodes in the ontology because the external nomenclature itself is genus- or clade-level rather than species-specific. For example:

  • DLA maps to Canis sp., while Calu maps specifically to Canis lupus
  • SLA maps to Sus sp., while Susc maps specifically to Sus scrofa
  • BoLA maps to Bos sp., while Bota maps specifically to Bos taurus
  • OLA maps to Ovis sp., while Ovar maps specifically to Ovis aries
  • CELA maps to Cetacea sp., while Tutr maps specifically to Tursiops truncatus

This distinction matters when interpreting parsed objects: an allele parsed from BoLA-... is attached to the generic cattle node unless the parse is explicitly constrained or rewritten to a descendant species.

MHC gene class assignments

Genes in species.yaml are organized by MHC class:

  • Ia: Classical class I (associates with B2M, presents peptides)
  • Ib: Non-classical class I (in MHC locus, associates with B2M)
  • Ic: Related MHC locus genes, no B2M association (e.g. MICA)
  • Id: Class I-related genes on other chromosomes
  • IIa: Classical class II alpha/beta chains presenting peptides
  • IIb: Accessory or non-classical class II proteins
  • other: Antigen processing genes (TAP1, TAP2, TAPBP, B2M)

Species-specific gene properties and families

species.yaml can attach source-backed properties to canonical genes. These properties inherit through the species tree and descendants can override them, so the same gene name may have different biology in different lineages. For example, human HLA-G is explicitly functional while MHC-G is a pseudogene in the macaque lineage. Missing metadata remains unknown rather than being inferred from the spelling of a gene.

The ontology can also define exact gene families. The jawed-vertebrate root defines TAP as the family containing TAP1 and TAP2. The strict parser does not promote TAP to a gene, but the species-aware parse_gene_class API can classify legacy family-level inputs without choosing a family member:

>>> info = mhcgnomes.parse_gene_class("SLA-TAP*1*01:01")
>>> (info.gene_name, info.mhc_class, info.non_mhc, info.source)
('TAP', 'other', True, 'ontology_family')

This classification uses the selected species' inherited ontology family; it does not infer TAP membership from a regex or a shared gene-name pattern.

Species prefix tiers

As mhcgnomes supports more species, short prefix codes increasingly collide. Codes like HLA/SLA/DLA, OrLA, and four-letter codes like Calu all hit collisions as coverage grows. We support multiple prefix tiers so that every species is always parseable:

Tier Form Example When used
Established short prefix 1–4 letters HLA, Gaga, Crpo Published in MHC literature or IPD-MHC. Preferred for display.
Novel 4+4 prefix First 4 of genus + first 4 of species OryzLati, StruCame Standard display prefix for species without an established literature prefix.
5+5 long prefix First 5 of genus + first 5 of species HomoSapie, OryziLatip Auto-generated alias for all binomial species. Always parseable.
Full latin name Concatenated genus + species HomoSapiens, ChrysemysPicta Always parseable as an alternative. Guaranteed collision-free.

All tiers are parsed case-insensitively. For example, these all parse to the same allele:

HLA-A*02:01          # established prefix
HomoSapi-A*02:01     # 4+4 novel prefix (auto-generated alias)
HomoSapie-A*02:01    # 5+5 long prefix (auto-generated alias)
HomoSapiens-A*02:01  # full latin name
Homo sapiens-A*02:01 # latin name with space

The 8-letter (4+4) novel prefix space greatly reduces collision probability compared to 4-letter codes, but only the full latin name is truly guaranteed to be unique. Since we don't yet know what naming conventions the scientific community will settle on for newer taxa, we support all tiers simultaneously.

Which prefixes are established vs generated: Comments in species.yaml document which prefixes are attested in MHC literature and which were generated by mhcgnomes. Established prefixes are never changed; generated prefixes are subject to replacement if a community convention emerges.

See the Curation Guide for the full prefix conflict resolution policy (source).

References

Development

Raw IPD-IMGT/HLA and IPD-MHC snapshots are not committed or bundled in releases. See External IMGT/IPD data for checksum-pinned downloads, local/CI caches, optional mirrors, offline use, and alias regeneration.

Local docs

./develop.sh
mkdocs serve
mkdocs build --strict

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mhcgnomes-3.38.0.tar.gz (183.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mhcgnomes-3.38.0-py3-none-any.whl (201.7 kB view details)

Uploaded Python 3

File details

Details for the file mhcgnomes-3.38.0.tar.gz.

File metadata

  • Download URL: mhcgnomes-3.38.0.tar.gz
  • Upload date:
  • Size: 183.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.6

File hashes

Hashes for mhcgnomes-3.38.0.tar.gz
Algorithm Hash digest
SHA256 44edb2bc8af384502beec8fb5dfc1cba6f7b68b3ac823c9f2dadcf18a67e7656
MD5 4bab598c4c5e6e195686f5fec5e29e32
BLAKE2b-256 fba3333b1bd9f7579f913b7dc98801eb8a07e6f51bab1314b55a1ce70a57ab99

See more details on using hashes here.

File details

Details for the file mhcgnomes-3.38.0-py3-none-any.whl.

File metadata

  • Download URL: mhcgnomes-3.38.0-py3-none-any.whl
  • Upload date:
  • Size: 201.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.6

File hashes

Hashes for mhcgnomes-3.38.0-py3-none-any.whl
Algorithm Hash digest
SHA256 4c5403824bdb1cbf5cbadf672a9fec729c572cc3dcd8573989ca3e7dc4bcc3ec
MD5 cfcd2001ff39453a2bcd9d614869c8f2
BLAKE2b-256 a78c4bcb83a4b27cd1440350dcc6c40d1f84de406590c484dcdf28b58d6ac9bb

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

3.38.0 This release

2 files

3.37.0

2 files

3.33.6

2 files

3.33.5

2 files

3.33.4

2 files

3.33.3

2 files

3.33.2

2 files

3.33.1

2 files

3.33.0

2 files

3.32.0

2 files

3.31.1

2 files

3.31.0

2 files

3.30.0

2 files

3.29.1

2 files

3.28.0

2 files

3.27.0

2 files

3.26.0

2 files

3.25.0

2 files

3.24.0

2 files

3.23.0

2 files

3.22.0

2 files

3.21.0

2 files

3.20.0

2 files

3.19.0

2 files

3.18.0

2 files

3.17.1

2 files

3.17.0

2 files

3.16.0

2 files

3.15.0

2 files

3.14.1

2 files

3.14.0

2 files

3.13.0

2 files

3.12.2

2 files

3.12.1

2 files

3.12.0

2 files

3.11.0

2 files

3.10.0

2 files

3.9.2

2 files

3.9.0

2 files

3.8.0

2 files

3.7.0

2 files

3.6.0

2 files

3.5.0

2 files

3.4.0

2 files

3.3.0

2 files

3.2.0

2 files

3.1.1

2 files

3.1.0

2 files

3.0.2

2 files

3.0.1

2 files

3.0.0

2 files

2.0.9

2 files

2.0.8

2 files

2.0.7

2 files

2.0.6

2 files

2.0.5

2 files

2.0.4

2 files

2.0.3

2 files

2.0.2

2 files

2.0

2 files

1.8.6

2 files

1.8.4

2 files

1.7.0

2 files

1.6.1

2 files

1.6.0

2 files

1.5.0

2 files

1.4.0

2 files

1.3.2

2 files

1.3.1

2 files

1.3.0

2 files

1.2.2

2 files

1.2.1

2 files

1.2.0

2 files

1.1.0

2 files

1.0.7

2 files

1.0.6

2 files

1.0.5

2 files

1.0.4

2 files

1.0.3

2 files

1.0.2

2 files

1.0.1

2 files

1.0.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page