mhcgnomes: Parsing MHC nomenclature in the wild
Documentation site: https://pirl-unc.github.io/mhcgnomes/
MHCgnomes is a parsing library for multi-species MHC nomenclature which aims to correctly parse every name in IEDB, IMGT/HLA, IPD/MHC, and the allele lists for both NetMHCpan and NetMHCIIpan predictors. This allows for standardization between immune databases and tools, which often use different naming conventions.
Usage example
In [1]: mhcgnomes.parse("HLA-A0201")
Out[1]: Allele(
gene=Gene(
species=Species(name="Homo sapiens", mhc_prefix="HLA"),
name="A"),
allele_fields=("02", "01"),
annotations=(),
mutations=())
In [2]: mhcgnomes.parse("HLA-A0201").to_string()
Out[2]: 'HLA-A*02:01'
In [3]: mhcgnomes.parse("HLA-A0201").compact_string()
Out[3]: 'A0201'
The problem: MHC nomenclature is nuts
Despite the valiant efforts of groups such as the Comparative MHC Nomenclature Committee, the names of MHC alleles you might encounter in different datasets (or accepted by immunoinformatics tools) are frustratingly ill-specified. It's not uncommon to see dozens of different forms for the same allele.
For example, these all refer to the same MHC protein sequence:
- "HLA-A*02:01"
- "HLA-A02:01"
- "HLA-A:02:01"
- "HLA-A0201"
Additionally, for human alleles, the species prefix is often omitted:
- "A*02:01"
- "A*0201"
- "A02:01"
- "A:02:01"
- "A0201"
Annotations
Sometimes, alleles are bundled with modifier suffixes which specify the functionality or abundance of the MHC. Here's an example with an allele which is secreted instead of membrane-bound:
- "HLA-A*02:01:01S"
These are collected in the annotations field of an
Allele
result.
Multi-letter annotations are also used in some non-human systems. In particular,
Ps (pseudogene) and Sp (splice variant) appear as suffixes on allele fields,
e.g. Mamu-B*074:03Sp or Caja-B5*01:01Ps, and are parsed into the
annotations field as Sp or Ps respectively.
Note that Ps can also appear as part of a gene name (prefix or suffix) in
non-human primates, such as Caja-G2Ps*01. In those cases Ps is treated as
part of the gene name, not an allele annotation.
The suffix-specific annotation_pseudogene property remains separate from
species/gene ontology status. Gene.pseudogene_status is True, False, or
None when no status has been curated; Gene.is_pseudogene is true only for a
curated pseudogene. Allele.is_pseudogene combines gene-level status with an
explicit Ps suffix:
>>> mhcgnomes.parse("HLA-H*02:01").is_pseudogene
True
>>> mhcgnomes.parse("Mamu-G*01:01").is_pseudogene
True
>>> mhcgnomes.parse("HLA-G*01:01").is_pseudogene
False
>>> mhcgnomes.parse("Caja-B5*01:01Ps").annotation_pseudogene
True
Mutations
MHC proteins are sometimes described in terms of mutations to a known allele.
- "HLA-B*08:01 N80I mutant"
These mutations are collected in the mutations field of an
Allele result.
Beyond humans
To make things worse, several model organisms (like mice and rats) use archaic naming systems, where there is no notion of allele groups or four/six/eight digit alleles but every allele is simply given a name, such as:
- "H2-Kk"
- "RT1-9.5f"
In the above example "H2"/"RT1" correspond to species, "K"/"9.5" are the gene names and "k"/"f" are the allele names.
To make these even worse, the name of a species is subject to variation (e.g. "H2" vs. "H-2") as well as drift over time (e.g. ChLA -> MhcPatr -> Patr).
Serotypes, supertypes, haplotypes, and other named entities
Besides alleles there are also other named MHC related entities you'll encounter in immunological data. Closely related to alleles are serotypes, which effectively denote a grouping of alleles that are all recognized by the same antibody:
- "HLA-A2"
- "A2"
Supertypes are functional groupings based on shared peptide-binding specificity rather than serological reactivity (Sidney et al. 2008). These are parsed when the "supertype" keyword is present:
- "A2 supertype"
- "HLA-B44 supertype"
Class II heterodimers can be specified using dot notation, which is common in celiac disease literature:
- "DQ2.5" (equivalent to DQA1*05:01/DQB1*02:01)
- "DQ8.5"
In many datasets the exact allele is not known but an experiment might note the genetic background of a model animal, resulting in loose haplotype restrictions such as:
- "H2-k class I"
Yes, good luck disambiguating "H2-k" the haplotype from "H2-K" the gene, especially since capitalization is not stable enough to be relied on for parsing.
In some cases immunological data comes only with a denoted species (e.g. "mouse"), a gene (e.g. "HLA-A"), or an MHC class ("human class I"). MHCgnomes has a structured representation for all of these cases and more.
CLI
After installation, a mhcgnomes CLI is available:
mhcgnomes "HLA-A*02:01" "DQ2.5"
# or:
python -m mhcgnomes "HLA-A*02:01" "DQ2.5"
This prints a table with:
- input string
- parsed result type
- normalized and compact forms
- species/gene/MHC class
- parsed properties from
to_record()
You can also use machine-friendly output:
mhcgnomes --format tsv "HLA-A*02:01" "HLA-A2"
mhcgnomes --format json "HLA-A*02:01" "not a real allele"
By default, unparseable values are shown as ParseError rows.
Use strict mode to fail fast:
mhcgnomes --strict "not a real allele"
Parsing strategy
It is a fool's errand to curate all possible MHC allele names since that list grows daily as the MHC loci of more people (and non-human animals) are sequenced. Instead, MHCgnomes contains an ontology of curated species and genes and then attempts to parse any given string into multiple candidates of the following types:
The set of candidate interpretations for each string are then
ranked according to heuristic rules. For example, a string will be
preferentially interpreted as an Allele rather
than a Serotype
or Haplotype.
Parsing untrusted input
parse accepts anything that looks like MHC nomenclature, which is the right
default for a known-good allele list and the wrong one for free text. Scanning
curator notes or a spreadsheet column, "it parsed" is not the same as "this is
an MHC molecule": a stray HLA class I is a valid MhcClass, and a bare d
is a valid mouse Haplotype.
Say what you will accept, and let anything else come back as None:
from mhcgnomes import parse, Allele, Gene, Pair
result = parse(token, required_result_types=(Allele, Gene, Pair), raise_on_error=False)
if result is None:
continue # not an MHC molecule
HLA-A*02:01 -> Allele n/a -> None
H2-K*b -> Allele - -> None
Patr-AL -> Gene HLA class I -> None
HLA-DPB1*06:01/DPA1*01:03 -> Pair d -> None
This is more robust than filtering on the returned type yourself, since it also prevents a lower-ranked interpretation from being chosen in the first place.
Which result type means what
Several result types are legitimate MHC designations without being molecules,
which is why HLA-DR15 and BoLA-DR parse cleanly but yield no allele:
| Type | Means | Example |
|---|---|---|
Allele |
A specific allele | HLA-A*02:01 |
Gene |
A locus, no allele given | Patr-AL |
Pair |
Class II alpha/beta pairing | HLA-DPA1*01:03/DPB1*06:01 |
AlleleWithoutGene |
Allele name whose gene is unknown | BoLA-D18.4 |
Class2Locus |
A class II locus family | BoLA-DR |
Serotype |
Serological group, may cover many alleles | HLA-DR15 |
Haplotype |
Named haplotype | H2-b |
Supertype |
Functional grouping of alleles | HLA-A02 |
MhcClass |
Just a class, with or without species | HLA class I |
Species |
Only a species was named | HLA |
Was the species explicit, or guessed?
Species inference is right most of the time and load-bearing when it is wrong. A bare gene symbol can carry a species you never supplied, and the default species means a deliberately generic string still comes back with one:
>>> parse("Gaga-BLB2*02").species_source
'explicit'
>>> parse("BLB2*02").species_source # inferred from the gene name
'inferred'
>>> parse("MHC class II").species_source # fell back to default_species
'default'
result.species_from_input is the boolean form, true only for explicit — it
answers the question a caller actually has: did I supply this species, or did
the parser? To reject anything you did not name outright — the right rule when
validating curated data — ask for it at parse time:
>>> parse("BLB2*02", require_explicit_species=True, raise_on_error=False) is None
True
Provenance is not part of a result's identity: two alleles that differ only in how their species was determined still compare equal and hash the same. It is also computed only when asked, so it costs nothing if you never look.
How many digits per field?
Originally alleles for many genes were numbered with two digits:
- "HLA-MICB*01"
But as the number of identified alleles increased, the number of fields specifying a distinct protein increased to two. This became conventionally called a "four digit" format, since each field has two digits. Yet, as the number of identified alleles continued to increase, the number of digits per field has often increased from two to three:
- "MICB*002:01"
- "HLA-A00201"
- "A:002:01"
- "A*00201"
MHCgnomes normalizes allele field widths by zero-padding to each gene's canonical minimum (e.g. 3 digits for MICA/MICB). Coverage of per-gene field widths is still incomplete for some non-human species.
However, if databases such as IPD-MHC or IMGT-HLA recorded an older form of an allele, then MHCgnomes can optionally map it onto the modern version (including capturing differences in numbers of digits per field).
Species-directed parsing
species= constrains parsing to a single species. The final parsed object
must match that species exactly, or parsing fails. This is useful when you
know the organism and want to reject cross-species mismatches:
>>> mhcgnomes.parse("BoLA-DRB3*01:01", species="Bos taurus").to_string()
'Bota-DRB3*01:01'
>>> mhcgnomes.parse("HLA-A*02:01", species="Bos taurus", raise_on_error=False) is None
True
>>> mhcgnomes.parse("A*02:01", species="Homo sapiens").species.name
'Homo sapiens'
When the input uses an ancestor prefix (like BoLA for genus-level Bos sp.),
species= rewrites the result to the requested descendant species if valid.
default_species= is a less strict alternative — it provides a fallback
species hint for inputs that don't contain a species prefix, but does not
reject inputs that resolve to a different species:
>>> mhcgnomes.parse("A*02:01", default_species="Homo sapiens").species.name
'Homo sapiens'
>>> mhcgnomes.parse("DMA", default_species="Chelonia mydas").species.name
'Chelonia mydas'
Species and gene ontology
MHCgnomes maintains a curated ontology of species prefixes and MHC gene names
in YAML data files under mhcgnomes/data/. The key files are:
| File | Purpose |
|---|---|
species.yaml |
Canonical species entries with MHC prefixes, genes, classes, properties, and families |
gene_aliases.yaml |
Alternative gene spellings that normalize to canonical genes |
allele_aliases.yaml |
Retired or shorthand allele names that normalize to canonical alleles |
known_alleles.yaml |
Curated known allele labels per species/gene |
The species tree
Species entries form a tree through their parent links. A parent is a
containment claim — "this species is inside that group" — and it is taxonomic
wherever it can be: exactly one parent link in the whole ontology points at
another genus's node. Genes, and the rest of a species' curated data, are
inherited along these links.
Every MHC prefix owns a node, and an umbrella prefix covers everything beneath its node:
Bos sp. [BoLA] -> Bota, Boin, Bofr, Bogr, Bubu
Macaca sp. [RhLA] -> Mafa, Mamu, Mane, Masi, Math
Canis sp. [DLA] -> Calu, Cala, Caru, Casi, Caba
NHP is a node, not the primate order. IPD-MHC's NHP group is Non-Human
Primates, so it is the primate order
minus humans — paraphyletic, and not a taxon. It gets its own entry, sibling
to Homo sapiens:
Primata sp. [Primata] the primate order, and the genes all primates share
├── Homo sapiens [HLA]
└── NHP [NHP] the IPD-MHC group
└── 55 non-human primates
That keeps both questions as plain ancestry, with no second predicate:
>>> Species.get("Homo sapiens").compatible_with("Primata sp.")
True # humans are primates
>>> Species.get("Homo sapiens").compatible_with("NHP")
False # NHP-* cannot denote one
>>> parse("NHP-E*01:01", species="Homo sapiens", raise_on_error=False)
None # so this is refused, not converted
Because NHP is paraphyletic, any taxon added under Primata sp. has to sit
wholly inside or wholly outside it. A Homo sp. node holding H. sapiens and
H. neanderthalensis is fine as a sibling of NHP; a Hominidae sp. spanning
humans and the great apes is not, since it would have to pull Gorilla sp.,
Pan sp. and Pongo sp. out of the NHP umbrella.
Bubalus bubalis is under Bos sp. despite being a different genus. This
is the one edge that is deliberately not taxonomy. Water buffalo belongs to
Bubalus, a sister genus of Bos within Bovini, but IPD-MHC files water
buffalo in the BoLA group and the literature assigns buffalo class II sequences
to cattle loci by trans-species polymorphism — Bubu-DRB is the orthologue of
BoLA-DRB3. The edge is load-bearing: the entry declares only DQA, DQA1
and DQB itself, so Bubu-DRA, Bubu-DRB3, Bubu-DQA2 and Bubu-DQB1 all
parse by inheritance, and Bubu-DRB normalizes to Bubu-DRB3 because of it.
So before "correcting" a parent link that looks taxonomically wrong, check what parses through it.
Two further consequences are easy to get wrong:
Gnathostomata sp. is a universal root. Every species descends from it, so
"do these two share a common ancestor?" is true for any pair and useless as a
compatibility test — worse, it fails open. Only a direct ancestor or descendant
relation is meaningful. Species.compatible_with implements that rule:
>>> Species.get("Bos taurus").compatible_with("Bos sp.")
True # less specific, not contradictory
>>> Species.get("Macaca mulatta").compatible_with("Macaca fascicularis")
False # siblings
This is what you want when checking a declared species against the species
mhcgnomes derives from an allele, since the two legitimately differ in
specificity: BoLA belongs to the genus-level Bos sp., so a BoLA-N*013:01
allele on a sample curated as Bos taurus is compatible.
Compatibility can follow an edge that exists for naming. A Bubu-* allele
is compatible with a curated Bos sp., since Bos sp. denotes the BoLA group
and water buffalo is in it. That is the intended answer, not a wart.
An X sp. node is a group entry: it holds the prefix and the genes shared by
everything beneath it. Bos sp. owns BoLA and the cattle gene list; Bos taurus inherits both.
Species prefix conventions
Each species is identified by a short prefix (usually 2-4 characters) such as
HLA (human), H2 (mouse), Gaga (chicken), or Dare (zebrafish). The
parser uses these prefixes to identify species before parsing gene names and
allele fields.
Prefixes are matched case-insensitively after stripping punctuation. A leading
Mhc prefix (common in bird MHC literature, e.g. MhcTyal-DAB1*01:01) is
automatically stripped as a fallback when normal prefix matching fails.
Some historically important prefixes are not single-species codes. Prefixes
such as DLA, SLA, OLA, BoLA, and CELA are curated as umbrella taxon
nodes in the ontology because the external nomenclature itself is genus- or
clade-level rather than species-specific. For example:
DLAmaps toCanis sp., whileCalumaps specifically toCanis lupusSLAmaps toSus sp., whileSuscmaps specifically toSus scrofaBoLAmaps toBos sp., whileBotamaps specifically toBos taurusOLAmaps toOvis sp., whileOvarmaps specifically toOvis ariesCELAmaps toCetacea sp., whileTutrmaps specifically toTursiops truncatus
This distinction matters when interpreting parsed objects: an allele parsed
from BoLA-... is attached to the generic cattle node unless the parse is
explicitly constrained or rewritten to a descendant species.
MHC gene class assignments
Genes in species.yaml are organized by MHC class:
- Ia: Classical class I (associates with B2M, presents peptides)
- Ib: Non-classical class I (in MHC locus, associates with B2M)
- Ic: Related MHC locus genes, no B2M association (e.g. MICA)
- Id: Class I-related genes on other chromosomes
- IIa: Classical class II alpha/beta chains presenting peptides
- IIb: Accessory or non-classical class II proteins
- other: Antigen processing genes (TAP1, TAP2, TAPBP, B2M)
Species-specific gene properties and families
species.yaml can attach source-backed properties to canonical genes. These
properties inherit through the species tree and descendants can override them,
so the same gene name may have different biology in different lineages. For
example, human HLA-G is explicitly functional while MHC-G is a pseudogene in
the macaque lineage. Missing metadata remains unknown rather than being inferred
from the spelling of a gene.
The ontology can also define exact gene families. The jawed-vertebrate root
defines TAP as the family containing TAP1 and TAP2. The strict parser does
not promote TAP to a gene, but the species-aware parse_gene_class API can
classify legacy family-level inputs without choosing a family member:
>>> info = mhcgnomes.parse_gene_class("SLA-TAP*1*01:01")
>>> (info.gene_name, info.mhc_class, info.non_mhc, info.source)
('TAP', 'other', True, 'ontology_family')
This classification uses the selected species' inherited ontology family; it does not infer TAP membership from a regex or a shared gene-name pattern.
Species prefix tiers
As mhcgnomes supports more species, short prefix codes increasingly collide.
Codes like HLA/SLA/DLA, OrLA, and four-letter codes like Calu all hit
collisions as coverage grows. Three forms are supported so that every species
is always reachable:
| Tier | Form | Example | When used |
|---|---|---|---|
| Established short prefix | 1–4 letters | HLA, Gaga, Crpo |
Published in MHC literature or IPD-MHC. Preferred for display. |
| Full latin name | Concatenated genus + species | HomoSapiens, ChrysemysPicta |
The default generated form. Collision-free across every binomial in the ontology. |
| 4+4 shorthand | First 4 of genus + first 4 of species | TachAcul, AbraBram |
Compact shorthand, and the curated prefix of most species without a literature code. Emitted as an alias only where globally unique. |
All are parsed case-insensitively. These all resolve to the same allele:
HLA-A*02:01 # established prefix
HomoSapiens-A*02:01 # full latin name
HomoSapi-A*02:01 # 4+4 shorthand (auto-generated alias)
Homo sapiens-A*02:01 # latin name with space
Two limits are worth knowing:
4+4 is not collision-free. Three forms are derivable from two species
each — ChryPict from both Chrysemys picta and Chrysolophus pictus,
LaniColl from two shrikes, LeucLeuc from a dace and a crane. None is ever
auto-generated, but that only suppresses the generated copy: each is curated
by one side of its pair, so it still resolves, confidently and to that side
alone — ChryPict gives the turtle, LaniColl gives Lanius collurio,
LeucLeuc gives the crane. Issue #134 tracks whether they should instead be
demoted to context-only. Two of the three are still the canonical prefix of
their owner; the entries that had a contested form as their canonical prefix
now use the concatenated binomial instead — three entries did, in 3.42.0:
| species | was | now |
|---|---|---|
| Chrysemys picta | ChryPict |
ChrysemysPicta |
| Lanius collaris | LaniCola |
LaniusCollaris |
| Leuciscus leuciscus | LeucisLeucis |
LeuciscusLeuciscus |
Their normalized output changes accordingly, so parse("ChryPict-UA") now
round-trips as ChrysemysPicta-UA. The old spellings stay parseable.
Subspecies mint no generated alias. A trinomial entry such as Canis lupus
baileyi deliberately does not claim CanisLupus, which belongs to its parent
binomial, and no third-token variant is generated. Subspecies are reached by
their curated prefix (Caba) or by latin name.
A 5+5 form (HomoSapie) existed up to 3.41.0 as a leftover of the pre-v3.12
scheme, and was removed in 3.42.0: no bundled corpus name, and no species token
in the sibling mhcseqs dataset, ever used one. See issue #128.
This is a breaking parse change. 596 auto-generated 5+5 aliases stop
resolving, and so do ten that had been curated by hand as other prefixes
when the scheme was introduced:
MonopAlbus GaviaGange CaimaCroco CaimaLatir CaretCaret
ChrysPicta CasuaCasua CyaniCaeru PhasiColch CycluCarin
Anything written with a 5+5 form should move to the concatenated binomial —
HomoSapie-A*02:01 becomes HomoSapiens-A*02:01.
Where a prefix came from
Species.prefix_provenance says how an entry came by its prefix:
| value | meaning |
|---|---|
"designated" |
Published nomenclature — IPD-MHC or IMGT/HLA writes alleles with it. Curated with a source in species.yaml; never inferred. |
"generated" |
mhcgnomes derived it from the latin name. Provable by re-deriving it. |
"group label" |
Names a grouping rather than a species, and is never written on an allele: Aves, Galliformes, NHP. |
None |
Not established. |
>>> Species.get("HLA").prefix_provenance
'designated'
>>> Species.get("TachAcul").prefix_provenance
'generated'
>>> Species.get("NHP").prefix_provenance
'group label'
None is deliberately distinct from "designated": a prefix mhcgnomes did not
generate is not thereby proven to be in published use. Most short prefixes are
still unchecked — see issue #131.
Which prefixes are established vs generated: Comments in species.yaml
document which prefixes are attested in MHC literature and which were generated
by mhcgnomes. Established prefixes are never changed; generated prefixes are
subject to replacement if a community convention emerges.
See the Curation Guide for the full prefix conflict resolution policy (source).
References
- IPD-MHC: nomenclature requirements for the non-human major histocompatibility complex in the next-generation sequencing era
- Comparative MHC nomenclature: report from the ISAG/IUIS-VIC committee 2018
- ISAG/IUIS-VIC Comparative MHC Nomenclature Committee report, 2005
- Nomenclature for factors of the SLA system, update 2008
Development
Raw IPD-IMGT/HLA and IPD-MHC snapshots are not committed or bundled in releases. See External IMGT/IPD data for checksum-pinned downloads, local/CI caches, optional mirrors, offline use, and alias regeneration.
Local docs
./develop.sh
mkdocs serve
mkdocs build --strict
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file mhcgnomes-3.42.0.tar.gz.
File metadata
- Download URL: mhcgnomes-3.42.0.tar.gz
- Upload date:
- Size: 191.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.12.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b32e5ba34afb18e77a7b9ad7f2b02a4b535b1ecec818dbaa2a0bb0add7612ace
|
|
| MD5 |
7e630a1f50522868c88f74f547d402cb
|
|
| BLAKE2b-256 |
79b1d63580e51a04cf90fc156d9a33114be76639f42e47168fa320848e2bdcdb
|
File details
Details for the file mhcgnomes-3.42.0-py3-none-any.whl.
File metadata
- Download URL: mhcgnomes-3.42.0-py3-none-any.whl
- Upload date:
- Size: 208.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.12.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e6510a7b03eb09802f122b4718d6ba50f9a8a53a2ad023d5a559e9232debdc3e
|
|
| MD5 |
0cf99daa7fee07bad1179e31cf3897a5
|
|
| BLAKE2b-256 |
0e4970db6f198e1c664c33df113adb8008bd2e1532de0dc522cab52f4559b01c
|