hitlist
A curated, harmonized, ML-training-ready MHC ligand mass-spectrometry dataset.
hitlist ingests immunopeptidome data from IEDB, CEDAR, and paper supplementary tables (PRIDE/jPOSTrepo); partitions MS-eluted observations from in-vitro binding-assay measurements into two separate parquet files (so downstream consumers never silently conflate them); joins every MS observation to expert-curated sample metadata (HLA genotype, tissue, disease, perturbation, instrument); and ships both indexes as parquet + a pandas-friendly Python API.
What's in the two indexes
After hitlist build observations (snapshot of the shipping 1.10.x default build):
observations.parquet — MS-eluted immunopeptidome
| Total observations (MS-eluted, all species) | 4,053,693 |
| Unique peptides | 1,285,987 |
| Unique MHC alleles | 691 |
| MHC species covered | 21 |
| IEDB rows | 3,986,991 |
| CEDAR rows | 595 |
| Supplementary rows | 66,107 |
binding.parquet — in-vitro binding-assay measurements
| Total binding rows (peptide microarray, refolding, MEDi, qualitative-tier) | 895,785 |
| Unique peptides | 258,199 |
The two indexes share the schema (including gene annotations from the peptide-mappings sidecar), but supplementary curation is MS-only — binding is pure IEDB/CEDAR.
Human MHC-I breakdown
| Observations | 2,672,046 |
| Unique peptides | 748,386 |
| Mono-allelic (exact allele) | 579,096 obs / 300K peptides / 119 alleles |
| Multi-allelic with allele match | 450,399 obs |
Multi-allelic with class-pool (N of M alleles) |
784,370 obs |
Allele-resolved sample_mhc coverage |
74.8% |
What's curated
Those numbers rest on per-study human curation. Rather than trust the raw
IEDB/CEDAR cell_name, disease, and tissue fields — which carry mislabels,
free-text placeholders, and Other/Unknown sentinels — hitlist annotates each
study (keyed by PMID) in pmid_overrides.yaml. Known mislabels are corrected,
and every MS sample is tagged with the metadata a training pipeline actually
needs (HLA genotype, tissue, disease, perturbation, instrument) plus a reference
proteome for flanking and source-protein attribution. A few papers whose peptides
never reached IEDB are ingested directly from their PRIDE / jPOSTrepo
supplementary tables.
| What's curated | Count | Detail |
|---|---|---|
Curated PMIDs (pmid_overrides.yaml) |
215 | Cover 96.0% of all observations |
↳ ms_samples with per-sample metadata |
761 | HLA genotype, tissue, disease, instrument, plus the flat condition columns — knockout / cytokine / drug / infection / control role as scalars, not prose |
↳ ms_samples with 4-digit HLA typing |
579 | The subset with allele-level genotype |
| Supplementary CSVs ingested (PRIDE / jPOSTrepo) | 9 | 3 papers — listed below |
| Species reference proteomes | 22 | Ensembl ×4, UniProt ×18 |
| Viral reference proteomes | 33 viruses | 58 name aliases |
Supplementary papers ingested (data not in IEDB, read straight from PRIDE/jPOSTrepo — 9 CSVs total):
- Abelin 2019 — MAPTAC mono-allelic class I, class II, and DR-tissue (3)
- Gómez-Zepeda 2024 — JY, HeLa, Raji, SK-MEL-37, and plasma (5)
- Stražar 2023 — HLA-II (1)
Install
pip install hitlist
Quick start for ML training
# One-time: register IEDB + CEDAR downloads and build
hitlist data register iedb /path/to/mhc_ligand_full.csv
hitlist data register cedar /path/to/cedar-mhc-ligand-full.csv
hitlist build observations # a few minutes end-to-end;
# writes observations.parquet +
# binding.parquet + peptide_mappings.parquet
# Export training-ready CSVs
hitlist export training --include-evidence ms --class I --species "Homo sapiens" --mono-allelic \
--min-allele-resolution four_digit -o mono_allelic_classI.csv
hitlist export training --include-evidence ms --class II --species "Homo sapiens" \
-o classII_training.csv
# Presto-style flank-aware export: one row per (evidence row, peptide mapping)
hitlist export training --include-evidence both --class I --species "Homo sapiens" \
--explode-mappings -o presto_training.parquet
hitlist export training does not create a new canonical store. It composes the existing observations.parquet, binding.parquet, and peptide_mappings.parquet indexes into one training-facing export surface. The low-level indexes keep their semantic boundaries; the training export gives downstream consumers one obvious API/CLI path when they want model-ready tables.
Python API
hitlist build observations checks the source files and the curation inputs that define stored
annotations, including study overrides, tissue/cell-line metadata, and peptide-attribution CSVs.
Editing these inputs triggers a rebuild without --force. Upgrading from an older cache that
lacks curation fingerprints also triggers one rebuild.
Training-data export
from hitlist.export import generate_ms_observations_table
from hitlist.export import generate_training_table
# Mono-allelic human class I MS observations with ground-truth allele
mono_ms = generate_ms_observations_table(
mhc_class="I",
species="Homo sapiens",
is_mono_allelic=True,
min_allele_resolution="four_digit",
)
# Presto-style mapping-aware export: one row per (evidence row, peptide mapping)
presto = generate_training_table(
include_evidence="both",
mhc_class="I",
species="Homo sapiens",
map_source_proteins=True,
)
# columns now include: evidence_kind, evidence_row_id, protein_id, position,
# n_flank, c_flank, proteome, proteome_source
generate_observations_table() remains available as a backward-compatible alias.
Gene queries can mix symbols and Ensembl IDs: gene=["PRAME", "ENSG00000198681"] selects
peptides matching either gene. An additional peptide= filter narrows that union. Mapping-expanded
training exports retain only mappings for the selected genes, even when a peptide also maps to
an unrelated gene. Separately supplied gene_name and gene_id filters on the low-level loaders
remain conjunctive.
Allele-set filters use the same normalization in raw loaders and exports: mhc_allele_in_set="A*02:01"
and mhc_allele_in_set="HLA-A*02:01" select the same evidence. Explicit empty allele-set queries
raise ValueError consistently.
Species filters accept any variant — "Homo sapiens", "human", "homo_sapiens", "Homo sapiens (human)" all work.
MS, binding, training, and peptide-summary exports support independent source_species and
host_species filters, plus exclude_chimeric=True. The existing species filter selects MHC
species. For example, generate_training_table(species="human", source_species="mouse") selects
mouse-source peptides presented on human MHC. The matching CLI flags are --source-species,
--host-species, and --exclude-chimeric; quote multiword species names. Defaults retain all
species and chimeric systems. Source filtering and system flags both fall back to the raw
species column when source_organism is blank; raw source annotations remain available.
Raw observations loading
from hitlist.observations import (
load_ms_observations, # MS-eluted immunopeptidome
load_binding, # in-vitro binding-assay measurements
load_all_evidence, # union, tagged with an evidence_kind column
is_built, is_binding_built,
observations_cache_is_current, # current vs. legacy; never builds or prints
observations_path, binding_path,
)
# MS-elution (the default training-data path)
df = load_ms_observations() # everything (MS-eluted only)
df = load_ms_observations(mhc_class="I") # class I only
df = load_ms_observations(species="Homo sapiens") # human only
df = load_ms_observations(source="iedb") # filter by source
df = load_ms_observations(columns=["peptide", "mhc_restriction", "src_cancer"])
# Binding assays — same filter API, reads binding.parquet
bd = load_binding(mhc_class="I", mhc_restriction="HLA-A*02:01")
# Union — for affinity-predictor training, or UI flags that want both.
# Rows are tagged with evidence_kind ∈ {"ms", "binding"}.
both = load_all_evidence(gene_name="PRAME", mhc_class="I")
both["evidence_kind"].value_counts()
load_observations() remains available as a backward-compatible alias.
is_built() only says a parquet exists. To learn whether it is current — the same
verdict build_observations(force=False) reaches before skipping — call
observations_cache_is_current(). It returns True, False, or None when no IEDB/CEDAR
source is registered and validity cannot be judged. hitlist.mappings.mappings_cache_is_current()
answers the same for the peptide_mappings.parquet sidecar. Neither prints, writes, or
downloads, so a consumer can announce a rebuild before starting one instead of buffering
the builder's output and inferring afterwards.
Building / curation
from hitlist.builder import build_observations
from hitlist.curation import (
classify_ms_row,
normalize_species,
normalize_allele,
load_pmid_overrides,
)
from hitlist.supplement import scan_supplementary, load_supplementary_manifest
build_observations(with_flanking=True, use_uniprot_search=True, force=False)
normalize_species("human") # → "Homo sapiens"
normalize_allele("H-2Kb") # → "H2-K*b"
scan_supplementary() # DataFrame of curated paper-supplement peptides
Peptide → protein attribution and flanking context
hitlist build observations always produces three parquet files (use --no-mappings to skip peptide_mappings.parquet):
observations.parquet— one row per assay observationbinding.parquet— one row per binding-assay observationpeptide_mappings.parquet— one row per (peptide, protein, position)
They land in the data directory (hitlist data dirs; see
Where hitlist keeps data).
The mappings sidecar preserves multi-mapping so a peptide shared by MAGEA1/A4/A10/A12 keeps every paralog. Ensembl mappings include gene_biotype: the default index covers ordinary protein_coding genes plus the coding IG_V/D/J/C_gene and TR_V/D/J/C_gene biotypes, while excluding pseudogenes. These receptor records are germline segments; Ensembl does not contain a donor's recombined receptor, so peptides spanning a V(D)J junction cannot map. Observations additionally carry semicolon-joined identity columns:
| column | example |
|---|---|
gene_names |
MAGEA4;MAGEA10 |
gene_ids |
ENSG00000147381;ENSG00000124260 |
protein_ids |
P43359;P43363 |
n_source_proteins |
2 |
from hitlist.observations import load_ms_observations
from hitlist.mappings import load_peptide_mappings
# Central columns — fast for everyday filters (uses mappings sidecar for pushdown)
df = load_ms_observations(gene_name="PRAME")
# Long form for paralog / position / flank analysis
mappings = load_peptide_mappings(gene_name="MAGEA4")
# Receptor-derived mappings remain distinguishable from ordinary proteins
receptor_mappings = load_peptide_mappings(
gene_biotype=["IG_V_gene", "IG_D_gene", "IG_J_gene", "IG_C_gene",
"TR_V_gene", "TR_D_gene", "TR_J_gene", "TR_C_gene"]
)
# columns include: peptide, protein_id, gene_name, gene_id, gene_biotype,
# transcript_id, position, n_flank, c_flank, proteome
For ad-hoc queries without building the full table:
from hitlist.proteome import ProteomeIndex
idx = ProteomeIndex.from_ensembl(release=112, species="human") # coding + germline IG/TR
flanking = idx.map_peptides(["SLLMWITQC", "GILGFVFTL"], flank=10)
# Explicit compatibility mode for the historical protein-coding-only index:
ordinary_only = ProteomeIndex.from_ensembl(release=112, biotype="protein_coding")
Proteome registry / UniProt resolution
from hitlist.downloads import (
lookup_proteome, # org string → registry entry (dict)
fetch_species_proteome, # download FASTA and cache to <data dir>/proteomes/
resolve_proteome_via_uniprot, # direct UniProt REST lookup
list_proteomes, # manifest section
)
lookup_proteome("Mycobacterium tuberculosis")
# → {'kind': 'uniprot', 'proteome_id': 'UP000001584', ...} # H37Rv reference
Output schema — generate_ms_observations_table()
| Column | Meaning |
|---|---|
peptide |
Amino acid sequence |
mhc_restriction |
Allele from IEDB (may be "HLA class I" for multi-allelic studies) |
sample_mhc |
Experimental MHC candidates; may be a selected restriction or an unresolved study/class pool, not a complete genotype |
sample_mhc_origin |
sample (this arm's own candidates), class_pool (the study's class-wide union), blank (none reached the row), or not_applicable on a binding row in the training table. The fact sample_match_type cannot give you: that column says whether the study has a pool, not whether this row took it |
mhc_basis |
Source-reviewed selected_restriction or sample_typing. Blank on an exported row means either unreviewed or that the claim was dropped because sample_mhc is not that arm's — sample_mhc_origin tells you which |
mhc_genotype, mhc_genotype_cell, mhc_genotype_source |
Independently sourced cellular typing, its cell/donor identity and citation; never pooled across cells |
mhc_genotype_reported_loci, mhc_genotype_complete_loci |
Loci with molecular typing versus explicitly complete typing; missing loci remain unknown |
mhc_class |
Canonical I, II, or non-classical; molecule-derived when possible |
mhc_class_reported |
Source-reported class, retained verbatim for auditability |
mhc_class_source, mhc_class_corrected |
Whether class came from one molecule, a consistent donor set, or source fallback; whether it corrected the source |
mhc_species |
Canonical MHC species (mhcgnomes plus explicit per-study context for ambiguous names) |
mhc_species_source, mhc_species_context_disagrees |
Species-resolution provenance and explicit context-conflict signal |
restriction_evidence |
How the named peptide-to-MHC restriction was established: experimental, monoallelic, predicted, or unknown |
serotype, serotypes |
Canonical serotype and full membership, preserving the MHC species prefix (for example HLA-A2 or BoLA-A18). An allele can carry a locus serotype and a public epitope such as Bw4. |
serotype_source |
Whether the serotype is primary data or a projection: reported (the study typed serologically; no molecule was measured, which is why the allele fields are empty), derived (computed from a named molecule through mhcgnomes' membership table, so only as current as the installed mhcgnomes — the build records mhcgnomes_version in observations_meta.json), donor_set (a union over a donor's typed alleles, making the serotype a candidate rather than the restriction's identity), or empty (no serotype) |
is_monoallelic |
True if sample has a single transfected allele (721.221, C1R, K562, MAPTAC…) |
has_peptide_level_allele |
True if mhc_restriction is a specific allele (not "HLA class I") |
is_potential_contaminant |
True for MS-eluted peptides that failed NetMHCpan binding prediction |
sample_match_type |
How sample_mhc was populated (see below) |
matched_sample_count |
Number of curated samples for this PMID |
src_cancer, src_healthy_tissue, src_ebv_lcl, ... |
Mutually-exclusive biological source categories |
source |
iedb, cedar, or supplement |
source_organism, reference_title, cell_name, source_tissue, disease |
IEDB sample context |
instrument, instrument_type, acquisition_mode, fragmentation, labeling, ip_antibody |
MS acquisition from ms_samples curation |
gene_names, gene_ids, protein_ids, n_source_proteins |
Multi-mapping peptide → source-protein attribution (always populated; use peptide_mappings.parquet for long-form positions + flanks) |
sample_match_type — join provenance
| Value | Meaning | Training-grade? |
|---|---|---|
allele_match |
IEDB recorded an allele matching the curated experimental candidates | Subject to the row's restriction evidence and sample attribution |
single_sample_fallback |
Study has exactly one eligible sample; its reported candidates supply sample_mhc |
Candidate set for that sample, with source precision retained |
pmid_class_pool |
Class-based attribution; unresolved rows may carry the union of class-matching candidates | An unresolved pool must not be used as one sample's genotype |
unmatched |
No curated sample for this PMID, or all samples have mhc: unknown |
No — sample_mhc empty |
sample_attribution says how the row reached its sample, and one value is
worth reading before trusting an arm. elution_conditions_excluded means the
row's own deposited elution statement named arms that exclude the one its
allele matched — for a pan-class-II elution of a line whose class-II arms are
all transduced, a peptide deposited under the untransduced statement. Those
rows carry no sample_label and take the class pool, because the arm they came
from is named by the deposit and absent from curation (#565, #567). Treat them
as unattributed, not as the arm their allele happens to fit.
Biological source classification
Every observation is classified by mutually-exclusive biological source category:
| Category | Flag | Rule |
|---|---|---|
| Cancer | src_cancer |
Tumor tissue, cancer patient biofluids, or non-EBV cell lines |
| Adjacent to tumor | src_adjacent_to_tumor |
Surgically resected "normal" tissue (per-PMID override) |
| Activated APC | src_activated_apc |
Monocyte-derived DCs/macrophages with pharmacological activation |
| Healthy somatic | src_healthy_tissue |
Direct ex vivo, healthy donor, non-reproductive, non-thymic |
| Healthy thymus | src_healthy_thymus |
Direct ex vivo thymus (expected for CTAs, AIRE-mediated) |
| Healthy reproductive | src_healthy_reproductive |
Direct ex vivo testis, ovary (expected for CTAs) |
| EBV-LCL | src_ebv_lcl |
EBV-transformed B-cell lines |
| Cell line | src_cell_line |
Any cultured cell line |
Cancer-specific = src_cancer AND NOT src_healthy_tissue. Thymus, reproductive tissue, adjacent tissue, EBV-LCLs, and activated APCs do NOT disqualify a peptide from being cancer-specific.
CLI reference
Data management
hitlist data register <name> <path> [-d DESCRIPTION] # register a local file
hitlist data fetch <name> [--force] # download a known dataset (IEDB/CEDAR/viral FASTAs)
hitlist data refresh <name> # re-download
hitlist data info <name> # detailed metadata (JSON)
hitlist data path <name> # print the registered path
hitlist data remove <name> [--delete] # unregister (optionally delete file)
hitlist data list # show registered datasets
hitlist data available # show all known datasets
hitlist data dirs # every directory hitlist uses, and why
Where hitlist keeps data
hitlist data dirs is the complete answer; the short version:
| what | where |
|---|---|
built indexes (observations/binding/bulk_proteomics/line_expression parquets, manifest.json, downloaded proteomes) |
the data directory — see the resolution order below |
| mirrored data assets (paper-derived CSVs) | datacache's cache dir for the hitlist subdir; not moved by HITLIST_DATA_DIR |
proteome index cache (*.pkl) |
always ~/.hitlist/proteome_index_cache — not moved by HITLIST_DATA_DIR (#591) |
The data directory resolves in this order:
hitlist.downloads.set_data_dir()- the
HITLIST_DATA_DIRenv var — the directory itself, exactly as it always has been; whitespace is stripped,~is expanded, and an empty value means unset (it used to resolve to the process's working directory) - an existing, populated
~/.hitlist— the legacy location, still fully supported, so an install that already has a corpus there keeps using it (and says so once per process) - otherwise
datacache.get_data_dir(subdir="hitlist")—~/Library/Caches/hitliston macOS,~/.cache/hitliston Linux, the convention pyensembl and the rest of the openvax ecosystem already use
A ~/.hitlist counts as populated when it holds a regular file at the top
level, or a subdirectory holding a regular file. The test is structural on
purpose — a list of known artifact names would drift, and would already have
missed the user whose only data is gene_cache/hgnc_lookups.json.
Merely existing is not enough: older releases created ~/.hitlist on every
call to the path helper, and <data dir>/proteomes/ and <data dir>/gene_cache/
are still created eagerly, so plenty of installs have one holding nothing but
empty folders.
Nothing moves or rebuilds when you upgrade. To move an existing corpus, set
HITLIST_DATA_DIR to the new location and copy the old directory's entire
contents there. Moving only the parquets leaves manifest.json behind, which
keeps rule 3 pointed at a ~/.hitlist that no longer has any indexes in it.
Build the observations table
hitlist build observations [--force] # ~90s full scan with tqdm progress
hitlist build observations # always builds peptide_mappings.parquet
hitlist build observations --use-uniprot # broader proteome coverage via UniProt REST
hitlist build observations --no-mappings # skip mapping step (faster, no gene attribution)
hitlist build observations --no-fetch-proteomes # don't auto-download missing proteomes
hitlist build observations --proteome-release 112 # Ensembl release for human/mouse/rat
Proteome management
hitlist data fetch-proteomes [--min-observations N] [--use-uniprot] [--force]
hitlist data list-proteomes
Export
hitlist export ms [filters...] -o train.csv # MS immunopeptidome + sample metadata
hitlist export ms -o train.parquet # parquet output supported
hitlist export peptide-summary --gene PRAME --serotype A24 # per-peptide support for one allele/serotype
hitlist export binding [filters...] -o binding.csv # binding-assay index (separate from MS)
hitlist export training [filters...] -o training.csv # unified training export from canonical indexes
hitlist export peptide-counts --by class # species x class peptide counts
hitlist export peptide-counts --by study # peptide counts per study/PMID
hitlist samples [--class I|II] # per-sample conditions (promoted from `export samples`)
hitlist qc normalization # validate YAML alleles with mhcgnomes
hitlist qc mhc-tokens # unparseable MHC tokens across all modalities
hitlist qc resolution # allele-resolution histogram for IEDB/CEDAR
hitlist export ms is the canonical name for the MS observations export; hitlist export observations remains as a backward-compatible alias.
Canonical indexes and the training export
Each hitlist build observations writes three parquet files to the data
directory (hitlist data dirs):
observations.parquet— MS-eluted immunopeptidome (IEDB + CEDAR + curated supplementary).binding.parquet— binding-assay rows (peptide microarray, refolding, MEDi, and quantitative-tier measurements likePositive-High/Intermediate/Low).peptide_mappings.parquet— long-form peptide → protein/position/flank mappings.
The canonical indexes are never silently mixed. Supplementary data is MS-only.
Use hitlist export observations and hitlist export binding when you want
the raw evidence families separately. Use hitlist export training or
generate_training_table(...) when you want a composed model-facing export
with evidence_kind tagging and optional mapping explosion for flank-aware
training pipelines.
Filters on hitlist export observations
| Flag | Values |
|---|---|
--class |
I, II, non classical |
--species |
Any species variant (normalized via mhcgnomes) |
--mono-allelic / --multi-allelic |
Filter on is_monoallelic |
--instrument-type |
Orbitrap, timsTOF, TOF, QqQ, ... |
--acquisition-mode |
DDA, DIA, PRM |
--min-allele-resolution |
four_digit, two_digit, serological, class_only |
--mhc-allele |
Exact match on mhc_restriction after allele normalization. Repeatable / comma-separated. |
--restriction-evidence |
Restriction evidence: experimental, monoallelic, predicted, or unknown. Independent of allele-set provenance. Repeatable. |
--serotype-source |
Whether the serotype is primary data or a projection: reported (serologically typed study), derived (computed from a named molecule), donor_set (union over a donor's typed alleles). Pair with --serotype to keep a serotype query on measured typings. Repeatable. |
--gene |
Symbol, Ensembl ID, or old alias (HGNC synonym lookup). Repeatable / comma-separated. Requires the mappings sidecar (default-on at build). |
--gene-name |
Exact match on gene_name column (no HGNC lookup) |
--gene-id |
Exact match on gene_id column (ENSG) |
--peptide |
Exact match on peptide sequence. Repeatable / comma-separated. |
--serotype |
MHC serotype: use a species prefix for non-human names (BoLA-A18, Patr-DR1); HLA names may omit it (A24, B57, DR15, Bw4, Bw6). Matches any serotype the allele belongs to, so --serotype Bw4 returns A*24:02, B*27:05, B*57:01, etc. HLA split serotypes are rolled into their broad parent, so --serotype A24 also matches A*24:03 (serotype A2403). Repeatable / comma-separated. |
--exclude-class-label-suspect |
Drop rows where the curated class disagrees with peptide length severely enough to be flagged suspect or implausible (mhc_class_label_severity). |
--exclude-class-label-implausible |
Strict-cleaning variant — drops only implausible rows (class-I ≥18aa or ≤7aa, class-II ≤4 or ≥45aa). Keeps borderline + suspect tiers, useful when bulged class-I 15-17aa peptides should be retained. |
--apm-only |
Filter to peptide rows from samples where any APM gene was perturbed (apm_perturbed=True). Reflects the sample's own condition; the parent study's perturbation panel is carried separately in study_apm_perturbed / study_apm_genes. |
--output / -o |
.csv or .parquet |
For filtering a frame already in memory, hitlist.normalize_serotype_query() uses
the same spelling rules as the loaders and exports: "a2" becomes "HLA-A2",
and "bola-a18" becomes "BoLA-A18".
All filters are pushed down to the parquet reader (pyarrow), so --gene PRAME reads
only the matching row groups — typically milliseconds rather than a full table scan.
Examples:
hitlist export ms --gene PRAME --class I -o prame_classI.csvhitlist export ms --gene "MART-1"(HGNC resolves toMLANA)hitlist export ms --mhc-allele HLA-A*02:01 --mono-allelichitlist export ms --serotype A24(locus-specific)hitlist export ms --serotype Bw4(public epitope — A23/24/25/32, B13/27/44/51/52/53/57/58)
hitlist export peptide-summary
Collapses the MS observations to one row per peptide, scoring how strongly each peptide is supported on a single target allele or serotype. Answers questions like "which PRAME peptides might be presented on A24 in cancers?"
hitlist export peptide-summary --gene PRAME --serotype A24 -o prame_a24.csv
hitlist export peptide-summary --gene PRAME --mhc-allele HLA-A*24:02
Requires exactly one of --mhc-allele / --serotype (a single value) plus a
gene/peptide scope filter. Each peptide's support is split into ranked tiers,
strongest first:
| Bucket column | Meaning |
|---|---|
n_mono_exact_rows |
Mono-allelic elution on the exact target allele (strongest) |
n_multi_exact_rows |
Multi-allelic elution including the exact target allele |
n_mono_serotype_rows / n_multi_serotype_rows |
Elution on a same-serotype allele |
n_class_only_sample_allele_rows |
Class-only row whose experimental candidates include the target allele |
n_class_only_sample_serotype_rows |
Class-only row whose experimental candidates include a matching serotype |
n_unknown_allele_rows |
Class-only row with no usable reported candidates (weakest) |
best_support names the strongest tier with any rows. A cancer / healthy /
adjacent / other source breakdown (n_cancer_rows, ...) comes along for free,
so you can prioritize peptides seen in tumors over healthy tissue.
Filters on hitlist export binding
Same shape as the observations filters minus the MS-specific ones
(--mono-allelic, --instrument-type, --acquisition-mode). --source
accepts only iedb or cedar — supplementary data is MS-only and never
appears in the binding index.
hitlist export binding --gene PRAME --class I -o prame_binding.csv
hitlist export binding --mhc-allele HLA-A*02:01 --serotype Bw4
Filters on hitlist export training
hitlist export training exposes the shared pMHC filters plus two export-shape controls:
--include-evidence ms|binding|bothchooses which canonical evidence families to compose.--explode-mappingsexpands the output to one row per(evidence row, peptide mapping)withprotein_id,gene_biotype,position,n_flank,c_flank,proteome, andproteome_source.
MS-specific filters (--mono-allelic, --instrument-type, --acquisition-mode) apply only to the MS slice. Binding rows never gain fake sample context; they remain tagged as evidence_kind="binding" with sample_match_type="not_applicable".
hitlist export training --include-evidence both --gene PRAME --class I -o prame_training.csv
hitlist export training --include-evidence ms --mono-allelic --class I -o mono_ms.csv
hitlist export training --include-evidence both --explode-mappings -o presto_training.parquet
Sample-level expression anchors (issue #140)
For line-like ms_samples (C1R, 721.221, JY and sibling EBV-LCLs, HAP1,
HeLa, HEK293, THP-1, SaOS-2, A375, K562, GM12878, and common engineered
derivatives) hitlist.line_expression resolves every sample to a stable
expression backend + key via a 6-tier fallback hierarchy:
- exact-line RNA / transcript quant (registry hit with shipped data)
- parent-line / engineered-derivative RNA (e.g. HeLa.ABC-KO → HeLa)
- line-family class anchor (EBV-LCL → GM12878; mono-allelic host → K562)
- cancer-type surrogate (caller-supplied
pirlygenesbackend) - broad tissue / lineage surrogate (HPA)
no_expression_anchor
Four provenance columns — expression_backend, expression_key,
expression_match_tier, expression_parent_key — are written on every
row so downstream tooling can distinguish "exact JY RNA" from "generic
EBV-LCL stand-in" from "melanoma cohort surrogate" instead of treating
them as equally trustworthy.
Tier 1 requires that the profiled material is the registered line, so an
engineered sample never resolves there. Engineering comes from curation, not
from parsing the label: any intervention in condition_knockout_genes,
condition_knockdown_genes, condition_overexpression_genes,
condition_genetic_variants, condition_transfection or
condition_transduction, or an introduced-MHC condition_mhc_context
(monoallelic, mhc_transfectant, mhc_coexpression). Such a sample
resolves at tier 2 against the same data even when its label matches the
reference line's alias — HeLa-CIITA (Mock), SaOS-2 + TP53 R175H,
K562 transfectant DPB1*01:01/DPA1*02:01 — with expression_parent_key
naming the material that supplied the RNA. resolve_sample_expression_anchor
derives this itself when given the pmid and sample_label of a curated arm;
pass engineered= to override. A row that names no arm — an observation
the export could not attribute to one of HAP1's wild-type and knockout arms,
say — is not known to be unmodified when the study has an engineered arm of
the same reference line as the row's anchor (lines are matched through the
registry, so aliases count), and also resolves at tier 2 with a reason saying
the arm is unresolved. An unattributed JY row keeps tier 1 in a study whose
only engineered arm is a Raji transfectant.
CLI:
hitlist export samples --with-expression-anchors -o sample_anchors.csv
hitlist export training \
--include-evidence ms --class I --mono-allelic \
--with-peptide-origin \
-o training_with_origin.parquet
hitlist export line-expression --line-key GM12878 --gene-name TP53
--with-peptide-origin attaches, per row, the argmax-TPM candidate gene
(peptide_origin_gene, peptide_origin_tpm) for the peptide in the
sample's resolved line. When transcript-level TPM is available the
score is the sum across only those transcripts whose translation
actually contains the peptide — isoforms that splice out the peptide
contribute zero, and a transcript is counted once regardless of how many
times the peptide appears in its protein.
Fetch and build the optional DepMap 24Q4 gene/transcript expression bundle (about 4.7 GB). It provides RNA anchors for HeLa, A375, SaOS-2, THP-1 and K562, and HAP1 transcript rows. This release has no HEK293 model entry. HAP1's DepMap gene row ships with the package, so HAP1 and its engineered KO panel resolve without the bundle (tier 1 and tier 2 respectively).
hitlist data fetch depmap
The bundle includes the model, profile and default-profile mappings needed
for transcript rows. Existing registered copies are reused. Individual
files can still be fetched or registered under depmap_rna,
depmap_rna_transcript, depmap_models, depmap_profiles and
depmap_default_profiles. RNA anchors report only sources with available
rows; missing optional data falls back to the next applicable tier.
line_expression.parquet is stamped with a content fingerprint of every
packaged build input — the anchor registry, sources.yaml, the packaged CSVs
and the ENSG → HGNC symbol map — plus an index build version. Reads compare
the stamp with the installed release and never rewrite the file:
- current: the index is used as built.
- no stamp (every index written before this release) or same build
version with different packaged inputs: reads use this release's packaged
sources plus the index's downloaded (e.g. DepMap) rows, re-enriched, for
the (line, source) pairs the current registry still lists and whose
provenance meets the current builder's invariants. An upgrade or an anchors
edit therefore does not cost a DepMap user HeLa or K562, while rows the
registry no longer lists (HAP1 under
DepMap_24Q4_gene) are left out, as are rows from before DepMap default-profile selection (#357): an index with noprofile_idcolumn, or a transcript row without itsPR-…profile. - another build version, or an unreadable stamp: packaged sources only.
Without a usable index, the packaged sources are served in the same row shape
a build writes (source metadata, parent_line_key, and gene symbols from the
packaged map), so gene_name filters reach them. The first read of a
non-current index warns once, naming what was left out and the cheapest
rebuild: hitlist data fetch depmap when the index held downloaded rows and
every DepMap file is still registered (a purely local rebuild); otherwise a
local hitlist.builder.build_line_expression(), with the exact DepMap files
a fetch would still download. line_expression_index_is_current() checks
without warning; is_line_expression_built() only says a file exists.
A note on mono-allelic curation
Mono-allelic is a PMID-level flag, not a per-sample property. Curation
lives in pmid_overrides.yaml (see mono_allelic_host and ms_samples).
Rows can legitimately have is_monoallelic=True with an empty
mhc_restriction — for example, supplementary contaminant peptides under
a mono-allelic PMID override carry the flag but not a per-row allele. If
your downstream pipeline needs a strict "mono-allelic AND has allele"
subset, post-filter on
is_monoallelic & mhc_restriction.str.startswith("HLA-").
Reports
hitlist report [--class I|II] [--output report.txt]
QC diagnostics
hitlist qc # run all checks, print summary
hitlist qc resolution # allele-resolution histogram
hitlist qc normalization # YAML alleles whose normalize_allele output drifts
hitlist qc mhc-tokens # invalid/parser-gap/unknown MHC tokens across modalities
hitlist qc cross-reference # alleles in YAML but not data, and reverse
hitlist qc discrepancies [--by sample] # per-PMID curation drift signals
hitlist qc plan [--top 10] # ranked next-PMID-to-curate roadmap
hitlist qc proteome-coverage # per-source-organism proteome registry coverage
hitlist qc proteome-coverage --missing-only --min-rows 100
Development
./develop.sh # install in dev mode
./format.sh # ruff format
./lint.sh # ruff check + format check
./test.sh # pytest with coverage (~3 min)
./deploy.sh # lint + test + build + upload to PyPI
When local memory cannot support the corpus tests, the Release build GitHub workflow runs the same gates on a hosted runner. It requires the full CI corpus, installs the current development heads of the six sibling libraries, and records their resolved revisions with the source commit and artifact checksums. The workflow builds artifacts; PyPI authentication stays on the maintainer's machine.
From a clean, up-to-date main checkout, dispatch the workflow and wait for its
successful manual run. A pull-request validation run cannot authorize publication.
gh workflow run release-build.yml --ref main
gh run list --workflow release-build.yml
Set release_run_id to that successful manual run's ID, then download into an
empty directory and verify before uploading. The verifier rejects a different
source commit or version, a non-main checkout, failed or PR workflow runs, and
modified distributions. Publish the downloaded wheel and sdist without rebuilding:
set -e
release_dir=$(mktemp -d)
gh run download "$release_run_id" \
--name "hitlist-release-$(git rev-parse HEAD)" --dir "$release_dir"
python scripts/release_artifacts.py verify "$release_dir"
python scripts/check_distribution_license.py "$release_dir"
twine check "$release_dir"/*.whl "$release_dir"/*.tar.gz
twine upload "$release_dir"/*.whl "$release_dir"/*.tar.gz
./deploy.sh --build-only retains the lint, full test, build and license gates,
and stops before upload. It does not mark a release as published.
See docs/pmid-curation.md for the curation YAML format and per-study overrides.
Metadata
Release files for hitlist 1.63.11
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| hitlist-1.63.11.tar.gz | 2.9 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| hitlist-1.63.11-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 5.4 MB
Release files / hitlist-1.63.11.tar.gz
| Download URL | hitlist-1.63.11.tar.gz |
|---|---|
| Size | 2.9 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
eb96138eb654c5e0b597e9330beb3fbfe385bf82fafb8619c1907ea7f72fc5b1
|
|
BLAKE2b-256 checksum How to use checksums |
cca1566bfa23e499b1e0411679f919a835cce807576dae27a4940a5b6104b814
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.6
|
Release files / hitlist-1.63.11-py3-none-any.whl
| Download URL | hitlist-1.63.11-py3-none-any.whl |
|---|---|
| Size | 2.5 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6703dcd6fa00ae670c3c1e52018c5127eda07b2a5f4e2b943ffbc090f3d95047
|
|
BLAKE2b-256 checksum How to use checksums |
2e3bc7bce769883868152f64d20f0a48608f935bb6c7503cdb710f3ee512daa9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.6
|