Skip to main content

oncoref

Tests PyPI

Curated cancer reference data — cancer-type ontology, tumor mutational burden (TMB), incidence/mortality, checkpoint-inhibitor (ICI) response, per-cohort RNA-seq expression, Human Protein Atlas (HPA) normal-tissue expression, and HPA-derived cancer-testis antigen references — behind one small Python API, a data fetch/cache CLI, and a set of reference plots.

Role in the stack

The openvax/PIRL tools are split by ownership boundary, not by file format. oncoref is the base layer: downstream packages can depend on it, but it never imports its consumers.

  • oncoref is the empirical base: what is true, measured, or canonically named about cancer types, cohorts, genes, and reference datasets. It owns gene identity and canonicalization, the cancer-type ontology/registry, expression reference data and normalization, epidemiology, TMB, ICI/anti-PD-1 response, and source-anchored cancer-testis antigen (CTA) facts. If a row has an n, a confidence interval, a source cohort, or a PMID/DOI anchoring a measurement, it is usually an oncoref fact.
  • pirlygenes owns curated gene sets and panels: which genes are useful for a purpose. That includes lineage/family/compartment/supertype panels, discriminators, surfaceome, tumor microenvironment (TME) and stem-cell markers, response-signature panels, target-to-therapy registries, and other opinionated selections keyed to oncoref cancer codes and gene IDs. An empty set can be a valid pirlygenes answer.
  • trufflepig owns per-sample interpretation: quality-control (QC) narration, library-prep/source warnings, deconvolution, scoring, and rule tables that fire against one tumor sample.

Adoption is staged: consumers can delegate parity-clean primitives while retaining their own curated artifacts and compatibility APIs. Shared facts and identifiers should be fixed here; purpose-specific panels and per-sample rules stay downstream.

Core model

  • Canonical identities first. Cancer facts key on the cancer-type registry; expression cohorts and evidence scopes are explicit rather than inferred from names. Gene APIs resolve to a canonical Ensembl space.
  • Small facts in the wheel, large matrices on demand. Ontology, gene, burden, response, and provenance tables install with the package. The large expression bundle and HPA datasets use versioned download caches.
  • Provenance is part of the result. Reference rows expose source scope, sample counts, versions, and review status. Missing or ineligible data stays distinguishable from a valid empty result.
  • Consumer policy stays downstream. oncoref exposes HPA-derived CTA calls and empirical expression facts; it does not turn them into therapy panels or one-sample decisions.

Install

pip install oncoref

Optional integrations are explicit extras:

pip install 'oncoref[genome]'  # pyensembl-backed transcript and gene lookup
pip install 'oncoref[plots]'   # reference plotting

Quick start

The flat oncoref namespace remains available for compatibility and quick interactive use. For new code, prefer the semantic submodules in the API guide; they make it clearer whether you are working with the cancer ontology, cohorts, ICI response, CTA coverage, generic antigen-panel coverage, or CTA-specific peptides.

import oncoref as od

od.resolve_cancer_type("prostate")        # -> "PRAD"
od.cancer_type_info("SARC_RMS_ARMS")      # full registry record + burden + tmb
od.cancer_tmb("LUAD_EGFR")                # 6.9  (inherited from LUAD)
od.cancer_burden("pancreas", metric="us_mortality_pct")
od.burden_category("SARC_OS")             # -> "bone_and_joint" (incidence/mortality bucket)
od.cancer_ici_response("SKCM")            # 42% objective response rate
od.cancer_ici_response("SKCM", regimen="PD-1+CTLA-4")   # 57.6  (pin a regimen)

# Cancer-testis antigens (HPA-derived tissue-restriction):
od.cta_gene_names()                       # expressed CTA symbols (MAGEA4, CT83, …)
od.cta_evidence()                         # full HPA restriction table
od.cta_clinical_target_evidence()         # explicit clinical/canonical tier + leak flags

# Per-cohort expression percentiles (downloads the data bundle on first use):
od.cohort_gene_percentiles("PRAD")        # per-gene p0…p100 vector (within-cohort)
od.within_sample_top_fraction("PRAD")     # per-gene frac of samples top-5% (within-sample)

Domain map

Question Preferred modules
What does this cancer code mean? oncoref.cancer_ontology, oncoref.cohorts
What expression reference is available? oncoref.expression, oncoref.source_matrices
How is expression normalized or filtered? oncoref.normalization, oncoref.gene_families
What is the canonical gene identity? oncoref.gene_ids, oncoref.genome, oncoref.proteoforms
What is the TMB, burden, or ICI response evidence? oncoref.tmb, oncoref.incidence, oncoref.ici_response
Which genes meet the HPA-derived CTA definition? oncoref.cta, oncoref.cta_coverage, oncoref.cta_peptides
Where is a dataset cached or downloaded? oncoref.catalog, oncoref.data_bundle, oncoref.hpa

The API guide explains each domain, its provenance contract, and the distinction between reader, builder, and compatibility APIs.

Command line

oncoref cancer-type prostate     # registry info as JSON
oncoref tmb LUAD_EGFR            # 6.9
oncoref ici SKCM                # 42  (--regimen to pin, --all-regimens to compare)
oncoref burden pancreas --metric us_mortality_pct
oncoref cta --count             # number of expressed CTAs
oncoref plot apd1-vs-tmb --out apd1_vs_tmb.png
oncoref plot patient-coverage --gene-set cta --out coverage_out
oncoref plot cta-curation --out cta_curation_out

Data cache

oncoref data list               # every wheel/bundle/HPA/source dataset
oncoref data status bundle      # expression-bundle cache state (no download)
oncoref data metadata           # package/data/cache/release contract JSON
oncoref data fetch bundle       # download the large expression bundle
oncoref data fetch hpa          # HPA RNA / immunohistochemistry / single-cell data
oncoref data dir bundle         # where the data bundle is cached
oncoref data prune --yes        # delete stale bundle version caches
oncoref version

Development

./develop.sh   # editable install with dev extras
./format.sh    # ruff format
./lint.sh      # ruff check + format --check
./test.sh      # lint + pytest with coverage

License

Apache 2.0.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

oncoref-1.8.165.tar.gz (3.0 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

oncoref-1.8.165-py3-none-any.whl (2.9 MB view details)

Uploaded Python 3

File details

Details for the file oncoref-1.8.165.tar.gz.

File metadata

  • Download URL: oncoref-1.8.165.tar.gz
  • Upload date:
  • Size: 3.0 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.6

File hashes

Hashes for oncoref-1.8.165.tar.gz
Algorithm Hash digest
SHA256 f055e5b4f9afcb85a10c43b12324edfa2a498acdac8e5f66e27aaebaa7597677
MD5 27f68e9ea65ccba285a8f5a34f78310d
BLAKE2b-256 db24582ed13fc7644449c102dc56185bf7a6a886e95985251af9586ab496dd85

See more details on using hashes here.

File details

Details for the file oncoref-1.8.165-py3-none-any.whl.

File metadata

  • Download URL: oncoref-1.8.165-py3-none-any.whl
  • Upload date:
  • Size: 2.9 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.6

File hashes

Hashes for oncoref-1.8.165-py3-none-any.whl
Algorithm Hash digest
SHA256 152128678bc96a5bb18f67e279bbf53d5f526951450503c39b8c6f7bf15c3805
MD5 8cd6820fd70f759c8031d1dfbd81e6fc
BLAKE2b-256 15302fd9e8bce6fac71f3bc03cf342f798db9547c21741a752d0c5f39d74d423

See more details on using hashes here.

Release history Release notifications | RSS feed

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page