polygenic - the polygenic scores toolkit
Basic info
Downloads
Index
Summary
Polygenic is a toolkit for a wide range of polygenic scores analysis tasks. The most important use cases include computing scores for samples in vcf files, building scores for GWAS results or fetching scores from repositories.
Diplotyping Algorithm
We begin by reading individual genetic variants (genotypes) from the patient's VCF file, where each genotype carries two alleles — one per chromosome. The system supports phasing using a custom reference panel to resolve which alleles sit on the same chromosome; when phased data is available, the algorithm preserves this linkage information, while for unphased data both chromosomes are treated symmetrically. We define haplotypes as specific combinations of co-inherited variants that form recognized gene versions, such as pharmacogenomic star alleles, where each haplotype definition distinguishes core defining variants (weighted at 1.0) from supportive sub-lineage variants (weighted at 0.05). The algorithm scores every candidate haplotype against the patient's alleles, then selects the best-matching haplotype for the first chromosome — keeping all candidates within a 2% margin of the top score. The matched alleles are then "claimed" by that haplotype, and the remaining unmatched alleles (leftovers) are passed into a second round, where up to 100 candidate haplotypes are re-scored against only those residual alleles to identify the second haplotype. Each first/second haplotype pair is ranked by combined match percentage and filtered by total missing data, producing the final diplotype call — e.g., CYP2D6 *1/*4. Every variant in the result carries a source label — direct genotyping, LD proxy, imputation flag from VCF, allele-frequency-based imputation, reference, or missing — enabling granular quality control and full traceability of the diplotype call.
Missing variant handling (reference fill + --no-ref-fallback / --no-indeterminate)
By default, variants called ./. in the input VCF are filled with homozygous reference (source: "reference") so a best-effort call and scores always compute. Pass --no-ref-fallback for strict behavior: ./. is then treated as missing (source: "missing") and subtracted from both the numerator and the denominator of the haplotype match score (the PharmCAT Named Allele Matcher approach of dropping missing positions), which can leave a sparse panel with no confident call. (There is no --ref-fallback flag — filling is the default as of 2.5.19.)
Filling alone could assert a wild-type call you can't justify, so two safeguards ride alongside it:
-
Indeterminate verdict (on by default; disable with
--no-indeterminate). The reference fill is recorded assource: "reference"and treated as uncovered by the verdict — i.e. it is computed on the pre-fill state. The headline call is"Indeterminate"only when ALL defining (function-level /core) variants are uncovered (missing, low-quality, or reference-filled). Otherwise the best-effort call is emitted, and which star alleles cannot be ruled in/out is reported per-allele inundeterminable(an allele is listed when any of its defining variants is uncovered and it is not contradicted by a covered slot). The best-effort call for a triggered (all-uncovered) case is kept ascall_filled. -
reference_defining_variants(per-model annotation). Positions where the GRCh38 reference base itself defines an allele are never reference-filled, and when unassayed they surface inundeterminable. The canonical case is CYP2C19: rs3758581 (chr10:94842866) has referenceA, butGis the global major allele and PharmVar v5 defines every "real" star allele (*1, *2, *17 …) as requiringGthere, reserving CYP2C19*38 (= the empty *1.001) for the minority carrying referenceA(see PharmVar GeneFocus: CYP2C19). Without the annotation, filling an unassayed rs3758581 withAwould collapse every sample toward *38; the annotation keeps the real alleles callable from their own variants and flags the rs3758581-dependent alleles instead. (To author such a model, list the position underreference_defining_variants— see the guidance inCLAUDE.md.)
Match-confidence gate (--top-n)
After scoring all candidate haplotype pairs, the algorithm returns a haplotype_id when either of these holds:
- the best pair's per-chromosome
max_percent_matchis≥ 50%(high-confidence call), or - the total scored-candidate pool size is
≤ top_n(default 15) — a small, non-ambiguous candidate space is itself a confidence signal.
Otherwise haplotype_id is None (the caller refuses to guess). This is more conservative than PharmCAT, which always returns the top-scoring diplotype with deterministic tie-breaking, and roughly comparable to Aldy 4, which reports "no more than three diplotypes" when it cannot fully disambiguate. Tune --top-n higher (laxer) or lower (stricter) depending on how much ambiguity is acceptable downstream. Set --top-n 0 to rely solely on the 50% threshold.
CYP2C19*38 worked example
On a clinical panel that does not assay rs3758581, cyp2c19-pharmvar-5.1.8.yml lists it under reference_defining_variants, so it is never reference-filled. A sample's actual informative variants drive the call, and the rs3758581-dependent alleles (*38 vs *1) are reported in undeterminable rather than miscalled:
| Sample | Non-ref CYP2C19 variants | Call | Note |
|---|---|---|---|
| Sample A | none | *1-like / not a confident *38 |
*1 vs *38 cannot be distinguished without rs3758581 → flagged in undeterminable |
| Sample B | rs12769205, rs4244285 (het) | contains *2 ✓ |
driven by rs4244285 |
| Sample C | rs12248560 (hom) | *17/*17 ✓ |
driven by rs12248560 |
| Sample D | rs12769205, rs4244285 (het) | contains *2 ✓ |
driven by rs4244285 |
The point of the annotation: were rs3758581 blindly filled with reference A/A, it would fail the G required by *1/*2/*17 … and collapse every sample to the empty *1.001 (= *38). Because it is not filled, the real variants still call *2/*17, and the inability to tell *1 from *38 surfaces as an undeterminable entry — never a spurious confident *38.
Installation
Requirements
pip install polygenic-pgx is enough — there is no system dependency and no C compiler
needed. Index access goes through pysam, which bundles htslib and ships manylinux wheels.
Older releases (before 2.5.34) required the tabix binary on PATH, because rsID-keyed
models resolved genotypes by shelling out to it; that is no longer the case, and neither is the
pytabix build step those instructions existed for.
The input VCF must still be bgzipped and tabix-indexed (a .tbi next to the .vcf.gz).
If you need to create one, any htslib build will do:
tabix -p vcf sample.vcf.gz # or: bcftools index -t sample.vcf.gz
With pip
Install for user account
python3 -m pip install --upgrade polygenic-pgx
Install globally
sudo -H python3 -m pip install polygenic-pgx
With conda
Run conda image
docker run -it conda/miniconda3 /bin/bash
Create python3.8 environment and install polygenic
yes | conda create --name py38 python=3.8
eval "$(conda shell.bash hook)"
conda activate py38
### should be 3.8
python --version
pip install polygenic-pgx
With docker
Large image with all data included
docker run intelliseq:polygenictk:2.1.0 *command*
Thin image with just polygenic package installed
docker run intelliseq:polygenic:2.1.0 *command*
Quick start guide
mkdir polygenic && cd polygenic # create working directory
wget https://downloads.intelliseq.com/public/polygenic/gbe-INI78-bone-density.yml # download model
wget https://downloads.intelliseq.com/public/polygenic/illu_merged-imputed.vcf.gz # download genotypes
wget https://downloads.intelliseq.com/public/polygenic/illu_merged-imputed.vcf.gz.tbi # download position index
wget https://downloads.intelliseq.com/public/polygenic/illu_merged-imputed.vcf.gz.idx.db # download rsid index
docker run -v $(pwd):/data intelliseq/polygenic:latest --vcf /data/illu_merged-imputed.vcf.gz --model /data/gbe-INI78-bone-density.yml --output-directory /data # compute model
Manual
Tools
pgs-compute
usage: pgstk [-h] -i VCF [-m MODEL [MODEL ...]] [-p PARAMETERS] [-s SAMPLE_NAME] [-o OUTPUT_DIRECTORY] [-n OUTPUT_NAME_APPENDIX] [-l LOG_FILE] [--af AF] [--af-field AF_FIELD]
[-v] [--print]
pgs-compute computes polygenic scores for genotyped sample in vcf format
optional arguments:
-h, --help show this help message and exit
-i, --vcf VCF vcf.gz file with genotypes
-m, --model MODEL [MODEL ...]
path to .yml model (can be specified multiple times with space as separator)
-p, --parameters PARAMETERS
parameters json (to be used in formula models)
-s, --sample-name SAMPLE_NAME
sample name in vcf.gz to calculate
-o, --output-directory OUTPUT_DIRECTORY
output directory
-n, --output-name-appendix OUTPUT_NAME_APPENDIX
appendix for output file names
-l, --log-file LOG_FILE
path to log file
--af AF vcf file containing allele freq data
--af-field AF_FIELD name of the INFO field to be used as allele frequency
-v, --version show program's version number and exit
--print Print output to stdout
Arguments
Required
--vcfvcf.gz file with genotypes (tabix index should be available)--modelpath to model file
Optional
--log_filelog file--out_dirdirectory for result jsons--populationpopulation code--models_pathpath to a directory containing models--afan indexed vcf.gz file containing allele freq data--versionprints version of package
Building models in yml
Index: Model structure Model types Parameters
Model structure
Core structure
Models have two properties which is model and description. model is a specification of computation to be performed and description is additional information to be included in the result.
model:
description:
Object keys
Each object that is not collection has a set of predefined keys (required or optional) that can be used for computation. For example: diplotype_model object has a required diplotypes key.
diplotype_model:
diplotypes:
The computation is first delegated to key specified objects and later aggregated by the top level object itself.
Collections
There is special category of objects that don't have predefined keys but are collections. Each key within collection becomes element of collection. Collections are easy to recognize, because they are specified in plural form like diplotypes or variants. Each element of collection will be defined as singular object of collection type. For example key in variants collection will becomes objects of variant type.
variants:
rs7041: {diplotype: C/C}
rs4588: {diplotype: T/T}
Variants
Variants can be identified by rsid. Variant value will be computed basing on information provided: diplotype or effect_allele.
Accepted sets of fields are:
- diplotypes
diplotypesymbol
- score
effect_alleleeffect_sizesymbol
Model types
There are currently implemented four types of models:
score_modeldiplotype_modelhaplotype_modelformula_modelThe type of model can be specified at the top of yml structure or within themodelfield.
Specification of model type at the top of yml structure
diplotype_model:
description:
Specification of model type within the model field
model:
diplotype_model:
description:
Parameters
External parameters can be used in formula_model through @parameters keyword.
Example parameters file in .json format:
{"sex": "F"}
Path to file can be provided as argument to polygenic tool:
--parameters /path/to/parameters.json
Example of use of parameters in the formula_model:
formula_model:
formula:
value: "@female.score_model.value if @parameters.sex == 'F' else @male.score_model.value"
male:
score_model:
variants:
...
female:
score_model:
variants:
Example models
Example diplotype model
This example diplotype model is based on Randolph 2014.
diplotype_model:
diplotypes:
1/1:
variants:
rs7041: {diplotype: C/C}
rs4588: {diplotype: T/T}
1/1s:
variants:
rs7041: {diplotype: C/C}
rs4588: {diplotype: T/G}
1/1f:
variants:
rs7041: {diplotype: C/A}
rs4588: {diplotype: T/G}
1/2:
variants:
rs7041: {diplotype: C/A}
rs4588: {diplotype: T/T}
1s/1s:
variants:
rs7041: {diplotype: C/C}
rs4588: {diplotype: G/G}
1s/1f:
variants:
rs7041: {diplotype: C/A}
rs4588: {diplotype: G/G}
1s/2:
variants:
rs7041: {diplotype: C/A}
rs4588: {diplotype: G/T}
1f/1f:
variants:
rs7041: {diplotype: A/A}
rs4588: {diplotype: G/G}
1f/2:
variants:
rs7041: {diplotype: A/A}
rs4588: {diplotype: G/T}
2/2:
variants:
rs7041: {diplotype: A/A}
rs4588: {diplotype: T/T}
description:
pmid: 24447085
genes: [GC]
result_diplotype_choice:
1/1: Moderate
1/1s: High
1/1f: High
1/2: Low
1s/1s: Very high
1s/1f: Very high
1s/2: Moderate
1f/1f: Very high
1f/2: Moderate
2/2: Very low
Example haplotype model
Haplotype model can be used for HLA and PGx.
To define haplotype models a list of alleles is required (called variants in this case, to be consistent with othe rypes of models). Each allele has associated list of defining mutations (alternative SNV alles) defined by Gnomad ID along with ref, alt and effect_allele properties. One star allele should be empty (containing only reference SNV alleles). The algorithm will utilised any phasing information in the vcf.
haplotype_model:
variants:
CYP2D6*1.001:
CYP2D6*1.002:
22-42126963-C-T: {ref: "C", alt: "T", effect_allele: "T"}
CYP2D6*1.003:
22-42128813-G-A: {ref: "G", alt: "A", effect_allele: "A"}
CYP2D6*1.004:
22-42128216-G-T: {ref: "G", alt: "T", effect_allele: "T"}
CYP2D6*1.005:
22-42128922-A-G: {ref: "A", alt: "G", effect_allele: "G"}
CYP2D6*1.006:
22-42129726-A-C: {ref: "A", alt: "C", effect_allele: "C"}
22-42129950-A-C: {ref: "A", alt: "C", effect_allele: "C"}
22-42130482-C-A: {ref: "C", alt: "A", effect_allele: "A"}
For copy-number star alleles (CYP2D6 *5/*1xN, CYP2C19 *36/*37) the model gains a
copy_number: block and structural: haplotypes — see docs/pgx-cnv.md
for the VCF contract and YAML schema.
Example score model with categories rescaling
score_model:
variants:
rs10012: {effect_allele: G, effect_size: 0.369215857410143}
rs1014971: {effect_allele: T, effect_size: 0.075546961392531}
rs10936599: {effect_allele: C, effect_size: 0.086359830674748}
rs11892031: {effect_allele: C, effect_size: -0.552841968657781}
rs1495741: {effect_allele: A, effect_size: 0.05307844348342}
rs17674580: {effect_allele: C, effect_size: 0.187520720836463}
rs2294008: {effect_allele: T, effect_size: 0.08278537031645}
rs798766: {effect_allele: T, effect_size: 0.093421685162235}
rs9642880: {effect_allele: G, effect_size: 0.093421685162235}
categories:
High risk: {from: 1.371624087, to: 2.581880425, scale_from: 2, scale_to: 3}
Potential risk: {from: 1.169616034, to: 1.371624087, scale_from: 1, scale_to: 2}
Average risk: {from: -0.346748358, to: 1.169616034, scale_from: 0, scale_to: 1}
Low risk: {from: -1.657132197, to: -0.346748358, scale_from: -1, scale_to: 0}
description:
about:
genes: []
result_statement_choice:
Average risk: Avg
Potential risk: Pot
High risk: Hig
Low risk: Low
science_behind_the_test:
test_type: Polygenic Risk Score
trait: Breast cancer
trait_authors:
- taken from the PGS catalog
trait_copyright: Intelliseq all rights reserved
trait_explained: None
trait_heritability: None
trait_pgs_id: PGS000001
trait_pmids:
- 25855707
trait_snp_heritability: None
trait_title: Breast_Cancer
trait_version: 1.0
what_you_can_do_choice:
Average risk:
High risk:
Low risk:
what_your_result_means_choice:
Average risk:
High risk:
Low risk:
Example Formula Model
formula_model:
formula:
brownexp: "math.exp(@brown.score_model.value - 2.0769)"
redexp: "math.exp(@red.score_model.value - 6.3953)"
blackexp: "math.exp(@black.score_model.value - 2.4029)"
sumexp: "@brownexp + @redexp + @blackexp"
brown_prob: "@brownexp / (1 + @sumexp)"
red_prob: "@redexp / (1 + @sumexp)"
black_prob: "@blackexp / (1 + @sumexp)"
blonde_prob: "1 - (@brown_prob + @red_prob + @black_prob)"
brown:
score_model:
variants:
rs796296176: {effect_allele: CA, effect_size: 1.2522}
rs11547464: {effect_allele: A, effect_size: -0.61155}
rs885479: {effect_allele: T, effect_size: 0.2937}
rs1805008: {effect_allele: T, effect_size: -0.50143}
rs1805005: {effect_allele: T, effect_size: 0.21172}
rs1805006: {effect_allele: A, effect_size: 1.9293}
rs1805007: {effect_allele: T, effect_size: -0.32318}
rs1805009: {effect_allele: C, effect_size: 0.60861}
rs1805009: {effect_allele: A, effect_size: 0.25624}
rs2228479: {effect_allele: A, effect_size: -0.054143}
rs1110400: {effect_allele: C, effect_size: -0.56315}
rs28777: {effect_allele: C, effect_size: 0.52168}
rs16891982: {effect_allele: C, effect_size: 0.75284}
rs12821256: {effect_allele: G, effect_size: -0.34957}
rs4959270: {effect_allele: A, effect_size: -0.19171}
rs12203592: {effect_allele: T, effect_size: 1.6475}
rs1042602: {effect_allele: T, effect_size: 0.16092}
rs1800407: {effect_allele: A, effect_size: -0.19111}
rs2402130: {effect_allele: G, effect_size: 0.35821}
rs12913832: {effect_allele: T, effect_size: 1.214}
rs2378249: {effect_allele: C, effect_size: 0.12669}
rs683: {effect_allele: C, effect_size: 0.21172}
red:
score_model:
variants:
rs796296176: {effect_allele: CA, effect_size: 25.508}
rs11547464: {effect_allele: A, effect_size: 2.5381}
rs885479: {effect_allele: T, effect_size: -0.20889}
rs1805008: {effect_allele: T, effect_size: 2.801}
rs1805005: {effect_allele: T, effect_size: 0.93493}
rs1805006: {effect_allele: A, effect_size: 3.65}
rs1805007: {effect_allele: T, effect_size: 3.4408}
rs1805009: {effect_allele: C, effect_size: 4.5868}
rs1805009: {effect_allele: A, effect_size: 22.107}
rs2228479: {effect_allele: A, effect_size: 0.62307}
rs1110400: {effect_allele: C, effect_size: 1.4453}
rs28777: {effect_allele: C, effect_size: 0.70401}
rs16891982: {effect_allele: C, effect_size: -0.41869}
rs12821256: {effect_allele: G, effect_size: -0.57964}
rs4959270: {effect_allele: A, effect_size: 0.24861}
rs12203592: {effect_allele: T, effect_size: 0.90233}
rs1042602: {effect_allele: T, effect_size: 0.45003}
rs1800407: {effect_allele: A, effect_size: -0.27606}
rs2402130: {effect_allele: G, effect_size: 0.28313}
rs12913832: {effect_allele: T, effect_size: -0.093776}
rs2378249: {effect_allele: C, effect_size: 0.76634}
rs683: {effect_allele: C, effect_size: -0.053427}
black:
score_model:
variants:
rs796296176: {effect_allele: CA, effect_size: 2.732}
rs11547464: {effect_allele: A, effect_size: -16.969}
rs885479: {effect_allele: T, effect_size: 0.39983}
rs1805008: {effect_allele: T, effect_size: -0.86062}
rs1805005: {effect_allele: T, effect_size: -0.0029013}
rs1805006: {effect_allele: A, effect_size: -16.088}
rs1805007: {effect_allele: T, effect_size: -1.3757}
rs1805009: {effect_allele: C, effect_size: 0.060631}
rs1805009: {effect_allele: A, effect_size: 3.9824}
rs2228479: {effect_allele: A, effect_size: 0.17012}
rs1110400: {effect_allele: C, effect_size: 0.29143}
rs28777: {effect_allele: C, effect_size: 0.82228}
rs16891982: {effect_allele: C, effect_size: 1.1617}
rs12821256: {effect_allele: G, effect_size: -0.89824}
rs4959270: {effect_allele: A, effect_size: -0.36359}
rs12203592: {effect_allele: T, effect_size: 1.997}
rs1042602: {effect_allele: T, effect_size: 0.065432}
rs1800407: {effect_allele: A, effect_size: -0.49601}
rs2402130: {effect_allele: G, effect_size: 0.26536}
rs12913832: {effect_allele: T, effect_size: 1.9391}
rs2378249: {effect_allele: C, effect_size: -0.089509}
rs683: {effect_allele: C, effect_size: 0.15796}
description:
name: HirisPlex
Description
Model keys glossary
model- generic model that can aggregate results of other model typesdiplotype_modelRequired keys:diplotypes
description- all properties to be included in the final results
Usecases
PGX
python3 -m pip install polygenic-pgx
pgstk pgs-compute --vcf [PATH_TO_VCF_GZ] --model cyp2d6 --print | jq .haplotype_model.haplotypes.match
License
Proprietary (contact@intelliseq.pl)
Updates
2.6.1
Haplotype assembly no longer chains across FORMAT/PS phase sets. The engine never read
FORMAT/PS: it treated any | genotype as globally phased and assigned allele 0 to haplotype 1
across phase-set boundaries. | order is only meaningful within a block, so this asserted
linkage the data does not contain. Measured on NA19207 CYP3A5, where two panels with identical
genotype evidence produced different calls (*3/*7.001 vs *1.001/*3.018) purely because
whatshap emitted one block as 1|0 in one run and 0|1 in the other - the correct answer was
luck. data_accessor now surfaces genotype["phase_set"], and only the largest block is trusted
for orientation; variants outside it take the unphased path, which tries both orders rather than
assuming one. Backward compatible: with no PS anywhere it falls back to the previous whole-VCF
behaviour.
pgstk model-diff: drift detection against a reference model set (BT-2379). Our models drift
from upstream without anyone noticing - RYR1 missing 7 variants and MT-RNR1 differing were both
found by spot-check, not by tooling. This is the detection; which source is authoritative is
a separate decision, and the tool never edits a model.
pgstk model-diff --against /path/to/freshly/built/models
pgstk model-diff --gene ryr1 --against <dir>
pgstk model-diff --against <dir> --json # for CI
Per gene it reports alleles only in ours, alleles missing from ours, alleles whose defining variant set differs (naming the variant keys), and variants whose representation differs - the same site spelled two ways, which is the indel left-alignment class already hit with CYP2A6 and which a plain set difference would report as two unrelated variants. Exits non-zero on any difference, so it can run as a release gate. A gene absent from the reference set is reported but does not fail the gate: the reference legitimately covers a different gene list.
Release automation (BT-2378). A git tag now drives the whole release:
.github/workflows/release.yml runs the suite on Python 3.8, checks that the tag matches
setup.py, verifies the manifest and model checksums, runs the clean-install gate, publishes to
PyPI via Trusted Publishing (no stored token), then builds, smoke-tests and pushes
intelliseqngs/polygenic-pgx:<version>. A new runtime Dockerfile builds the wheel from source
and asserts the models are inside the installed package - the WDL can pin a version instead of
wiring model resources in as task inputs. workflow_dispatch runs the same checks publishing
nothing. See docs/releasing.md.
Per-sample run summary (BT-2376). Every run now also writes
{sample}-pgx-summary.json, so a consumer reads one finished file instead of stitching
per-model JSONs. Additive - the per-model results are untouched; --no-summary disables it.
- Per gene: canonical
gene+gene_hgnc_id, the model's name/source/version/sha256/origin,call,call_core,status,possible_callsand thecopy_numberblock when present, covered/total/missing defining-variant counts, and the uncovered defining positions. - Run level: polygenic version,
manifest_digestnaming the exact model set, the effective parameters,models_not_assayed, per-statuscounts, and aschema_versionfrom day one. statusdistinguishescalled/indeterminate/no_call/not_assayedas four separate values. Collapsing them is the BT-2323 failure: a called MT-RNR1 result was reported as Indeterminate because the consumer could not tell "the engine refused to guess" from "this gene was never on the panel".- Every gene row carries the same keys whatever its status, so a consumer never has to branch before it can read a field.
- Fixed in passing: inside the model loop, a score model's
description.parameterswas assigned to the same name as the run parameters, so one score model could silently become the run parameters for every model after it in the same invocation.
Per-gene run parameters (BT-2375). Every run toggle now takes an optional gene list, so one
invocation can mix policies instead of the caller running pgs-compute again for the one gene
that needs a different flag.
--no-ref-fallback=cyp2c19,cyp2d6and the same shape for--no-safety-lock,--no-indeterminate,--no-ad-arrangement,--all-model-variants. Bare keeps today's global meaning, so nothing breaks. Prefer the--flag=valueform: with optional arity,--flag cyp2c19is ambiguous next to other options.- The same structure is accepted in
-p/--parameters({"per_gene": {"CYP2C19": {"ref_fallback": false}}}), and CLI and JSON combine. Gene names are matched loosely (mt-rnr1==MT-RNR1==mtrnr1) against the model's own gene. - A gene naming no model in the run raises, rather than quietly applying nothing - a typo that does nothing is indistinguishable from the flag working.
- Applied overrides are recorded at
description.parameter_overridesin the result. - Fixed in passing:
ref_fallbackgiven in-p/--parameterswas silently ignored, because the CLI default overwrote it unconditionally. A JSON-supplied value is now honoured.
Panel mode: one invocation calls every gene the VCF covers (BT-2374). -m is now optional.
With neither -m nor --genes, pgs-compute runs every built-in model the input VCF actually
assays, so the caller no longer loops per model.
- Coverage is decided from a per-model footprint (chromosome + min/max of its variant keys,
plus any
copy_number.region) newly stored inmanifest.json: one index query per model, never a trial run. The 10 rsID-keyed models carryprobe_rsidsinstead, because an rsID has no coordinate without a dbSNP build the package does not ship. --genes cyp2c19,cyp2d6keeps the subset case. A gene named explicitly always runs and is never filtered out, so asking for a gene the panel lacks gives a visibleIndeterminaterather than silence. A model whose footprint is unknown (e.g. one fromPOLYGENIC_MODELS_DIR, which has no manifest) also runs - "unknown" is not "absent".- Every decision is logged with its reason, because panel mode makes the gene set implicit and a panel change would otherwise change the report's gene list silently.
- Reads are now shared across models and samples.
VcfAccessorcaches records by variant id, so a position is read once per run instead of once per model per sample - 44x on repeated lookups. The rsID index is also opened once: it was reconnected on every single lookup, and leaked the handle each time (with sqlite3.connect(...)commits, it does not close). The per-modelDataAccessoris deliberately not shared, since its cache holds genotype interpretation that depends on that model'sref_fallbackandreference_defining_variants.
Output naming contract (BT-2377). Three guarantees so no consumer has to parse or rewrite a call string:
- One separator. An assembled diplotype is joined by
/, always. It used to be_for ordinary diplotypes and/only for copy-number calls, which is why the downstream pipeline carried a normalisation step commented "until polygenic uses a consistent separator"./is the safe choice because allele ids legitimately contain_(RYR1c.51_53del,c.992_994dup) and no allele id in any of the 40 packaged models contains/- asserted by reading every model, not assumed. Calls written by older versions are still accepted on input. haplotype_id_corecarries the same call with star sub-allele suffixes dropped (CYP2D6*2.002x2/CYP2D6*4.006x2->CYP2D6*2x2/CYP2D6*4x2). The xN multiplier survives - dropping it changes the activity score, which is the bug in the pipeline's own strip that this replaces. HGVS-style names are untouched (a blanket strip turnsc.51_53delintoc51_53del), a haploid collapse stays collapsed, and verdicts (Indeterminate) pass through.description.gene_hgnc_id(e.g.HGNC:2625) alongsidedescription.gene, so reports join on a stable identifier rather than a symbol. Sourced from a committedhgnc.json- no network at runtime; an unknown symbol yieldsnull, never a guess.
2.5.37
The Indeterminate verdict was unreachable for whole genes, and RYR1 was the one that mattered.
A wild-type call resting on allele-defining positions that were never covered was reported as a
confident normal result. On one validated 8-sample run every RYR1 row read
Reference/Reference -> uncertain susceptibility while 5-6 malignant-hyperthermia defining
positions sat at read depth 0-1.
Two independent causes, and fixing either alone changes nothing:
core(function-level) was decided by "the haplotype name contains a dot". That marks a sub-allele in star nomenclature (CYP2D6*1.001), but the PharmGKB/CPIC-derived models name haplotypes in HGVS - RYR1c.10042C>T, CACNA1Sc.3257G>A, MT-RNR1m.1095T>C- which always contains a dot. So 339 of RYR1's 340 haplotypes counted as sub-alleles, every variant became non-core, andindeterminate.defining_totalcollapsed to 0 - which the verdict is gated on. Now only a trailing numeric suffix on a star/rsID allele marks a sub-allele (is_suballele_name), which also keepsDPYD*rs112766203.1correctly non-core._is_wildtype_pairmatched only*1, soReference_Referencewas never recognised as a wild-type call in the first place. It now accepts aReferenceallele too.
Related: a model that declares only sub-alleles (CYP4F2 is just *1.001..*4.001; likewise
CYP2J2, CYP3A7, CYP2S1, CYP2W1, CYP3A43, CYP26A1, CYP2F1, CYP2R1) has nothing to contrast against,
so nothing was core there either. When no variant in a model is core, every variant is defining.
- Scoring is unchanged, and that is checked rather than assumed.
corealso drives the weighted haplotype match (weight * 0.05), butpercent_weight_matchis a ratio, so a flip of every genotype in a model cancels. Only uniform flips are safe, and all 12 affected models flip uniformly. An A/B of all 40 packaged models on one sample, old engine vs new on the same tree: 36 results byte-identical, exactly one call moves (RYR1 ->Indeterminate, withcall_filledkeeping the best-effort answer), and 3 bookkeeping-only changes wheredefining_totalbecomes honest (CACNA1S and CYP4F2 0 -> 2, G6PD 181 -> 182) without triggering. INFO/MODEL_FILLis carried into the genotype asmodel_fill/model_fill_depth. A pipeline that fills model sites from read depth before polygenic produces genotypes that read assource: genotyping, making a depth-verified assumption indistinguishable from a measurement. Recorded as its own field, deliberately not folded intosource:compute_qcrejects an unknown source, and a fill that passed its depth gate is real evidence, so no verdict moves.indeterminate's reason string now names the actual call instead of always saying*1/*1.
2.5.36
The packaged models now cover the whole PGx report (33 -> 40 models). Seven genes existed only in the deployed model streams, so anything switching a pipeline to package-provided models would have silently dropped them from every report.
- Added NAT2 (
nat2-pharmvar-6.2.25.yml, 129 sub-alleles / 59 core). The version was established by checksum: the deployedmodels-pharmvar/06.2026export IS PharmVar 6.2.25 - its CYP2A6 differs from ours only in indel left-alignment. - Added ABCG2, CACNA1S, CFTR, G6PD, RYR1, UGT1A1 (
*-pharmgkb-1.0.0.yml), vendored from the deployed PharmGKB-derived set; each header records the source path and the sha256 of the unchanged body. These are not star-allele models: G6PD names alleles as variant combinations and is X-linked, CFTR by ivacaftor responsiveness class, RYR1 has 340 single-variant alleles, and UGT1A1 carries the(TA)npromoter repeat (*28/*36/*37). - All their indels (CFTR 1, G6PD 10, RYR1 4, UGT1A1 3) were checked against GRCh38 with
bcftools normand were already left-aligned, so nothing was rewritten. - Every added model was validated against real output: NAT2 matches an external review sheet on
4/5 samples (the fifth is PharmVar vs legacy naming,
*5/*16vs*5/*5), and the six others reproduce the production star-allele JSON exactly. - Manifest fix: a model added by copying an existing file is now dated at the copy. git
reported NAT2 as a copy (
C096) of a file vendored elsewhere a day earlier, and the manifest dated the model before it existed in the package. A copy is not a move - the source still exists.
Older entries: see CHANGELOG.md, which ships with the source distribution and holds the full release history.
Metadata
Release files for polygenic-pgx 2.6.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| polygenic_pgx-2.6.1.tar.gz | 216.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| polygenic_pgx-2.6.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 414.1 kB
Release files / polygenic_pgx-2.6.1.tar.gz
| Download URL | polygenic_pgx-2.6.1.tar.gz |
|---|---|
| Size | 216.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
9c414345c0981ca342900f17697f1169e51e6574aa37156efed203a3f38c3b1a
|
|
BLAKE2b-256 checksum How to use checksums |
fd45688340e2f533dca522dcd36f4867e7368de9f270851deb5629894c34c60c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.3
|
Release files / polygenic_pgx-2.6.1-py3-none-any.whl
| Download URL | polygenic_pgx-2.6.1-py3-none-any.whl |
|---|---|
| Size | 197.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
4873f3dd62a0f97db40b6ab853498f604ace33c5f802225a4de414bade91eee2
|
|
BLAKE2b-256 checksum How to use checksums |
f8385c12cf2a1d00a72ec4459f6bf157dbd0bda124f289bbd81e4125bda7e4c1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.3
|