Skip to main content

PRSedm (Polygenic Risk Score Extension for Diabetes Mellitus) is a package for local and remote generation of Polygenic Risk Scores (PRS) for Diabetes Mellitus (DM)

Reason this release was yanked:

Older version with naming error

Project description

PRSedm

Graphical Abstract

Overview

PRSedm (Polygenic Risk Score Extension for Diabetes Mellitus) is a flexible and extendable open-source package for efficient local and remote (All of Us, UK Biobank, etc.) generation of published Polygenic Risk Scores (PRS) for Diabetes Mellitus (DM) and related cardiometabolic phenotypes.

PRS for Type 1 diabetes (T1D) and Type 2 diabetes (T2D), and more recent partitioned Polygenic Scores (pPS), have numerous applications as research and clinical tools.

PRSedm aims to introduce a new parallelized "one-liner" method to generate standardized PRS and pPS for DM robust to variables such as genotyping method, quality control, and imputation panel.

Installation

Dependencies

PRSedm requires the following packages:

  • Python (>=3.9)
  • Joblib (>=1.3.2)
  • Pandas (>=2.2.3)
  • Pysam (>=0.22.0)
  • Numpy* (2.x/1.x)

*Build with 1.x when deploying to RAP platforms with 1.x dependencies.

User Installation

PRSedm is available through a number of channels:
PIP: not yet available
Anaconda: not yet available
Build from source: not yet available

Usage

Command Line Interface

To call PRSedm from the command line:

prsedm --bcf <path_to_bcf_file> [options]
  • --vcf (required): Path to an indexed VCF (.gz/.bgz) or BCF file, or a text file mapping one file per contig.
  • --col: Genotype column to score (default: GT, options: GT for WGS, GP for imputed data).
  • --build: Genome build to use (default: hg38, options: hg19, hg38).
  • --prsflags (required): Comma-separated list of PRS to generate, e.g., PRS1,PRS2.
  • --impute (optional): Enable imputation (requires --refvcf) (default: 1).
  • --refvcf (optional): Reference directory (required if imputation is enabled).
  • --norm (optional): Perform fixed MinMax normalization (default: 1).
  • --parallel (optional): Enable parallel processing (default: 1).
  • --ntasks (optional): Number of tasks to use for parallel processing (default: CPU count).
  • --batch-size (optional): Number of variants per batch (default: 1).
  • --output: Path to save the output file (default: results.csv).

Python (recommended)

To call PRSedm from python:

import PRSedm
df = prsedm.gen_dm(vcf, col, build, prsflags, impute, refvcf, norm, parallel, ntasks, batch_size)

Single file per-chromosome loading

For --vcf and --refvcf you can point to a text a single file mapping per contiguous region formatted as such:

file1.chr1.vcf.gz   chr1
file2.chr2.vcf.gz   chr2
...

Research Analysis Platforms (RAP's)

Remote deployment to remote Research Analysis Platforms (RAP's) is possible via notebook wrappers:

  • All of Us - Notebook Here
  • UK Biobank - Notebook not yet available

Available PRS

PRS schematics are stored in a SQLlite database (variants.db) and accessed via a JSON metadata (prs_meta.json). These are automatically read and fully extendable if the user wishes to add additional PRS. The following PRS are available by default:

Type 1 Diabetes

Flag Method Variants Description PMID
t1dgrs2-sharp24 HLA Interaction + Partitioned 67 "GRS2x" updated PRS with widest compatibility and HLA-based risk pPS. TBA
t1dgrs2-qu22 HLA Interaction + Partitioned 71 Original "GRS2" PRS with the addition of 4 African ancestry SNPs from Onengut, proposed in Qu et al and utilized in eMERGE. 34997821
t1dgrs2-sharp21 HLA Interaction + Partitioned 67 Version of "GRS2" PRS designed for "TOPMED-R2" from 2021 GitHub. 35312757
t1d-onengut19-afr Additive 6 African-ancestry PRS proposed by Onengut in 2019, updated for modern compatibility. 30659077
t1dgrs2-sharp19 HLA Interaction + Partitioned 67 Original 1000 Genomes version of "GRS2" PRS as published, with limited modern compatibility.  30655379

Type 2 Diabetes

Flag Method Variants Description PMID
t2dp_suzuki24_ma Additive + Partitioned 1289 Multiancestry weighted Suzuki T2D index variant PRS, and pPS from hard-clustering analyses. 38374256
t2dp_suzuki24_<ancestry> Additive + Partitioned 1128 - 1285 As above but weighted for specific ancestries <eur/afr/safr/eas/sas/his> 38374256
t2dp_smith24_ma Additive + Partitioned 353 Multiancestry cluster-weighted Smith T2D index variant PRS, and pPS from soft-clustering analyses. 38443691
t2dp_smith24_<ancestry> Additive + Partitioned 25 - 490  As above but from ancestry-specific soft clustering <eur/afr/eas/amr>. 38443691
t2d_mahajan22_ma Additive 338 Older PRS from Mahajan et al composed of multiancestry index variants. 35551307
t2dp_udler18 Additive + Partitioned 67 T2D pPS from first soft-clustering analysis. 30240442

Other

Flag Phenotype Method Variants Description PMID
cdgrs_sharp_24 Celiac Disease HLA Interaction + Partitioned 42 Modernized Celiac disease PRS and pPS with similar model to "GRS2", utilized for combined screening. 32790217

Additional Features

HLA Interaction PRS (+GRS2x Update)

PRSedm features a complete algorithm for GRS which incorporate HLA interaction terms as previously published by us such as T1D-GRS2 (or just GRS2). A number of advancements have been added to improve the generation of HLA interaction, described as GRS2x.

HLA Type Estimation and LD Tiebreak

HLA alleles can be estimated by proxy (or tag) single nucleotide polymorphisms alone and predictions are output e.g. (DR3-DQ2.5/DR3-DQ2.5) Due to imperfect proxy SNPs >2 HLA calls can be made in interaction scores such as GRS2, a probablistic tiebreaker algorithm using a HLA reference frequencies (Klitz et al) now resolves impossible numbers of calls without excluding any samples.

Missing variant mean effect imputation (optional)

PRSedm optionally uses Hardy-Weinberg Equilibrium with a reference VCF/BCF legend (ensure you have variant frequency coded as 'AF', genotypes not required) to impute the mean effect size for missing SNPs, handle missing variants, and enable static normalization.

Minimum and Maximum Normalization (optional)

PRSedm hardcodes static normalization of minimum and maximum potential risk contribution (no risk alleles vs all risk alleles) creating a scale of 0-1. Static normalization with imputation ensures that PRS values translate to a common relative risk scale across datasets. Imputation must be enabled, or all variants must be present.

Development

Developed and maintained by Seth A. Sharp (ssharp@stanford.edu) at the Translational Genomics of Diabetes lab (Dr. Anna Gloyn), Stanford University, with collaboration from colleagues at the University of Exeter (Amber Luckett, Dr. Michael Weedon, Dr. Richard Oram), and MGH/Broad Institute (Dr. Aaron Deutsch, Dr. Miriam Udler).

License

This project is licensed under the MIT License (Non-Commercial).

  • Academic, research, and personal use are allowed.
  • Commercial use is prohibited without prior permission.

See the LICENSE file for full details.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

prsedm-1.0.1b0.tar.gz (345.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

prsedm-1.0.1b0-py3-none-any.whl (345.5 kB view details)

Uploaded Python 3

File details

Details for the file prsedm-1.0.1b0.tar.gz.

File metadata

  • Download URL: prsedm-1.0.1b0.tar.gz
  • Upload date:
  • Size: 345.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/5.1.1 CPython/3.12.7

File hashes

Hashes for prsedm-1.0.1b0.tar.gz
Algorithm Hash digest
SHA256 87c402437cd20004ec9f60dea12b77300acfb3d113a7c970adb62d923c89e864
MD5 f84bd7faf1d8efc94f0713f02db63efe
BLAKE2b-256 0c77a6680b1f03ee3f88d27b2945cc33f887109f964513c524e5522a18cca42f

See more details on using hashes here.

Provenance

The following attestation bundles were made for prsedm-1.0.1b0.tar.gz:

Publisher: python-publish.yml on sethsh7/PRSedm

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file prsedm-1.0.1b0-py3-none-any.whl.

File metadata

  • Download URL: prsedm-1.0.1b0-py3-none-any.whl
  • Upload date:
  • Size: 345.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/5.1.1 CPython/3.12.7

File hashes

Hashes for prsedm-1.0.1b0-py3-none-any.whl
Algorithm Hash digest
SHA256 ff1b5fd181dff22946dc22bc1e71d69adb38984130cfd91f12e16f59b47ea8c1
MD5 002b475469f62b82ba0693c5670b49fd
BLAKE2b-256 3c38aaab3ebeeee438fe13334e5307182002c9d76137822f4ab3d5347f9f52b5

See more details on using hashes here.

Provenance

The following attestation bundles were made for prsedm-1.0.1b0-py3-none-any.whl:

Publisher: python-publish.yml on sethsh7/PRSedm

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page