Skip to main content

A collection of handy tools for GWAS SumStats

Project description

GWASLab

image

badge Downloads badge_pip badge_commit_m

  • A handy Python-based toolkit for handling GWAS summary statistics (sumstats).
  • Each process is modularized and can be customized to your needs.
  • Sumstats-specific manipulations are designed as methods of a Python object, gwaslab.Sumstats.

Installation

install via pip

The latest version of GWASLab now supports Python 3.9, 3.10, 3.11, and 3.12.

pip install gwaslab

install in conda environment

Create a Python 3.9, 3.10, 3.11 or 3.12 environment and install gwaslab using pip:

conda env create -n gwaslab -c conda-forge python=3.12

conda activate gwaslab

pip install gwaslab

or create a new environment using yml file environment.yml

conda env create -n gwaslab -f environment.yml

install using docker (deprecated)

A docker file is available here for building local images.

Quick start

import gwaslab as gl

# load plink2 output
mysumstats = gl.Sumstats("sumstats.txt.gz", fmt="plink2")

# or load sumstats with auto mode (auto-detecting commonly used headers) 
# assuming ALT/A1 is EA, and frq is EAF
mysumstats = gl.Sumstats("sumstats.txt.gz", fmt="auto")

# or you can specify the columns:
mysumstats = gl.Sumstats("sumstats.txt.gz",
             snpid="SNP",
             chrom="CHR",
             pos="POS",
             ea="ALT",
             nea="REF",
             eaf="Frq",
             beta="BETA",
             se="SE",
             p="P",
             direction="Dir",
             n="N",
             build="19")

# manhattan and qq plot
mysumstats.plot_mqq()
...

Documentation and tutorials

Documentation and tutorials for GWASLab are avaiable at here.

Functions

Loading and Formatting

  • Loading sumstats by simply specifying the software name or format name, or specifying each column name.
  • Converting GWAS sumstats to specific formats:
    • LDSC / MAGMA / METAL / PLINK / SAIGE / REGENIE / MR-MEGA / GWAS-SSF / FUMA / GWAS-VCF / BED...
    • check available formats
  • Optional filtering of variants in commonly used genomic regions: Hapmap3 SNPs / High-LD regions / MHC region

Standardization & Normalization

  • Variant ID standardization
  • CHR and POS notation standardization
  • Variant POS and allele normalization
  • Genome build : Inference and Liftover

Quality control, Value conversion & Filtering

  • Statistics sanity check
  • Extreme value removal
  • Equivalent statistics conversion
    • BETA/SE , OR/OR_95L/OR_95U
    • P, Z, CHISQ, MLOG10P
  • Customizable value filtering

Harmonization

  • rsID assignment based on CHR, POS, and REF/ALT
  • CHR POS assignment based on rsID using a reference text file
  • Palindromic SNPs and indels strand inference using a reference VCF
  • Check allele frequency discrepancy using a reference VCF
  • Reference allele alignment using a reference genome sequence FASTA file

Visualization

  • Mqq plot: Manhattan plot, QQ plot or MQQ plot (with a bunch of customizable features including auto-annotate nearest gene names)
  • Miami plot: mirrored Manhattan plot
  • Brisbane plot: GWAS hits density plot
  • Regional plot: GWAS regional plot
  • Genetic correlation heatmap: ldsc-rg genetic correlation matrix
  • Scatter plot: variant effect size comparison
  • Scatter plot: allele frequency comparison
  • Scatter plot: trumpet plot (plot of MAF and effect size with power lines)

Visualization Examples

image image image image

Other Utilities

  • Read ldsc h2 or rg outputs directly as DataFrames (auto-parsing).
  • Extract lead variants given a sliding window size.
  • Extract novel loci given a list of known lead variants / or known loci obtained from GWAS Catalog.
  • Logging: keep a complete record of manipulations applied to the sumstats.
  • Sumstats summary: give you a quick overview of the sumstats.
  • ...

Issues

How to cite

  • GWASLab preprint: He, Y., Koido, M., Shimmori, Y., Kamatani, Y. (2023). GWASLab: a Python package for processing and visualizing GWAS summary statistics. Preprint at Jxiv, 2023-5. https://doi.org/10.51094/jxiv.370

Sample data used for tutorial

  • Sample GWAS data used in GWASLab is obtained from: http://jenger.riken.jp/ (Suzuki, Ken, et al. "Identification of 28 new susceptibility loci for type 2 diabetes in the Japanese population." Nature genetics 51.3 (2019): 379-386.).

Acknowledgement

Thanks to @sup3rgiu, @soumickmj and @gmauro for their contributions to the source codes.

Contacts

Project details


Release history Release notifications | RSS feed

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

gwaslab-3.6.11.tar.gz (20.9 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

gwaslab-3.6.11-py3-none-any.whl (20.9 MB view details)

Uploaded Python 3

File details

Details for the file gwaslab-3.6.11.tar.gz.

File metadata

  • Download URL: gwaslab-3.6.11.tar.gz
  • Upload date:
  • Size: 20.9 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.0

File hashes

Hashes for gwaslab-3.6.11.tar.gz
Algorithm Hash digest
SHA256 8d7e9712d911bf45a7178db2d004fe0439c82b3bc51448fb8de733dc4d0138d2
MD5 433bf46cc8e38d4e2c63ba9c546c3489
BLAKE2b-256 f1a06062aedcfd820fe75409ef53c5f01d6612e2054f860476876265a1a4a66f

See more details on using hashes here.

File details

Details for the file gwaslab-3.6.11-py3-none-any.whl.

File metadata

  • Download URL: gwaslab-3.6.11-py3-none-any.whl
  • Upload date:
  • Size: 20.9 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.0

File hashes

Hashes for gwaslab-3.6.11-py3-none-any.whl
Algorithm Hash digest
SHA256 3d5fdbf3964a4f9e3f4dfe9f13c06c4f31c2c75be6915f6c59e600e7a1dea047
MD5 5143a644d8cd477676ce0476fc191c34
BLAKE2b-256 44b7976eac9e8d99d3136febfc45f9bbb11ce8cd6715d70c9d0a6570608474c4

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page