Generic toolkit for processing DNA polymorphism data

These details have not been verified by PyPI

Project links

Homepage

License
- OSI Approved :: GNU General Public License v3 (GPLv3)
Operating System
- OS Independent
Programming Language
- Python :: 3

Project description

AdaGenes

AdaGenes is a generic toolkit for processing, annotating, filtering and transforming DNA polymorphism data.

Main features:

A powerful data object to store and edit DNA mutation data
Functionality to read and write files in common genomics file formats, including VCF, MAF, CSV/TSV, XLSX and plain text files
Effective variant filtering according to specific threshold or feature values
Liftover genome positions between hg38/GRCh38, hg19/GRCh37 and T2T-CHM13 reference genomes
Effective variant normalization in VCF and HGVS notation

Installation

AdaGenes is both usable as a Python package or directly from the command line. You can install AdaGenes in Python directly via PyPI:

pip install adagenes

Getting started

Reading files

Start by reading in a data file in one of the supported file formats in a biomarker frame with the read_file() function. adagenes automatically identifies the file type and inititates the corresponding file reader. You may also manually inititate a file reader and call its read_file() function:

import adagenes as ag

bframe = ag.read_file("data/somaticMutations.vcf")

# Print biomarker identifiers
print(bframe.get_ids())

# Print loaded variant data completely
print(bframe.data)

Instead of loading a variant file, you may also create a biomarker frame manually at genomic or protein level:

import adagenes as ag

# create biomarker frame based on variants at genomic level
bframe = ag.BiomarkerFrame(data=["chr7:g.140753336A>T"])

If the variant data has been parsed correctly, the data of the biomarker frame should be a nested JSON dictionary:

{
'chr7:140753336A>T': {'variant_data': {'CHROM': '7', 'POS': '140753336', 'ID': '.', 'REF': 'A', 'ALT': 'T', 'QUAL': '100', ... },
'chr1:2556664C>.': {'variant_data': {'CHROM': '1', 'POS': '2556664', 'ID': '.', ... } }
}

Liftover

Convert the genomic positions of variants between genome assemblies with the liftover function (GRCh37 / GRCh38 / T2T-CHM13):

For large variant files, you can use the AdaGenes process_file() function for stream-based processing:

import adagenes as ag

infile = "somaticMutations.vcf"
outfile = "somaticMutations.t2t.vcf"

client = ag.LiftoverClient(genome_version="hg19", target_genome="t2t")
ag.process_file(infile, outfile, client)

For small to medium sized variant files, you can load and edit the variant data as a biomarker frame:

import adagenes as ag

# Load a biomarker frame by defining the genome version (hg19/hg38/t2t)
infile = "somaticMutations.vcf"
bframe = ag.read_file(infile, genome_version="hg38")

# Liftover to another genome assemly
bframe_t2t = ag.liftover(bframe, target_genome="t2t")

# Write the new biomarker frame in T2T to a file
ag.write_file("somaticMutations.t2t.vcf", bframe_t2t)

Filter mutations

Annotate variants

Use Onkopus to annotate variants from the command line, e.g.

import adagenes as ag
import onkopus as op

bframe = ag.read_file("somaticMutations.vcf", genome_version="hg38")

bframe.data = op.PathogenicityClient(genome_version="hg38").process_data(bframe.data)

ag.write_file(bframe, "somaticMutations.annotated.vcf")

For further details on how to annotate variants, check out the Onkopus documentation.

Variant notations and normalization

Visualization

Annotate variants

You can easily annotate variant data by combining an AdaGenes biomarker frame with the Onkopus annotation framework:

pip install onkopus

Annotate the variant data of a biomarker frame by calling an Onkopus client directly on the bframe.data:

import adageness as av
import onkopus as op

genome_version="hg38"
bframe = av.read_file("somaticMutations.vcf", genome_version="hg38")

# Annotate with all Onkopus modules
bframe.data = op.annotate(bframe.data)

# Annotate with specific modules
bframe.data = op.AlphaMissenseClient(genome_version=genome_version).process_data(bframe.data)
bframe.data = op.GENCODEClient(genome_version=genome_version).process_data(bframe.data)

av.write_file("somaticMutations.annotated.avf",bframe)

Saving data

Write a biomarker frame to a file with write_file() in one of the supported file formats (.vcf,.maf,.csv):

import adagenes as ag

ag.write_file("/data/somaticMutations.annotated.maf", bframe, file_type="csv")

Dependencies

scikit-learn
pandas
matplotlib
plotly
pyliftover
blosum
openpyxl
requests

License

GPLv3

Project details

These details have not been verified by PyPI

Project links

Homepage

License
- OSI Approved :: GNU General Public License v3 (GPLv3)
Operating System
- OS Independent
Programming Language
- Python :: 3

Release history Release notifications | RSS feed

0.5.3

Sep 9, 2025

0.5.2

Sep 1, 2025

0.5.1

Jul 29, 2025

0.5.0

Jul 16, 2025

0.4.9

Jul 9, 2025

0.4.8

Jul 3, 2025

0.4.7

Jul 2, 2025

0.4.6

Jul 2, 2025

0.4.5

Jun 25, 2025

0.4.4

Jun 13, 2025

0.4.3

May 24, 2025

0.4.2

May 24, 2025

0.4.0

Mar 26, 2025

0.3.9

Mar 23, 2025

0.3.8

Mar 18, 2025

0.3.7

Feb 13, 2025

0.3.6

Feb 3, 2025

0.3.5

Feb 3, 2025

0.3.4

Jan 30, 2025

0.3.3

Jan 29, 2025

0.3.2

Jan 23, 2025

0.3.1

Jan 23, 2025

0.3.0

Jan 21, 2025

0.2.9

Jan 16, 2025

0.2.8

Jan 15, 2025

0.2.7

Jan 9, 2025

0.2.6

Jan 7, 2025

0.2.5

Dec 13, 2024

0.2.4

Nov 28, 2024

0.2.3

Nov 26, 2024

This version

0.2.2

Nov 25, 2024

0.2.1

Nov 19, 2024

0.2.0

Nov 15, 2024

0.1.9

Oct 8, 2024

0.1.8

Oct 8, 2024

0.1.7

Oct 8, 2024

0.1.6

Oct 7, 2024

0.1.5

Oct 1, 2024

0.1.4

Sep 26, 2024

0.1.2

Sep 26, 2024

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

adagenes-0.2.2.tar.gz (12.5 MB view details)

Uploaded Nov 25, 2024 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

adagenes-0.2.2-py3-none-any.whl (436.6 kB view details)

Uploaded Nov 25, 2024 Python 3

File details

Details for the file adagenes-0.2.2.tar.gz.

File metadata

Download URL: adagenes-0.2.2.tar.gz
Upload date: Nov 25, 2024
Size: 12.5 MB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/5.1.1 CPython/3.10.12

File hashes

Hashes for adagenes-0.2.2.tar.gz
Algorithm	Hash digest
SHA256	`aad965c89f4ce737a4f26a7b1851ba2c53c970b39de7dd1fa6e52b4cb2bc2acd`
MD5	`3b124749adc0594f9c14d8f6787db137`
BLAKE2b-256	`2ea556ef18cb234c133ddf3727d01d719ef88948a096e32efb83c2267ffc1fd8`

See more details on using hashes here.

File details

Details for the file adagenes-0.2.2-py3-none-any.whl.

File metadata

Download URL: adagenes-0.2.2-py3-none-any.whl
Upload date: Nov 25, 2024
Size: 436.6 kB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: twine/5.1.1 CPython/3.10.12

File hashes

Hashes for adagenes-0.2.2-py3-none-any.whl
Algorithm	Hash digest
SHA256	`0144778045e9d37cdff2746dbb0adf30cdca8718c3e92d82bebe54415324cc96`
MD5	`63b12d1b90cbacf42a5b8345c28cb976`
BLAKE2b-256	`f53cf43f1bfa70f759389ef7fe52b86a30df9037f909a6bd6623cd06a8ad92f0`

See more details on using hashes here.

adagenes 0.2.2

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

AdaGenes

Main features:

Installation

Getting started

Reading files

Liftover

Filter mutations

Annotate variants

Variant notations and normalization

Visualization

Annotate variants

Saving data

Dependencies

License

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes