Skip to main content

DNARecords

PyPI license example workflow codecov pylint Score semantic-release: angular

Genomics data ML ready.

Transform your vcf, bgen, etc. genomics datasets into a sample wise format so that you can use it for Deep Learning models.

Installation

DNARecords package has two main dependencies:

  • Hail, if you are transforming your genomics data into DNARecords
  • Tensorflow, if you are using a previously DNARecords dataset, for example, to train a DL model

As you may know, Tensorflow and Spark does not play very well together on a cluster with more than one machine.

However, dnarecords package needs to be installed only on the driver machine of a Hail cluster.

For that reason, we recommend following these installation tips.

On a dev environment

$ pip install dnarecords

For further details (or any trouble), review Local environments section.

On a Hail cluster or submitting a job to it

You will already have Pyspark installed and will not intend to install Tensorflow.

So, just install dnarecords without dependencies on the driver machine.

There will be no missing modules as soon as you use the classes and functionality intended for Spark.

$ /opt/conda/miniconda3/bin/python -m pip install dnarecords --no-deps

Note: assuming Hail python executable is /opt/conda/miniconda3/bin/python

On a Tensorflow environment or submitting a job to it

You will already have Tensorflow installed and will not intend to install Pyspark.

So, just install dnarecords without dependencies.

There will be no missing modules as soon as you use the classes and functionality intended for Tensorflow.

$ pip install dnarecords --no-deps

Working on Google Dataproc

Just use and initialization action that installs dnarecords without dependencies.

$ hailctl dataproc start dnarecords --init gs://dnarecords/dataproc-init.sh

Iy you need to work with other cloud providers, refer to Hail docs.

Usage

It is quite straightforward to understand the functionality of the package.

Given some genomics data, you can transform it into a DNARecords Dataset this way:

import dnarecords as dr


hl = dr.helper.DNARecordsUtils.init_hail()
hl.utils.get_1kg('/tmp/1kg')
mt = hl.read_matrix_table('/tmp/1kg/1kg.mt')
mt = mt.annotate_entries(dosage=hl.pl_dosage(mt.PL))

dnarecords_path = '/tmp/dnarecords'
writer = dr.writer.DNARecordsWriter(mt.dosage)
writer.write(dnarecords_path, sparse=True, sample_wise=True, variant_wise=True,
             tfrecord_format=True, parquet_format=True,
             write_mode='overwrite', gzip=True)

Given a DNARecords Dataset, you can read it as Tensorflow Datasets this way:

import dnarecords as dr


dnarecords_path = '/tmp/dnarecords'
reader = dr.reader.DNARecordsReader(dnarecords_path)
samplewise_ds = reader.sample_wise_dataset()
variantwise_ds = reader.variant_wise_dataset()

Or, given a DNARecords Dataset, you can read it as Pyspark DataFrames this way:

import dnarecords as dr


dnarecords_path = '/tmp/dnarecords'
reader = dr.reader.DNASparkReader(dnarecords_path)
samplewise_df = reader.sample_wise_dnarecords()
variantwise_df = reader.variant_wise_dnarecords()

We will provide more examples and integrations soon.

Contributing

Interested in contributing? Check out the contributing guidelines. Please note that this project is released with a Code of Conduct. By contributing to this project, you agree to abide by its terms.

License

dnarecords was created by Atray Dixit, Andrés Mañas Mañas, Lucas Seninge. It is licensed under the terms of the MIT license.

Credits

dnarecords was created with cookiecutter and the py-pkgs-cookiecutter template.

Release files for dnarecords 0.2.7

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for dnarecords 0.2.7
File Size Uploaded
dnarecords-0.2.7.tar.gz 14.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for dnarecords 0.2.7
File Interpreter ABI Platform
dnarecords-0.2.7-py3-none-any.whl Python 3 none any Details

Total release size: 29.4 kB

Release files / dnarecords-0.2.7.tar.gz

Download URL dnarecords-0.2.7.tar.gz
Size 14.6 kB
Tags Source
SHA-256 checksum
How to use checksums
7075face459d81ebd0b729052313a4a75859fe83f171ee24543028707d80d882
BLAKE2b-256 checksum
How to use checksums
3a6f1f8f84f787074e0df87efa189e6b9205ccf0b7972f18374e2b0b344718f8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.1 CPython/3.9.13

Release files / dnarecords-0.2.7-py3-none-any.whl

Download URL dnarecords-0.2.7-py3-none-any.whl
Size 14.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a08f6a9aa4da9f6df0f8a389662e14f9f53546c5ce395ddc378b780e541b53a6
BLAKE2b-256 checksum
How to use checksums
5ebfff9c3870e5b798a353f0a317748957896f7faccc311572c0a845a3169fad
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.1 CPython/3.9.13

Release history Release notifications | RSS feed

This release

0.2.7 This release

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page