Genomic Arrays based on TileDB
GenomicArrays is a Python package for converting genomic data from BigWig format to TileDB arrays.
Installation
Install the package from PyPI
pip install genomicarrays
Quick Start
Build a GenomicArray
Building a GenomicArray generates 3 TileDB files in the specified output directory:
feature_annotation: A TileDB file containing input feature intervals.sample_metadata: A TileDB file containing sample metadata, each BigWig file is considered a sample.- A matrix TileDB file named by the
layer_matrix_nameparameter. This allows the package to store multiple different matrices, e.g. 'coverage', 'some_computed_statistic', for the same interval, and sample metadata attributes.
The organization is inspired by the SummarizedExperiment data structure. The TileDB matrix file is stored in a features X samples orientation.
To build a GenomicArray from a collection of BigWig files:
import numpy as np
import tempfile
import genomicarrays as garr
# Create a temporary directory, this is where the
# output files are created. Pick your location here.
tempdir = tempfile.mkdtemp()
# List BigWig paths
bw_dir = "your/biwig/dir"
files = os.listdir(bw_dir)
bw_files = [f"{bw_dir}/{f}" for f in files]
features = pd.DataFrame({
"seqnames": ["chr1", "chr1"],
"starts": [1000, 2000],
"ends": [1500, 2500]
})
# Build GenomicArray
dataset = garr.build_genomicarray(
files=bw_files,
output_path=tempdir,
features=features,
# Specify a fasta file to extract sequences
# for each region in features
genome_fasta="path/to/genome.fasta",
# agg function to summarize mutiple values
# from bigwig within an input feature interval.
feature_annotation_options=garr.FeatureAnnotationOptions(
aggregate_function = np.nanmean
),
# for parallel processing multiple bigwig files
num_threads=4
)
Query a GenomicArrayDataset
Users have the option to reuse the dataset object retuned when building the arrays or by creating a GenomicArrayDataset object by initializing it to the path where the files were created.
# Create a GenomicArrayDataset object from the existing dataset
dataset = GenomicArrayDataset(dataset_path=tempdir)
# Query data for the first 10 regions across all samples
coverage_data = dataset[0:10, :]
print(expression_data.matrix)
print(expression_data.feature_annotation)
## output 1
array([[1. , 0.5],
[1. , 0.5],
[1. , 0.5],
[1. , 0.5],
[1. , 0.5],
[1. , 0.5],
[1. , 0.5],
[1. , 0.5],
[1. , 0.5],
[1. , 0.5],
[1. , nan]], dtype=float32)
## output 2
seqnames starts ends genarr_feature_index
0 chr1 300 315 0
1 chr1 320 335 1
2 chr1 340 355 2
3 chr1 360 375 3
4 chr1 380 395 4
5 chr1 400 415 5
6 chr1 420 435 6
7 chr1 440 455 7
8 chr1 460 475 8
9 chr1 480 495 9
10 chr1 500 515 10
Note
This project has been set up using PyScaffold 4.6. For details and usage information on PyScaffold see https://pyscaffold.org/.
Metadata
Release files for GenomicArrays 0.2.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| genomicarrays-0.2.2.tar.gz | 113.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| GenomicArrays-0.2.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 133.4 kB
Release files / genomicarrays-0.2.2.tar.gz
| Download URL | genomicarrays-0.2.2.tar.gz |
|---|---|
| Size | 113.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
27a974dbd1d46907f460e9fcebe483815ba88b55a5adc9bc378235794ff19a97
|
|
BLAKE2b-256 checksum How to use checksums |
c7806d162c08865d481ac8356294b672fc70afd471ca52c72a28e6e2ed890987
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.12.8
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jan 30, 2025.
Transparency logRelease files / GenomicArrays-0.2.2-py3-none-any.whl
| Download URL | GenomicArrays-0.2.2-py3-none-any.whl |
|---|---|
| Size | 19.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a97007a61f2e0bc0efd569f8f1877704687c0a3a3b5563c6495df8e6bb04c9d5
|
|
BLAKE2b-256 checksum How to use checksums |
4b20b47f2639674db48f5f0a639044046f139b7dadb18fade8595351df99e0f7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.12.8
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jan 30, 2025.
Transparency log