Skip to main content

A collection of datasets and predictors for benchmarking miRNA target site prediction algorithms

Project description

miRNA target site prediction Benchmarks

Installation

miRBench package can be easily installed using pip:

pip install miRBench

Default installation allows access to the datasets. To use predictors and encoders, you need to install additional dependencies.

Dependencies for predictors and encoders

To use miRBench with predictors and encoders, install the following dependencies:

  • numpy
  • biopython
  • viennarna
  • torch
  • tensorflow
  • typing-extensions

To install the miRBench package with all dependencies into a virtual environment, you can use the following commands:

python3.8 -m venv mirbench_venv
source mirbench_venv/bin/activate
pip install miRBench
pip install numpy==1.24.3 biopython==1.83 viennarna==2.7.0 torch==1.9.0 tensorflow==2.13.1 typing-extensions==4.5.0

Note: This installation is for running predictors on the CPU. If you want to use GPU, you need to install version of torch and tensorflow with GPU support.

Examples

List all available datasets

The dataset module is responsible for access to the benchmark datasets described in the miRBench paper.

from miRBench.dataset import list_datasets

list_datasets()
['AGO2_CLASH_Hejret2023',
 'AGO2_eCLIP_Klimentova2022',
 'AGO2_eCLIP_Manakov2022']

Not all datasets are available with all splits. To get available splits, use the full option.

list_datasets(full=True)
{'AGO2_CLASH_Hejret2023': {'splits': ['train', 'test']},
 'AGO2_eCLIP_Klimentova2022': {'splits': ['test']},
 'AGO2_eCLIP_Manakov2022': {'splits': ['train', 'test', 'leftout']}}

Get dataset

from miRBench.dataset import get_dataset_df

dataset_name = "AGO2_CLASH_Hejret2023"
df = get_dataset_df(dataset_name, split="test")
df.head()
gene noncodingRNA label
0 AGATATGTATTCAGCTTGTCTTCAAATACGGCCAAGCAGAAAATGTTTTA CACTGCATTCCTGCTTGGCCCAG 1
1 ATTCCTTGGGGGATGGTTTGGGCCGAATGGGGAGTGGAATATTTGACATT CACTGCATTCCTGCTTGGCCCAG 1
2 TGAATCAACCCACAGAACCCCCTCCTAAACCCGTTTTCCCACCCACTGCT TTGGAGGCGTGGGTTTT 1
3 GGAGTCTGGAGTCAAACCCAGAGCAGCTGCAGGCCATGAGGCACATTGTT AAAGCAAATGTTGGGTGAACGGC 0
4 CAGCTGTGTACAGCGCCATCTCTCTGCCTTCTGTTGCCCCTCACTCACCA AATAGCTCAGAATGTCAGTTCTG 0

Depending on the dataset version, additional annotation columns may be provided (e.g. genomic coordinates, transcript features, conservation scores, etc). These columns are useful for downstream analyses but are not required for model inference.

If you want to get just a path to the dataset, use the get_dataset_path function:

from miRBench.dataset import get_dataset_path

dataset_path = get_dataset_path(dataset_name, split="test")
dataset_path
/home/user/.miRBench/datasets/20540907/AGO2_CLASH_Hejret2023/test/dataset.tsv

List all available tools

from miRBench.predictor import list_predictors

list_predictors()
['CnnMirTarget_Zheng2020',
 'RNACofold',
 'miRNA_CNN_Hejret2023',
 'miRBind_Klimentova2022',
 'TargetNet_Min2021',
 'Seed8mer',
 'Seed7mer',
 'Seed6mer',
 'Seed6merBulgeOrMismatch',
 'TargetScanCnn_McGeary2019',
 'InteractionAwareModel_Yang2024',
 'miRBenchCNN_Manakov',
 'miRBenchCNN_HejretCorrected']

Encode dataset

The encoder module is responsible for encoding data into the format expected by a predictor module. The main function of the module is get_encoder(predictor_name) which returns an instance of an encoder object implemented for a specified predictor. The encoder expects data as a Pandas DataFrame with columns named noncodingRNA and gene. Specifying custom column names is possible when calling the encoder. The returned data format differs for every encoder and is specific to the predictor.

from miRBench.encoder import get_encoder

tool = 'miRBind_Klimentova2022'
encoder = get_encoder(tool)

input = encoder(df)

Get predictions

The predictor module is responsible for predicting miRNA-binding site interaction. The main function of the module is get_predictor(predictor_name) which downloads the specified predictor to /home/user/.miRBench/models/20612339/<predictor_name>/<predictor_file> and returns an instance of the predictor object. The predictor object expects data encoded by a corresponding encoder and returns an array of predictions.

from miRBench.predictor import get_predictor

predictor = get_predictor(tool)

predictions = predictor(input)
predictions[:10]
array([0.6899161 , 0.15220629, 0.07301956, 0.43757868, 0.34360734,
       0.20519172, 0.0955029 , 0.79298246, 0.14150576, 0.05329492],
      dtype=float32)

Citing miRBench

If you use miRBench in your research, please cite the following article:

Sammut, Stephanie, et al. miRBench: novel benchmark datasets for microRNA binding site prediction that mitigate against prevalent microRNA frequency class bias. Bioinformatics 41.Supplement_1 (2025): i542-i551.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mirbench-1.0.3.tar.gz (17.8 kB view details)

Uploaded Source

File details

Details for the file mirbench-1.0.3.tar.gz.

File metadata

  • Download URL: mirbench-1.0.3.tar.gz
  • Upload date:
  • Size: 17.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.0.0 CPython/3.12.3

File hashes

Hashes for mirbench-1.0.3.tar.gz
Algorithm Hash digest
SHA256 c95cfa9a325356727a10961a7fd4d1edfb3dabb11de909ef9f96165819bba653
MD5 e456ea1c3b4395eb6b354bac3674d432
BLAKE2b-256 f8681f91189958e58d65ba637e6bb6fa5baa04af1dce1d5a440f043202b6cb24

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page