Skip to main content

A package for benchmarking biological data spliting methods

Project description

SAEA Benchmark

A Python package to benchmark sequence splitting algorithms using:

  • BLAST-based sequence identity calculation
  • ML model performance evaluation

Installation

This package can be installed via pip.

pip install saea-benchmark

Usage

Instantiate an experiment object.

from saea_benchmark.experiment import BenchmarkExperiment

# Initialize experiment
benchmark = BenchmarkExperiment(
   full_fasta_file_path='*.fasta', # Path to the fasta file containing all sequences
   split_file_path='*.csv',        # Path to the csv file for the split
   suite=['blast', 'model']
)

Prepare arguments for configuring an experiment.

args_blast = {
   "out_file": "blast_output.csv", # Path to save the identities, None for not saving them
   "max_identity_threshold": 0.8,  # Threshold for the "% below threshold" metric
   "n_procs": 8                    # Number of threads for sequence alignment
}

args_model = {
   "train_val_adata_path": '*.h5ad', # Path to the .h5ad file for train/validation
   "test_adata_path": '*.h5ad',      # Path to the .h5ad file for test
   "metric_name": 'gorodkin',        # Metric for hyperparam tuning, gorodkin, mcc or f1
   "n_trials": 10,                   # Number of tuning steps
   "random_state": 42,               # Random state
   "allow_logging": False            # Whether optuna logs are displayed
}

Run experiments.

# Run separately
blast_results = benchmark.run_blast_benchmark(**args_blast)
model_results = benchmark.run_model_benchmark(**args_model)

# Run together sequentially
merged_results = benchmark.run_suite(
    {
        'blast': args_blast,
        'model': args_model
    },
    concurrent=True # Run two suites concurrently
)

Save the results of an experiment.

benchmark.save_results('*.json') # Save results to a json file

Generate a plot to visualize the experiment.

import matplotlib.pyplot as plt
from saea_benchmark.create_plot import visualize_fold_analysis

fig = visualize_fold_analysis(
   full_adata_path='*.h5ad',      # .h5ad file storing embedding of all sequences
   split_df_path='*.csv',         # Path to the csv file for the split
   dataset_name='Cyc',            # Name of the dataset displayed on the title
   experiment_json_path='*.json', # File saving the results of benchmarking, default to None (no benchmarking results displayed)
   figsize=(14, 6),               # Figure size for matplotlib, default to (14, 6)
   width_ratios=(0.5, 0.5),       # Ratio of left/right halves of the image, default to (0.5, 0.5)
   scatter_alpha=0.5,             # Transparency for the scatter plot, default to 0.5
   table_col_widths=(0.15, 0.25, 0.3, 0.3, 0.3, 0.3), # Column widths for the table, default to (0.15, 0.25, 0.3, 0.3, 0.3, 0.3)
   table_font_size: int = 12,     # Font size for text in the table, default to 12
   table_scale=(0.65, 6),         # Scale for width/height of the table, default to (0.65, 6)
   bar_alpha=0.7,                 # Transparency for the bar plot, default to 0.7
   bar_width=0.25,                # Width of the bar, default to 0.25
   dpi=300,                       # Resolution of displayed/saved image, default to 300
   save_path=None,                # Path for saving the image, default to None (not saving the image)
)
plt.show()

Other methods can be viewed at this notebook.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

saea_benchmark-0.1.11.tar.gz (17.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

saea_benchmark-0.1.11-py3-none-any.whl (18.6 kB view details)

Uploaded Python 3

File details

Details for the file saea_benchmark-0.1.11.tar.gz.

File metadata

  • Download URL: saea_benchmark-0.1.11.tar.gz
  • Upload date:
  • Size: 17.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.10.16

File hashes

Hashes for saea_benchmark-0.1.11.tar.gz
Algorithm Hash digest
SHA256 5be98378153fb588b616818c678d96145d97689beb6db2486e9be8db05aa8395
MD5 b2a2e6661870405ba6ca41c97356a20a
BLAKE2b-256 d05869301bffe959db30f4b277f874948646c2e1eb0d0c933d53e8f83d1d8b33

See more details on using hashes here.

File details

Details for the file saea_benchmark-0.1.11-py3-none-any.whl.

File metadata

File hashes

Hashes for saea_benchmark-0.1.11-py3-none-any.whl
Algorithm Hash digest
SHA256 32d47e4f2712125e97099042277a524492dbeae7b7daa021307ed00934a00357
MD5 651ea32f2e3b3110ee69987ded1510be
BLAKE2b-256 96b7e123425018ae905f5b2ea283744b4b05a5bef29f2f9cbbbffc2066ed031d

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page