Skip to main content

A package for benchmarking biological data spliting methods

Project description

SAEA Benchmark

A Python package to benchmark sequence splitting algorithms using:

  • BLAST-based sequence identity calculation
  • ML model performance evaluation

Installation

This package can be installed via pip.

pip install saea-benchmark

Usage

Instantiate an experiment object.

from saea_benchmark.experiment import BenchmarkExperiment

# Initialize experiment
benchmark = BenchmarkExperiment(
   full_fasta_file_path='*.fasta', # Path to the fasta file containing all sequences
   split_file_path='*.csv',        # Path to the csv file for the split
   suite=['blast', 'model']
)

Prepare arguments for configuring an experiment.

args_blast = {
   "out_file": "blast_output.csv", # Path to save the identities, None for not saving them
   "max_identity_threshold": 0.8,  # Threshold for the "% below threshold" metric
   "n_procs": 8                    # Number of threads for sequence alignment
}

args_model = {
   "train_val_adata_path": '*.h5ad', # Path to the .h5ad file for train/validation
   "test_adata_path": '*.h5ad',      # Path to the .h5ad file for test
   "metric_name": 'gorodkin',        # Metric for hyperparam tuning, gorodkin, mcc or f1
   "n_trials": 10,                   # Number of tuning steps
   "random_state": 42,               # Random state
   "allow_logging": False            # Whether optuna logs are displayed
}

Run experiments.

# Run separately
blast_results = benchmark.run_blast_benchmark(**args_blast)
model_results = benchmark.run_model_benchmark(**args_model)

# Run together sequentially
merged_results = benchmark.run_suite(
    {
        'blast': args_blast,
        'model': args_model
    },
    concurrent=True # Run two suites concurrently
)

Save the results of an experiment.

benchmark.save_results('*.json') # Save results to a json file

Generate a plot to visualize the experiment.

import matplotlib.pyplot as plt
from saea_benchmark.create_plot import visualize_fold_analysis

fig = visualize_fold_analysis(
   full_adata_path='*.h5ad',      # .h5ad file storing embedding of all sequences
   split_df_path='*.csv',         # Path to the csv file for the split
   dataset_name='Cyc',            # Name of the dataset displayed on the title
   experiment_json_path='*.json', # File saving the results of benchmarking, default to None (no benchmarking results displayed)
   figsize=(14, 6),               # Figure size for matplotlib, default to (14, 6)
   width_ratios=(0.5, 0.5),       # Ratio of left/right halves of the image, default to (0.5, 0.5)
   scatter_alpha=0.5,             # Transparency for the scatter plot, default to 0.5
   table_col_widths=(0.15, 0.25, 0.3, 0.3, 0.3, 0.3), # Column widths for the table, default to (0.15, 0.25, 0.3, 0.3, 0.3, 0.3)
   table_font_size: int = 12,     # Font size for text in the table, default to 12
   table_scale=(0.65, 6),         # Scale for width/height of the table, default to (0.65, 6)
   bar_alpha=0.7,                 # Transparency for the bar plot, default to 0.7
   bar_width=0.25,                # Width of the bar, default to 0.25
   dpi=300,                       # Resolution of displayed/saved image, default to 300
   save_path=None,                # Path for saving the image, default to None (not saving the image)
)
plt.show()

Other methods can be viewed at this notebook.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

saea_benchmark-0.1.7.tar.gz (17.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

saea_benchmark-0.1.7-py3-none-any.whl (18.3 kB view details)

Uploaded Python 3

File details

Details for the file saea_benchmark-0.1.7.tar.gz.

File metadata

  • Download URL: saea_benchmark-0.1.7.tar.gz
  • Upload date:
  • Size: 17.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.10.16

File hashes

Hashes for saea_benchmark-0.1.7.tar.gz
Algorithm Hash digest
SHA256 05d4ea1642b494b0b19830b42556d8c37e1a6b7ff06cbd4dffdf58f324fc8418
MD5 89760360d08e74cc512911fb514ab56c
BLAKE2b-256 bd5756ca93b2ab43e1561e90badf8bfabcccb64384a6b48701033a512191b0f6

See more details on using hashes here.

File details

Details for the file saea_benchmark-0.1.7-py3-none-any.whl.

File metadata

  • Download URL: saea_benchmark-0.1.7-py3-none-any.whl
  • Upload date:
  • Size: 18.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.10.16

File hashes

Hashes for saea_benchmark-0.1.7-py3-none-any.whl
Algorithm Hash digest
SHA256 e8fff27214403b1d952ed4b1e526cad12b609b31cea91f228aeb67056bda0c61
MD5 ce0c44cec5fba44c7397125017261875
BLAKE2b-256 40ef6239fb10098b8bbb04bbd3a299094f7a5f41f7cc4f7ea7932de90633183e

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page