Skip to main content

protclust logo

protclust

PyPI version Tests Coverage License: MIT Python Version

A Python library for working with protein sequence data, providing:

  • Clustering capabilities via MMseqs2
  • Machine learning dataset creation with cluster-aware splits

Requirements

This library requires MMseqs2, which must be installed and accessible via the command line. MMseqs2 can be installed using one of the following methods:

Installation Options for MMseqs2

  • Homebrew:

    brew install mmseqs2
    
  • Conda:

    conda install -c conda-forge -c bioconda mmseqs2
    
  • Docker:

    docker pull ghcr.io/soedinglab/mmseqs2
    
  • Static Build (AVX2, SSE4.1, or SSE2):

    wget https://mmseqs.com/latest/mmseqs-linux-avx2.tar.gz
    tar xvfz mmseqs-linux-avx2.tar.gz
    export PATH=$(pwd)/mmseqs/bin/:$PATH
    

MMseqs2 must be accessible via the mmseqs command in your system's PATH. If the library cannot detect MMseqs2, it will raise an error.

Installation

Installation

You can install protclust using pip:

pip install protclust

Or if installing from source, clone the repository and run:

pip install -e .

For development purposes, also install the testing dependencies:

pip install pytest pytest-cov pre-commit ruff

Features

Sequence Clustering and Dataset Creation

import pandas as pd
from protclust import clean, cluster, split, set_verbosity

# Enable detailed logging (optional)
set_verbosity(verbose=True)

# Example data
df = pd.DataFrame({
    "id": ["seq1", "seq2", "seq3", "seq4"],
    "sequence": ["ACDEFGHIKL", "ACDEFGHIKL", "MNPQRSTVWY", "MNPQRSTVWY"]
})

# Clean data
clean_df = clean(df, sequence_col="sequence")

# Cluster sequences
clustered_df = cluster(clean_df, sequence_col="sequence", id_col="id")

# Split data into train and test sets
train_df, test_df = split(clustered_df, group_col="cluster_representative", test_size=0.3)

print("Train set:\n", train_df)
print("Test set:\n", test_df)

# MILP-based splitting with property balancing
from protclust import milp_split
train_df, test_df = milp_split(
    clustered_df,
    group_col="cluster_representative",
    test_size=0.3,
    balance_cols=["molecular_weight", "hydrophobicity"]
)

Parameters

Common parameters for clustering functions:

  • df: Pandas DataFrame containing sequence data
  • sequence_col: Column name containing sequences
  • id_col: Column name containing unique identifiers
  • min_seq_id: Minimum sequence identity threshold (0.0-1.0, default 0.3)
  • coverage: Minimum alignment coverage (0.0-1.0, default 0.5)
  • cov_mode: Coverage mode (0-3, default 0)
  • cluster_mode: Clustering algorithm (0: Set-Cover, 1: Connected component, 2: Greedy by length, default 0)
  • cluster_steps: Number of cascaded clustering steps for large datasets (default 1)
  • test_size: Desired fraction of data in test set (default 0.2)
  • random_state: Random seed for reproducibility
  • tolerance: Acceptable deviation from desired split sizes (default 0.05)

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

  1. Fork the repository
  2. Create your feature branch (git checkout -b feature/amazing-feature)
  3. Run tests (pytest tests/)
  4. Commit your changes (git commit -m 'Add some amazing feature')
  5. Push to the branch (git push origin feature/amazing-feature)
  6. Open a Pull Request

License

This project is licensed under the MIT License - see the LICENSE file for details.

Citation

If you use protclust in your research, please cite:

@software{protclust,
  author = {Michael Scutari},
  title = {protclust: Protein Sequence Clustering and ML Dataset Creation},
  url = {https://github.com/michaelscutari/protclust},
  version = {0.2.0},
  year = {2025},
}

Metadata

Release files for protclust 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for protclust 0.2.0
File Size Uploaded
protclust-0.2.0.tar.gz 37.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for protclust 0.2.0
File Interpreter ABI Platform
protclust-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 51.3 kB

Release files / protclust-0.2.0.tar.gz

Download URL protclust-0.2.0.tar.gz
Size 37.0 kB
Tags Source
SHA-256 checksum
How to use checksums
230e305b7336015a9db9c795d4602334571bc5cff6c92504e70e72f383a13379
BLAKE2b-256 checksum
How to use checksums
14f4244e75efcc50983b9ad27c86d5f03a9d7ff72344dc8c7e741dd3ffa41f97
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.1

Release files / protclust-0.2.0-py3-none-any.whl

Download URL protclust-0.2.0-py3-none-any.whl
Size 14.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5d0a91ccad77a5dfaf2258c40448abf4b8dc0cf14946b616bb8356de89fad36f
BLAKE2b-256 checksum
How to use checksums
b80d2e89ec6371d9251f1f2a03dfff633abbdf2301d6a3b3939328a46d2e2517
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.1

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page