protclust
A Python library for working with protein sequence data, providing:
- Clustering capabilities via MMseqs2
- Machine learning dataset creation with cluster-aware splits
Requirements
This library requires MMseqs2, which must be installed and accessible via the command line. MMseqs2 can be installed using one of the following methods:
Installation Options for MMseqs2
-
Homebrew:
brew install mmseqs2
-
Conda:
conda install -c conda-forge -c bioconda mmseqs2
-
Docker:
docker pull ghcr.io/soedinglab/mmseqs2
-
Static Build (AVX2, SSE4.1, or SSE2):
wget https://mmseqs.com/latest/mmseqs-linux-avx2.tar.gz tar xvfz mmseqs-linux-avx2.tar.gz export PATH=$(pwd)/mmseqs/bin/:$PATH
MMseqs2 must be accessible via the mmseqs command in your system's PATH. If the library cannot detect MMseqs2, it will raise an error.
Installation
Installation
You can install protclust using pip:
pip install protclust
Or if installing from source, clone the repository and run:
pip install -e .
For development purposes, also install the testing dependencies:
pip install pytest pytest-cov pre-commit ruff
Features
Sequence Clustering and Dataset Creation
import pandas as pd
from protclust import clean, cluster, split, set_verbosity
# Enable detailed logging (optional)
set_verbosity(verbose=True)
# Example data
df = pd.DataFrame({
"id": ["seq1", "seq2", "seq3", "seq4"],
"sequence": ["ACDEFGHIKL", "ACDEFGHIKL", "MNPQRSTVWY", "MNPQRSTVWY"]
})
# Clean data
clean_df = clean(df, sequence_col="sequence")
# Cluster sequences
clustered_df = cluster(clean_df, sequence_col="sequence", id_col="id")
# Split data into train and test sets
train_df, test_df = split(clustered_df, group_col="cluster_representative", test_size=0.3)
print("Train set:\n", train_df)
print("Test set:\n", test_df)
# MILP-based splitting with property balancing
from protclust import milp_split
train_df, test_df = milp_split(
clustered_df,
group_col="cluster_representative",
test_size=0.3,
balance_cols=["molecular_weight", "hydrophobicity"]
)
Parameters
Common parameters for clustering functions:
df: Pandas DataFrame containing sequence datasequence_col: Column name containing sequencesid_col: Column name containing unique identifiersmin_seq_id: Minimum sequence identity threshold (0.0-1.0, default 0.3)coverage: Minimum alignment coverage (0.0-1.0, default 0.5)cov_mode: Coverage mode (0-3, default 0)cluster_mode: Clustering algorithm (0: Set-Cover, 1: Connected component, 2: Greedy by length, default 0)cluster_steps: Number of cascaded clustering steps for large datasets (default 1)test_size: Desired fraction of data in test set (default 0.2)random_state: Random seed for reproducibilitytolerance: Acceptable deviation from desired split sizes (default 0.05)
Contributing
Contributions are welcome! Please feel free to submit a Pull Request.
- Fork the repository
- Create your feature branch (
git checkout -b feature/amazing-feature) - Run tests (
pytest tests/) - Commit your changes (
git commit -m 'Add some amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
License
This project is licensed under the MIT License - see the LICENSE file for details.
Citation
If you use protclust in your research, please cite:
@software{protclust,
author = {Michael Scutari},
title = {protclust: Protein Sequence Clustering and ML Dataset Creation},
url = {https://github.com/michaelscutari/protclust},
version = {0.2.0},
year = {2025},
}
Metadata
Release files for protclust 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| protclust-0.2.0.tar.gz | 37.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| protclust-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 51.3 kB
Release files / protclust-0.2.0.tar.gz
| Download URL | protclust-0.2.0.tar.gz |
|---|---|
| Size | 37.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
230e305b7336015a9db9c795d4602334571bc5cff6c92504e70e72f383a13379
|
|
BLAKE2b-256 checksum How to use checksums |
14f4244e75efcc50983b9ad27c86d5f03a9d7ff72344dc8c7e741dd3ffa41f97
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.1
|
Release files / protclust-0.2.0-py3-none-any.whl
| Download URL | protclust-0.2.0-py3-none-any.whl |
|---|---|
| Size | 14.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5d0a91ccad77a5dfaf2258c40448abf4b8dc0cf14946b616bb8356de89fad36f
|
|
BLAKE2b-256 checksum How to use checksums |
b80d2e89ec6371d9251f1f2a03dfff633abbdf2301d6a3b3939328a46d2e2517
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.1
|