Random Forest Enzyme Prediction
Project description
RAEP: Rapid Enzyme/Non-Enzyme Prediction
RAEP (Rapid Enzyme/Non-Enzyme Prediction) is an efficient enzyme/non-enzyme prediction tool for protein sequences, built on multi-physicochemical property features and the XGBoost machine learning algorithm.
Features
- Efficient Prediction : Achieves fast and accurate enzyme/non-enzyme classification using optimized feature extraction and the XGBoost model.
- Multi-Mode Support : Supports single-sequence prediction, multi-sequence batch prediction, and FASTA file batch prediction.
- Rich Feature Set : Utilizes multi-physicochemical property pseudo-amino acid composition (Pseudo-AAC), CTD features, and windowed amino acid composition.
- User-Friendly : Offers a concise Python API that is easy to integrate into existing projects.
- Multi-Process Optimization : Employs multi-process parallel processing in the feature extraction step to improve processing efficiency for large-scale datasets.
Installation Methods
Install from PyPI (Recommended)
pip install raep
Install from Source Code
# 克隆仓库(如果有)
git clone <repository-url>
cd RAEP
# 开发模式安装
pip install -e .
Dependencies
- joblib: Used for model saving/loading and parallel processing.
- numpy: Numerical computation.
- pandas: Data processing.
- scikit-learn: Machine learning utilities and evaluation metrics.
- xgboost: Implementation of the gradient boosting tree algorithm.
Basic Usage
Import and Initialization
from raep_package import RAEP
# 默认初始化(使用内置模型)
predictor = RAEP()
# 使用自定义模型路径初始化
# predictor = RAEP(model_path="path/to/your/model.pkl")
Single Sequence Prediction
# 预测单个蛋白质序列
sequence = "MVLSPADKTNVKAAWGKVGAHAGEYGAEALERMFLSFPTTKTYFPHFDLSHGSAQVKGHGKKVADALTNAVAHVDDMPNALSALSDLHAHKLRVDPVNFKLLSHCLLVTLAAHLPAEFTPAVHASLDKFLASVSTVLTSKYR"
prediction, probability = predictor.predict_sequence(sequence)
print(f"预测结果: {'酶' if prediction == 1 else '非酶'}")
print(f"预测概率: 非酶={probability[0]:.4f}, 酶={probability[1]:.4f}")
Batch Prediction for Multiple Sequences
# 批量预测多个序列
sequences = [
"MASMTGGQQMGRGSEF",
"MKVLVVLVLLAVLVLA",
"MDEKTTGWRGGHVA"
]
results = predictor.predict_sequences(sequences)
for i, (pred, prob) in enumerate(results, 1):
print(f"序列 {i}: {'酶' if pred == 1 else '非酶'} (酶概率: {prob[1]:.4f})")
Batch Prediction for FASTA Files
# 从FASTA文件批量预测
fasta_path = "test_sequences.fasta"
results = predictor.predict_fasta(fasta_path)
print(f"预测结果 ({len(results)} 个序列):")
for i, (pred, prob) in enumerate(results, 1):
print(f"序列 {i}: {'酶' if pred == 1 else '非酶'} (酶概率: {prob[1]:.4f})")
API Reference
RAEP Class
Initialization
RAEP(model_path=None)
Purpose : Instantiates the RAEP predictor, automatically loads the model and initializes feature extraction parameters (e.g., LAG=10, W=0.05), ensuring consistency in subsequent prediction workflows.
- Parameters :
model_path: Optional, path to a custom model file. If not provided, the built-inenzyme_xgb_model.pklmodel will be used.
Methods
predict_sequence
predict_sequence(sequence)
Purpose : Predicts whether a single protein sequence is an enzyme.
-
Parameters:
sequence: String, the protein sequence to be predicted.
-
Return Value:
- Tuple
(prediction, probability):prediction: Integer, the predicted class (0 = Non-enzyme, 1 = Enzyme).probability:List ,containing two floats, representing the probabilities of the sequence being a non-enzyme and an enzyme, respectively.
- Tuple
predict_sequences
predict_sequences(sequences)
Purpose : Performs batch prediction for multiple protein sequences.
-
Parameters :
sequences: List containing multiple protein sequence strings.
-
Return Value:
- List where each element is a tuple
(prediction, probability)corresponding to the prediction result of the input sequence.
- List where each element is a tuple
predict_fasta
predict_fasta(fasta_path)
Purpose : Performs batch prediction for protein sequences from a FASTA file.
-
Parameters:
fasta_path: String, path to the FASTA file.
-
Return Value:
- List where each element is a tuple
(prediction, probability)corresponding to the prediction result of the sequences in the FASTA file.
- List where each element is a tuple
Usage Examples
Refer to the example_usage.py file in the project, which contains complete usage examples:
python example_usage.py
Notes
- Sequence Format Requirements: Input sequences should only contain single-letter codes (uppercase) for the 20 standard amino acids.
- Sequence Length: The tool automatically processes sequences of different lengths, but excessively short sequences (e.g., < 10 amino acids) may affect prediction accuracy.
- Multi-Process Processing: The feature extraction process uses multi-processing acceleration by default, which automatically adjusts based on the number of CPU cores in the system.
- Model File: Ensure the model file exists and is accessible, especially when using a custom model path.
Troubleshooting
Common Issues
- Failed to import the RAEP package: Ensure the package is correctly installed in the current Python environment.
- Model loading failure: Verify that the model file path is correct and the file exists at the specified location.
- Prediction errors: Check if the input sequence format is valid and contains only standard amino acid characters.
- Performance issues: For extremely large datasets, consider processing in batches to avoid memory overflow.
Getting Help
If you encounter any problems, please contact the author:
- Author: DHY
- Email: dhy.scut@outlook.com
License
This project is licensed under the MIT License. See the LICENSE file for details.
Acknowledgments
The development of this project is supported by several open-source tools, especially machine learning libraries such as XGBoost and scikit-learn.
Version: 0.0.4 Last updated: 2025
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file raep-0.0.4.tar.gz.
File metadata
- Download URL: raep-0.0.4.tar.gz
- Upload date:
- Size: 6.6 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.10.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ad02ca5125aa3029fed4c76618ba27023f5ec82be0aef860a2efe30fa5aeeea3
|
|
| MD5 |
e3f82f27dc8197035e740bb98f22090e
|
|
| BLAKE2b-256 |
db981f1fbf347b34a8183a1d4fd1720b32cd460a54477bb39e6d6146fbc4a558
|
File details
Details for the file raep-0.0.4-py3-none-any.whl.
File metadata
- Download URL: raep-0.0.4-py3-none-any.whl
- Upload date:
- Size: 6.7 MB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.10.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
acdc0bd963eaeee36554787c907d80ad91c8c25af225c81d67d4888ff6b8e364
|
|
| MD5 |
bc54285dd82c030c688e2e66cbe3df20
|
|
| BLAKE2b-256 |
2d252a79d9a38f817f64878f8115eb4b299bbdbdb46c11ba7cec2e1cb4a2a0c8
|