conda create -n quickprot -c conda-forge -c bioconda \
biopython perl perl-uri miniprot td2 pip
conda activate quickprot
pip install rare-quickprot
alternatively using the conda-environment.yml file included in this repo:
conda env create -f conda-environment.yml
QuickProt User Guide
Update
- 2026/04/28
- Major Update: The version number has been updated to 1.9.0
- BUSCO results have improved significantly!!!
- The runtime logic and results have been optimized, resulting in more comprehensive gene predictions.
- Partially overlapping genes (<0.2 overlap) are now permitted.
- Compatibility with newer versions of Python has been improved.
- 2026/04/13
- extract_sequence_from_gff3.py now supports GFF3 file from NCBI. Some gene models in NCBI GFF3 contain in-frame stop codons; when translated into proteins, these are now converted to X (using the -cx option).
- 2025/09/16
- Added add_type_gff3.py script.
- Optimized prediction of stop codon.
- Four new options have been added, namely "-c", "-ps", "-ms", and "-an", for quality control of protein mapping, rational use to reduce pseudogenes.
- 2025/09/08
- Added gtf_genome_to_cdna_fasta.py and gtf_genome_to_cdna_fasta.py script.
- Input files now support .gz compressed files.
- Now QuickProt can output more running details.
- 2025/09/06
- Provides information about the running process.
- Added filter_repeatPeps_from_gff3.py script for removing repeat proteins from gff3 file.
- 2025/05/28
- Provide -ORFSoftware TD2 option, you can use TD2 as a tool for ORF prediction.
- Optimization of genetic code options.
What is QuickProt?
The QuickProt algorithm is a homology-based method for predicting gene models across entire genomes, designed to rapidly construct a non-redundant set of gene models. As illustrated in Figure 1, its core principle is analogous to the blotting method. It primarily employs miniprot (v0.18), to align homologous protein sequences to the genome, delineates high-alignment regions to assemble pseudo-transcripts (lacking UTR regions), and predicts coding regions within these pseudo-transcripts using TransDecoder (v5.7.1). Subsequently, low-quality gene models are filtered out and chimeric gene models are dissected, ultimately generating a high-accuracy, non-redundant gene set.
Installation:
Before use, you need to install Perl, Python, and biopython.
Python3 >= 3.8, perl >= 5
For ease of use, miniprot (v0.18) and TransDecoder (v5.7.1) software are integrated into QuickProt.
wget https://github.com/thecgs/quickprot/archive/refs/tags/quickprot-v1.8.0.tar.gz
tar -zxvf quickprot-v1.8.0.tar.gz
cd quickprot-v1.8.0
./quickprot -h
Note:
# if you need to use --mask optional of qucikprot.py script, and you need to install biopython
pip install biopython
# if you need to use sort_gff3.py script, and you need to install natsort.
pip install natsort
# if you need to use -ORFSoftware TD2, and you need to install TD2
pip install TD2
Usage:
To quickly run QuickProt software. like this,
./quickprot -q protein.fasta -g genome.fasta
This pipeline can improve busco missing result, but you need to download compleasm software.
## step1. running quickprot software
./quickprot.py -q protein.fasta -g genome.fasta -p quickprot.raw
## step2. running compleasm software
compleasm.py run -a genome.fasta -o ./ -l your_lineage
## step3. to update raw gff3 of step1 from compleasm result
./script/update_gff3_from_minibusco.py -r quickprot.raw.longest.gff3 -m ./your_lineage/miniprot_output.gff -g genome.fasta -o improve_busco.gff3
## step4. merge step1 and step3 gff3 result
cat quickprot.raw.longest.gff3 improve_busco.gff3 > genome.longest.gff.tmp
## step5. to sort by chromosomes or scaffold and gene start position and to rename gff3
./script/sort_gff3.py genome.longest.gff.tmp | ./script/rename_gff3.py - -o genome.longest.gff3 -p QUICKPROT; rm genome.longest.gff.tmp
## step6. extract protein sequence from genome and gff file
./bin/TransDecoder-5.7.1/util/gff3_file_to_proteins.pl --gff3 genome.longest.gff3 --fasta genome.fasta --seqType prot > genome.longest.pep.fasta
## step7. extract CDS sequence from genome and gff file
./bin/TransDecoder-5.7.1/util/gff3_file_to_proteins.pl --gff3 genome.longest.gff3 --fasta genome.fasta --seqType CDS > genome.longest.cds.fasta
This step can help you remove repeat proteins (e.g. ENV, Gag, Pol, RT, RH, INT, etc.), but you need to download diamond.
./script/filter_repeatPeps_from_gff3.py -q genome.longest.pep.fasta -g genome.longest.gff3
## results
## retain.gff3 —— Gene model without repeat proteins
## discard.gff3 —— Gene model of repeat proteins
Run with Singularity
Download the Singularity image here
singularity exec -B PATH -e quickprot.v1.9.0.sif quickprot.py -q protein.fasta -g genome.fasta
Test run
If you clone this repository, you can perform a test run of QuickProt using pytest. This may be useful if you wish to contribute to development.
First, it is convenient to set up a Conda environment for QuickProt testing:
git clone https://github.com/thecgs/quickprot.git
conda env create -f quickprot/test/environment.yml
conda activate quickprot-pytest
To perform the test run, simply clone and navigate to the repository, and run pytest.
cd quickprot
pytest -s
The included test datasets in test/data are FASTA files containing:
- A collection of approximately 8,000 Saccharomyces protein sequences downloaded from UniProt (query proteins, Swiss-prot reviewed with tax ID 4930)
- Chromosome 1 of Saccharomyces cerevisiae from Ensembl (genome, softmasked).
When the test is executed, pytest will carry out annotation of the toy genome sequence with the toy protein query set and check for equality of the output files with a set of reference outputs found in test/data/output using MD5 checksums. If all test outputs are equal to the references, the test run will pass, if not, it will fail.
If you include the option --TD2, pytest will carry out two test runs, one using TransDecoder and a second one using TD2.
pytest -s --TD2
Note: TD2's outputs are not as deterministic, so only the miniprot output and transcript assembly will be checked
Cite QuickProt:
If you use QuickProt, please cite:
Guisen Chen, Hehe Du, Zhenjie Cao, Ying Wu, Chen Zhang, Yongcan Zhou, Jingqun Ao, Yun Sun, Zihao Yuan. 2026. “ QuickProt: A Fast and Accurate Homology-Based Protein Annotation Tool for Non-Model Organisms to Advance Comparative Genomics.” Molecular Ecology Resources 26, no. 2: e70097. https://doi.org/10.1111/1755-0998.70097.
Release files for rare-quickprot 1.10.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| rare_quickprot-1.10.1.tar.gz | 73.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| rare_quickprot-1.10.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 236.2 kB
Release files / rare_quickprot-1.10.1.tar.gz
| Download URL | rare_quickprot-1.10.1.tar.gz |
|---|---|
| Size | 73.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
bf342d7970258a3aa046accb296a8ec21fd7c2acfe1e72f93674f1db9eba1d98
|
|
BLAKE2b-256 checksum How to use checksums |
3475109900e4bf32e742da08579e09b34cb9ebd194f96a6247a11b82484d4341
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.16
|
Release files / rare_quickprot-1.10.1-py3-none-any.whl
| Download URL | rare_quickprot-1.10.1-py3-none-any.whl |
|---|---|
| Size | 163.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0d8cfaa9bcc112e61c1000fecaaf66d412ae2f3d5f576a8a985432c56b461b57
|
|
BLAKE2b-256 checksum How to use checksums |
d8372d871b7e1cacc76cfcc9740d5dc3a3d135b09bf632305680f4c908267f75
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.16
|