conda create -n quickprot -c conda-forge -c bioconda \
biopython perl perl-uri miniprot td2 pip
conda activate quickprot
pip install rare-quickprot
alternatively using the conda-environment.yml file included in this repo:
conda env create -f conda-environment.yml
QuickProt User Guide
Update
- 2026/04/28
- Major Update: The version number has been updated to 1.9.0
- BUSCO results have improved significantly!!!
- The runtime logic and results have been optimized, resulting in more comprehensive gene predictions.
- Partially overlapping genes (<0.2 overlap) are now permitted.
- Compatibility with newer versions of Python has been improved.
- 2026/04/13
- extract_sequence_from_gff3.py now supports GFF3 file from NCBI. Some gene models in NCBI GFF3 contain in-frame stop codons; when translated into proteins, these are now converted to X (using the -cx option).
- 2025/09/16
- Added add_type_gff3.py script.
- Optimized prediction of stop codon.
- Four new options have been added, namely "-c", "-ps", "-ms", and "-an", for quality control of protein mapping, rational use to reduce pseudogenes.
- 2025/09/08
- Added gtf_genome_to_cdna_fasta.py and gtf_genome_to_cdna_fasta.py script.
- Input files now support .gz compressed files.
- Now QuickProt can output more running details.
- 2025/09/06
- Provides information about the running process.
- Added filter_repeatPeps_from_gff3.py script for removing repeat proteins from gff3 file.
- 2025/05/28
- Provide -ORFSoftware TD2 option, you can use TD2 as a tool for ORF prediction.
- Optimization of genetic code options.
What is QuickProt?
The QuickProt algorithm is a homology-based method for predicting gene models across entire genomes, designed to rapidly construct a non-redundant set of gene models. As illustrated in Figure 1, its core principle is analogous to the blotting method. It primarily employs miniprot (v0.18), to align homologous protein sequences to the genome, delineates high-alignment regions to assemble pseudo-transcripts (lacking UTR regions), and predicts coding regions within these pseudo-transcripts using TransDecoder (v5.7.1). Subsequently, low-quality gene models are filtered out and chimeric gene models are dissected, ultimately generating a high-accuracy, non-redundant gene set.
Installation:
Before use, you need to install Perl, Python, and biopython.
Python3 >= 3.8, perl >= 5
For ease of use, miniprot (v0.18) and TransDecoder (v5.7.1) software are integrated into QuickProt.
wget https://github.com/thecgs/quickprot/archive/refs/tags/quickprot-v1.8.0.tar.gz
tar -zxvf quickprot-v1.8.0.tar.gz
cd quickprot-v1.8.0
./quickprot -h
Note:
# if you need to use --mask optional of qucikprot.py script, and you need to install biopython
pip install biopython
# if you need to use sort_gff3.py script, and you need to install natsort.
pip install natsort
# if you need to use -ORFSoftware TD2, and you need to install TD2
pip install TD2
Usage:
To quickly run QuickProt software. like this,
./quickprot -q protein.fasta -g genome.fasta
This pipeline can improve busco missing result, but you need to download compleasm software.
## step1. running quickprot software
./quickprot.py -q protein.fasta -g genome.fasta -p quickprot.raw
## step2. running compleasm software
compleasm.py run -a genome.fasta -o ./ -l your_lineage
## step3. to update raw gff3 of step1 from compleasm result
./script/update_gff3_from_minibusco.py -r quickprot.raw.longest.gff3 -m ./your_lineage/miniprot_output.gff -g genome.fasta -o improve_busco.gff3
## step4. merge step1 and step3 gff3 result
cat quickprot.raw.longest.gff3 improve_busco.gff3 > genome.longest.gff.tmp
## step5. to sort by chromosomes or scaffold and gene start position and to rename gff3
./script/sort_gff3.py genome.longest.gff.tmp | ./script/rename_gff3.py - -o genome.longest.gff3 -p QUICKPROT; rm genome.longest.gff.tmp
## step6. extract protein sequence from genome and gff file
./bin/TransDecoder-5.7.1/util/gff3_file_to_proteins.pl --gff3 genome.longest.gff3 --fasta genome.fasta --seqType prot > genome.longest.pep.fasta
## step7. extract CDS sequence from genome and gff file
./bin/TransDecoder-5.7.1/util/gff3_file_to_proteins.pl --gff3 genome.longest.gff3 --fasta genome.fasta --seqType CDS > genome.longest.cds.fasta
This step can help you remove repeat proteins (e.g. ENV, Gag, Pol, RT, RH, INT, etc.), but you need to download diamond.
./script/filter_repeatPeps_from_gff3.py -q genome.longest.pep.fasta -g genome.longest.gff3
## results
## retain.gff3 —— Gene model without repeat proteins
## discard.gff3 —— Gene model of repeat proteins
Run with Singularity
Download the Singularity image here
singularity exec -B PATH -e quickprot.v1.9.0.sif quickprot.py -q protein.fasta -g genome.fasta
Cite QuickProt:
If you use QuickProt, please cite:
Guisen Chen, Hehe Du, Zhenjie Cao, Ying Wu, Chen Zhang, Yongcan Zhou, Jingqun Ao, Yun Sun, Zihao Yuan. 2026. “ QuickProt: A Fast and Accurate Homology-Based Protein Annotation Tool for Non-Model Organisms to Advance Comparative Genomics.” Molecular Ecology Resources 26, no. 2: e70097. https://doi.org/10.1111/1755-0998.70097.
Release files for rare-quickprot 1.10.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| rare_quickprot-1.10.0.tar.gz | 71.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| rare_quickprot-1.10.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 234.1 kB
Release files / rare_quickprot-1.10.0.tar.gz
| Download URL | rare_quickprot-1.10.0.tar.gz |
|---|---|
| Size | 71.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
02b53579ccc7678a1a9cea0688c85b58fad81c85c4e855ea61600ebc4bfda800
|
|
BLAKE2b-256 checksum How to use checksums |
cb1541f177f0d6edac1f46c7ebc05d9a1334302f7766e634f609d7377ff8d09c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.14
|
Release files / rare_quickprot-1.10.0-py3-none-any.whl
| Download URL | rare_quickprot-1.10.0-py3-none-any.whl |
|---|---|
| Size | 162.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
bfaeaa978886d0ad9f60ea894c4959d361ce9bec42fe9ca4a9cee75124bdb0de
|
|
BLAKE2b-256 checksum How to use checksums |
169cf8c4490acd29a88a1f191949ff9a1ef6ae5bda0712e005dc1e4fc021f986
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.14
|