Skip to main content
conda create -n quickprot -c conda-forge -c bioconda \
  biopython perl perl-uri miniprot td2 pip
conda activate quickprot
pip install rare-quickprot

alternatively using the conda-environment.yml file included in this repo:

conda env create -f conda-environment.yml

QuickProt User Guide

Update

  • 2026/04/28
    1. Major Update: The version number has been updated to 1.9.0
    2. BUSCO results have improved significantly!!!
    3. The runtime logic and results have been optimized, resulting in more comprehensive gene predictions.
    4. Partially overlapping genes (<0.2 overlap) are now permitted.
    5. Compatibility with newer versions of Python has been improved.
  • 2026/04/13
    1. extract_sequence_from_gff3.py now supports GFF3 file from NCBI. Some gene models in NCBI GFF3 contain in-frame stop codons; when translated into proteins, these are now converted to X (using the -cx option).
  • 2025/09/16
    1. Added add_type_gff3.py script.
    2. Optimized prediction of stop codon.
    3. Four new options have been added, namely "-c", "-ps", "-ms", and "-an", for quality control of protein mapping, rational use to reduce pseudogenes.
  • 2025/09/08
    1. Added gtf_genome_to_cdna_fasta.py and gtf_genome_to_cdna_fasta.py script.
    2. Input files now support .gz compressed files.
    3. Now QuickProt can output more running details.
  • 2025/09/06
    1. Provides information about the running process.
    2. Added filter_repeatPeps_from_gff3.py script for removing repeat proteins from gff3 file.
  • 2025/05/28
    1. Provide -ORFSoftware TD2 option, you can use TD2 as a tool for ORF prediction.
    2. Optimization of genetic code options.

What is QuickProt?

The QuickProt algorithm is a homology-based method for predicting gene models across entire genomes, designed to rapidly construct a non-redundant set of gene models. As illustrated in Figure 1, its core principle is analogous to the blotting method. It primarily employs miniprot (v0.18), to align homologous protein sequences to the genome, delineates high-alignment regions to assemble pseudo-transcripts (lacking UTR regions), and predicts coding regions within these pseudo-transcripts using TransDecoder (v5.7.1). Subsequently, low-quality gene models are filtered out and chimeric gene models are dissected, ultimately generating a high-accuracy, non-redundant gene set.

Schema of quickprot algorithm

Fig1. Schema of QuickProt algorithm

Installation:

Before use, you need to install Perl, Python, and biopython.

Python3 >= 3.8, perl >= 5

For ease of use, miniprot (v0.18) and TransDecoder (v5.7.1) software are integrated into QuickProt.

wget https://github.com/thecgs/quickprot/archive/refs/tags/quickprot-v1.8.0.tar.gz
tar -zxvf quickprot-v1.8.0.tar.gz
cd quickprot-v1.8.0
./quickprot -h

Note:

# if you need to use --mask optional of qucikprot.py script, and you need to install biopython
pip install biopython

# if you need to use sort_gff3.py script, and you need to install natsort.
pip install natsort

# if you need to use -ORFSoftware TD2, and you need to install TD2
pip install TD2

Usage:

To quickly run QuickProt software. like this,

./quickprot -q protein.fasta -g genome.fasta

This pipeline can improve busco missing result, but you need to download compleasm software.

## step1. running quickprot software
./quickprot.py -q protein.fasta -g genome.fasta -p quickprot.raw

## step2. running compleasm software
compleasm.py run -a genome.fasta -o ./ -l your_lineage

## step3. to update raw gff3 of step1 from compleasm result
./script/update_gff3_from_minibusco.py -r quickprot.raw.longest.gff3 -m ./your_lineage/miniprot_output.gff -g genome.fasta -o improve_busco.gff3

## step4. merge step1 and step3 gff3 result
cat quickprot.raw.longest.gff3 improve_busco.gff3 > genome.longest.gff.tmp

## step5. to sort by chromosomes or scaffold and gene start position and to rename gff3
./script/sort_gff3.py genome.longest.gff.tmp | ./script/rename_gff3.py - -o genome.longest.gff3 -p QUICKPROT; rm genome.longest.gff.tmp

## step6. extract protein sequence from genome and gff file
./bin/TransDecoder-5.7.1/util/gff3_file_to_proteins.pl --gff3 genome.longest.gff3 --fasta genome.fasta --seqType prot  > genome.longest.pep.fasta

## step7. extract CDS sequence from genome and gff file
./bin/TransDecoder-5.7.1/util/gff3_file_to_proteins.pl --gff3 genome.longest.gff3 --fasta genome.fasta --seqType CDS  > genome.longest.cds.fasta

This step can help you remove repeat proteins (e.g. ENV, Gag, Pol, RT, RH, INT, etc.), but you need to download diamond.

./script/filter_repeatPeps_from_gff3.py -q genome.longest.pep.fasta -g genome.longest.gff3

## results
## retain.gff3 —— Gene model without repeat proteins
## discard.gff3 —— Gene model of repeat proteins

Run with Singularity

Download the Singularity image here

singularity exec -B PATH -e quickprot.v1.9.0.sif quickprot.py -q protein.fasta -g genome.fasta

Test run

If you clone this repository, you can perform a test run of QuickProt using pytest. This may be useful if you wish to contribute to development.

First, it is convenient to set up a Conda environment for QuickProt testing:

git clone https://github.com/thecgs/quickprot.git
conda env create -f quickprot/test/environment.yml
conda activate quickprot-pytest

To perform the test run, simply clone and navigate to the repository, and run pytest.

cd quickprot
pytest -s

The included test datasets in test/data are FASTA files containing:

  • A collection of approximately 8,000 Saccharomyces protein sequences downloaded from UniProt (query proteins, Swiss-prot reviewed with tax ID 4930)
  • Chromosome 1 of Saccharomyces cerevisiae from Ensembl (genome, softmasked).

When the test is executed, pytest will carry out annotation of the toy genome sequence with the toy protein query set and check for equality of the output files with a set of reference outputs found in test/data/output using MD5 checksums. If all test outputs are equal to the references, the test run will pass, if not, it will fail.

If you include the option --TD2, pytest will carry out two test runs, one using TransDecoder and a second one using TD2.

pytest -s --TD2

Note: TD2's outputs are not as deterministic, so only the miniprot output and transcript assembly will be checked

Cite QuickProt:

If you use QuickProt, please cite:

Guisen Chen, Hehe Du, Zhenjie Cao, Ying Wu, Chen Zhang, Yongcan Zhou, Jingqun Ao, Yun Sun, Zihao Yuan. 2026. “ QuickProt: A Fast and Accurate Homology-Based Protein Annotation Tool for Non-Model Organisms to Advance Comparative Genomics.” Molecular Ecology Resources 26, no. 2: e70097. https://doi.org/10.1111/1755-0998.70097.

Release files for rare-quickprot 1.10.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for rare-quickprot 1.10.1
File Size Uploaded
rare_quickprot-1.10.1.tar.gz 73.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for rare-quickprot 1.10.1
File Interpreter ABI Platform
rare_quickprot-1.10.1-py3-none-any.whl Python 3 none any Details

Total release size: 236.2 kB

Release files / rare_quickprot-1.10.1.tar.gz

Download URL rare_quickprot-1.10.1.tar.gz
Size 73.0 kB
Tags Source
SHA-256 checksum
How to use checksums
bf342d7970258a3aa046accb296a8ec21fd7c2acfe1e72f93674f1db9eba1d98
BLAKE2b-256 checksum
How to use checksums
3475109900e4bf32e742da08579e09b34cb9ebd194f96a6247a11b82484d4341
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release files / rare_quickprot-1.10.1-py3-none-any.whl

Download URL rare_quickprot-1.10.1-py3-none-any.whl
Size 163.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0d8cfaa9bcc112e61c1000fecaaf66d412ae2f3d5f576a8a985432c56b461b57
BLAKE2b-256 checksum
How to use checksums
d8372d871b7e1cacc76cfcc9740d5dc3a3d135b09bf632305680f4c908267f75
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release history Release notifications | RSS feed

1.10.2

2 release files

This release

1.10.1 This release

2 release files

1.10.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page