Skip to main content
conda create -n quickprot -c conda-forge -c bioconda \
  biopython perl perl-uri miniprot td2 pip
conda activate quickprot
pip install rare-quickprot

alternatively using the conda-environment.yml file included in this repo:

conda env create -f conda-environment.yml

QuickProt User Guide

Update

  • 2026/04/28
    1. Major Update: The version number has been updated to 1.9.0
    2. BUSCO results have improved significantly!!!
    3. The runtime logic and results have been optimized, resulting in more comprehensive gene predictions.
    4. Partially overlapping genes (<0.2 overlap) are now permitted.
    5. Compatibility with newer versions of Python has been improved.
  • 2026/04/13
    1. extract_sequence_from_gff3.py now supports GFF3 file from NCBI. Some gene models in NCBI GFF3 contain in-frame stop codons; when translated into proteins, these are now converted to X (using the -cx option).
  • 2025/09/16
    1. Added add_type_gff3.py script.
    2. Optimized prediction of stop codon.
    3. Four new options have been added, namely "-c", "-ps", "-ms", and "-an", for quality control of protein mapping, rational use to reduce pseudogenes.
  • 2025/09/08
    1. Added gtf_genome_to_cdna_fasta.py and gtf_genome_to_cdna_fasta.py script.
    2. Input files now support .gz compressed files.
    3. Now QuickProt can output more running details.
  • 2025/09/06
    1. Provides information about the running process.
    2. Added filter_repeatPeps_from_gff3.py script for removing repeat proteins from gff3 file.
  • 2025/05/28
    1. Provide -ORFSoftware TD2 option, you can use TD2 as a tool for ORF prediction.
    2. Optimization of genetic code options.

What is QuickProt?

The QuickProt algorithm is a homology-based method for predicting gene models across entire genomes, designed to rapidly construct a non-redundant set of gene models. As illustrated in Figure 1, its core principle is analogous to the blotting method. It primarily employs miniprot (v0.18), to align homologous protein sequences to the genome, delineates high-alignment regions to assemble pseudo-transcripts (lacking UTR regions), and predicts coding regions within these pseudo-transcripts using TransDecoder (v5.7.1). Subsequently, low-quality gene models are filtered out and chimeric gene models are dissected, ultimately generating a high-accuracy, non-redundant gene set.

Schema of quickprot algorithm

Fig1. Schema of QuickProt algorithm

Installation:

Before use, you need to install Perl, Python, and biopython.

Python3 >= 3.8, perl >= 5

For ease of use, miniprot (v0.18) and TransDecoder (v5.7.1) software are integrated into QuickProt.

wget https://github.com/thecgs/quickprot/archive/refs/tags/quickprot-v1.8.0.tar.gz
tar -zxvf quickprot-v1.8.0.tar.gz
cd quickprot-v1.8.0
./quickprot -h

Note:

# if you need to use --mask optional of qucikprot.py script, and you need to install biopython
pip install biopython

# if you need to use sort_gff3.py script, and you need to install natsort.
pip install natsort

# if you need to use -ORFSoftware TD2, and you need to install TD2
pip install TD2

Usage:

To quickly run QuickProt software. like this,

./quickprot -q protein.fasta -g genome.fasta

This pipeline can improve busco missing result, but you need to download compleasm software.

## step1. running quickprot software
./quickprot.py -q protein.fasta -g genome.fasta -p quickprot.raw

## step2. running compleasm software
compleasm.py run -a genome.fasta -o ./ -l your_lineage

## step3. to update raw gff3 of step1 from compleasm result
./script/update_gff3_from_minibusco.py -r quickprot.raw.longest.gff3 -m ./your_lineage/miniprot_output.gff -g genome.fasta -o improve_busco.gff3

## step4. merge step1 and step3 gff3 result
cat quickprot.raw.longest.gff3 improve_busco.gff3 > genome.longest.gff.tmp

## step5. to sort by chromosomes or scaffold and gene start position and to rename gff3
./script/sort_gff3.py genome.longest.gff.tmp | ./script/rename_gff3.py - -o genome.longest.gff3 -p QUICKPROT; rm genome.longest.gff.tmp

## step6. extract protein sequence from genome and gff file
./bin/TransDecoder-5.7.1/util/gff3_file_to_proteins.pl --gff3 genome.longest.gff3 --fasta genome.fasta --seqType prot  > genome.longest.pep.fasta

## step7. extract CDS sequence from genome and gff file
./bin/TransDecoder-5.7.1/util/gff3_file_to_proteins.pl --gff3 genome.longest.gff3 --fasta genome.fasta --seqType CDS  > genome.longest.cds.fasta

This step can help you remove repeat proteins (e.g. ENV, Gag, Pol, RT, RH, INT, etc.), but you need to download diamond.

./script/filter_repeatPeps_from_gff3.py -q genome.longest.pep.fasta -g genome.longest.gff3

## results
## retain.gff3 —— Gene model without repeat proteins
## discard.gff3 —— Gene model of repeat proteins

Run with Singularity

Download the Singularity image here

singularity exec -B PATH -e quickprot.v1.9.0.sif quickprot.py -q protein.fasta -g genome.fasta

Cite QuickProt:

If you use QuickProt, please cite:

Guisen Chen, Hehe Du, Zhenjie Cao, Ying Wu, Chen Zhang, Yongcan Zhou, Jingqun Ao, Yun Sun, Zihao Yuan. 2026. “ QuickProt: A Fast and Accurate Homology-Based Protein Annotation Tool for Non-Model Organisms to Advance Comparative Genomics.” Molecular Ecology Resources 26, no. 2: e70097. https://doi.org/10.1111/1755-0998.70097.

Release files for rare-quickprot 1.10.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for rare-quickprot 1.10.0
File Size Uploaded
rare_quickprot-1.10.0.tar.gz 71.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for rare-quickprot 1.10.0
File Interpreter ABI Platform
rare_quickprot-1.10.0-py3-none-any.whl Python 3 none any Details

Total release size: 234.1 kB

Release files / rare_quickprot-1.10.0.tar.gz

Download URL rare_quickprot-1.10.0.tar.gz
Size 71.5 kB
Tags Source
SHA-256 checksum
How to use checksums
02b53579ccc7678a1a9cea0688c85b58fad81c85c4e855ea61600ebc4bfda800
BLAKE2b-256 checksum
How to use checksums
cb1541f177f0d6edac1f46c7ebc05d9a1334302f7766e634f609d7377ff8d09c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.14

Release files / rare_quickprot-1.10.0-py3-none-any.whl

Download URL rare_quickprot-1.10.0-py3-none-any.whl
Size 162.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
bfaeaa978886d0ad9f60ea894c4959d361ce9bec42fe9ca4a9cee75124bdb0de
BLAKE2b-256 checksum
How to use checksums
169cf8c4490acd29a88a1f191949ff9a1ef6ae5bda0712e005dc1e4fc021f986
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.14

Release history Release notifications | RSS feed

1.10.2

2 release files

1.10.1

2 release files

This release

1.10.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page