Skip to main content
conda create -n quickprot -c conda-forge -c bioconda \
  biopython perl perl-uri miniprot td2 pip
conda activate quickprot
pip install rare-quickprot

alternatively using the conda-environment.yml file included in this repo:

conda env create -f conda-environment.yml

QuickProt User Guide

Update

  • 2026/04/28
    1. Major Update: The version number has been updated to 1.9.0
    2. BUSCO results have improved significantly!!!
    3. The runtime logic and results have been optimized, resulting in more comprehensive gene predictions.
    4. Partially overlapping genes (<0.2 overlap) are now permitted.
    5. Compatibility with newer versions of Python has been improved.
  • 2026/04/13
    1. extract_sequence_from_gff3.py now supports GFF3 file from NCBI. Some gene models in NCBI GFF3 contain in-frame stop codons; when translated into proteins, these are now converted to X (using the -cx option).
  • 2025/09/16
    1. Added add_type_gff3.py script.
    2. Optimized prediction of stop codon.
    3. Four new options have been added, namely "-c", "-ps", "-ms", and "-an", for quality control of protein mapping, rational use to reduce pseudogenes.
  • 2025/09/08
    1. Added gtf_genome_to_cdna_fasta.py and gtf_genome_to_cdna_fasta.py script.
    2. Input files now support .gz compressed files.
    3. Now QuickProt can output more running details.
  • 2025/09/06
    1. Provides information about the running process.
    2. Added filter_repeatPeps_from_gff3.py script for removing repeat proteins from gff3 file.
  • 2025/05/28
    1. Provide -ORFSoftware TD2 option, you can use TD2 as a tool for ORF prediction.
    2. Optimization of genetic code options.

What is QuickProt?

The QuickProt algorithm is a homology-based method for predicting gene models across entire genomes, designed to rapidly construct a non-redundant set of gene models. As illustrated in Figure 1, its core principle is analogous to the blotting method. It primarily employs miniprot (v0.18), to align homologous protein sequences to the genome, delineates high-alignment regions to assemble pseudo-transcripts (lacking UTR regions), and predicts coding regions within these pseudo-transcripts using TransDecoder (v5.7.1). Subsequently, low-quality gene models are filtered out and chimeric gene models are dissected, ultimately generating a high-accuracy, non-redundant gene set.

Schema of quickprot algorithm

Fig1. Schema of QuickProt algorithm

Installation:

Before use, you need to install Perl, Python, and biopython.

Python3 >= 3.8, perl >= 5

For ease of use, miniprot (v0.18) and TransDecoder (v5.7.1) software are integrated into QuickProt.

wget https://github.com/thecgs/quickprot/archive/refs/tags/quickprot-v1.8.0.tar.gz
tar -zxvf quickprot-v1.8.0.tar.gz
cd quickprot-v1.8.0
./quickprot -h

Note:

# if you need to use --mask optional of qucikprot.py script, and you need to install biopython
pip install biopython

# if you need to use sort_gff3.py script, and you need to install natsort.
pip install natsort

# if you need to use -ORFSoftware TD2, and you need to install TD2
pip install TD2

Usage:

To quickly run QuickProt software. like this,

./quickprot -q protein.fasta -g genome.fasta

This pipeline can improve busco missing result, but you need to download compleasm software.

## step1. running quickprot software
./quickprot.py -q protein.fasta -g genome.fasta -p quickprot.raw

## step2. running compleasm software
compleasm.py run -a genome.fasta -o ./ -l your_lineage

## step3. to update raw gff3 of step1 from compleasm result
./script/update_gff3_from_minibusco.py -r quickprot.raw.longest.gff3 -m ./your_lineage/miniprot_output.gff -g genome.fasta -o improve_busco.gff3

## step4. merge step1 and step3 gff3 result
cat quickprot.raw.longest.gff3 improve_busco.gff3 > genome.longest.gff.tmp

## step5. to sort by chromosomes or scaffold and gene start position and to rename gff3
./script/sort_gff3.py genome.longest.gff.tmp | ./script/rename_gff3.py - -o genome.longest.gff3 -p QUICKPROT; rm genome.longest.gff.tmp

## step6. extract protein sequence from genome and gff file
./bin/TransDecoder-5.7.1/util/gff3_file_to_proteins.pl --gff3 genome.longest.gff3 --fasta genome.fasta --seqType prot  > genome.longest.pep.fasta

## step7. extract CDS sequence from genome and gff file
./bin/TransDecoder-5.7.1/util/gff3_file_to_proteins.pl --gff3 genome.longest.gff3 --fasta genome.fasta --seqType CDS  > genome.longest.cds.fasta

This step can help you remove repeat proteins (e.g. ENV, Gag, Pol, RT, RH, INT, etc.), but you need to download diamond.

./script/filter_repeatPeps_from_gff3.py -q genome.longest.pep.fasta -g genome.longest.gff3

## results
## retain.gff3 —— Gene model without repeat proteins
## discard.gff3 —— Gene model of repeat proteins

Run with Singularity

Download the Singularity image here

singularity exec -B PATH -e quickprot.v1.9.0.sif quickprot.py -q protein.fasta -g genome.fasta

Test run

If you clone this repository, you can perform a test run of QuickProt using pytest. This may be useful if you wish to contribute to development.

First, it is convenient to set up a Conda environment for QuickProt testing:

git clone https://github.com/thecgs/quickprot.git
conda env create -f quickprot/test/environment.yml
conda activate quickprot-pytest

To perform the test run, simply clone and navigate to the repository, and run pytest.

cd quickprot
pytest -s

The included test datasets in test/data are FASTA files containing:

  • A collection of approximately 8,000 Saccharomyces protein sequences downloaded from UniProt (query proteins, Swiss-prot reviewed with tax ID 4930)
  • Chromosome 1 of Saccharomyces cerevisiae from Ensembl (genome, softmasked).

When the test is executed, pytest will carry out annotation of the toy genome sequence with the toy protein query set and check for equality of the output files with a set of reference outputs found in test/data/output using MD5 checksums. If all test outputs are equal to the references, the test run will pass, if not, it will fail.

If you include the option --TD2, pytest will carry out two test runs, one using TransDecoder and a second one using TD2.

pytest -s --TD2

Note: TD2's outputs are not as deterministic, so only the miniprot output and transcript assembly will be checked

Cite QuickProt:

If you use QuickProt, please cite:

Guisen Chen, Hehe Du, Zhenjie Cao, Ying Wu, Chen Zhang, Yongcan Zhou, Jingqun Ao, Yun Sun, Zihao Yuan. 2026. “ QuickProt: A Fast and Accurate Homology-Based Protein Annotation Tool for Non-Model Organisms to Advance Comparative Genomics.” Molecular Ecology Resources 26, no. 2: e70097. https://doi.org/10.1111/1755-0998.70097.

Release files for rare-quickprot 1.10.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for rare-quickprot 1.10.2
File Size Uploaded
rare_quickprot-1.10.2.tar.gz 74.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for rare-quickprot 1.10.2
File Interpreter ABI Platform
rare_quickprot-1.10.2-py3-none-any.whl Python 3 none any Details

Total release size: 241.8 kB

Release files / rare_quickprot-1.10.2.tar.gz

Download URL rare_quickprot-1.10.2.tar.gz
Size 74.9 kB
Tags Source
SHA-256 checksum
How to use checksums
d907d27493309ff3ae261d3f49c1a142113694f373498b2586259063bf4b1335
BLAKE2b-256 checksum
How to use checksums
7013c37ad4a186ef7ae02c66973f5369239e322ec8c48560017ad1facaddd141
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.14

Release files / rare_quickprot-1.10.2-py3-none-any.whl

Download URL rare_quickprot-1.10.2-py3-none-any.whl
Size 166.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0dec424466cff781804495018ceda7705cbab4bd3c18af0df868906f95a91791
BLAKE2b-256 checksum
How to use checksums
7d22e720222e6f5a2112cf7844a962136cd022655b7ed8fcb56b27103e261c65
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.14

Release history Release notifications | RSS feed

This release

1.10.2 This release

2 release files

1.10.1

2 release files

1.10.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page