Skip to main content

gffkit

gffkit is a lightweight toolkit for GFF/GTF annotation integration and analysis. Version 0.6.0 provides six standalone utilities:

  1. detect-bridge: detect suspicious merged-gene artifacts caused by bridge transcripts.
  2. complement: complement/merge annotations, with optional region-swap mode.
  3. add-utr: reconstruct five_prime_UTR and three_prime_UTR features from exon/CDS coordinates.
  4. rename-sort: rename gene/transcript/child IDs with a prefix and sort the final GFF3.
  5. get_longest_transcript: retain one longest transcript isoform per gene.
  6. stat: calculate basic gene-annotation statistics for one GFF or a directory of GFF files.

The integrate workflow continues to use only the original four integration utilities. get_longest_transcript and stat are independent commands and are not automatically run by integrate.

Installation

pip install gffkit

Quick start

Full integration pipeline

gffkit integrate \
  --annotation-a EviAnn.gff3 \
  --annotation-b ANNEVO.gff3 \
  --outdir gffkit_out \
  --prefix sample \
  -t 8

Outputs:

  • gffkit_out/sample.suspicious.tsv
  • gffkit_out/sample.merged.gff3
  • gffkit_out/sample.final.withUTR.gff3.pre_rename.gff3
  • gffkit_out/sample.final.withUTR.gff3
  • gffkit_out/sample.final.withUTR.gff3.id_map.tsv

In integrate, --prefix sample is also used for final ID renaming. Final gene IDs are written like sample_C01g00001, transcript IDs like sample_C01g00001.t1, and child IDs like sample_C01g00001.t1.exon1.

Step-by-step usage

# 1. Detect suspicious merged genes in Annotation A
gffkit detect-bridge -i EviAnn.gff3 -o suspicious.tsv -t 8

# 2. Use A as the global reference, but switch to B in suspicious regions
gffkit complement \
  --ref EviAnn.gff3 \
  --add ANNEVO.gff3 \
  --swap_region_tsv suspicious.tsv \
  --swap_region_flank 100 \
  --output merged.gff3 \
  -t 8

# 3. Add UTR features
gffkit add-utr -i merged.gff3 -o final.annotation.withUTR.pre_rename.gff3

# 4. Rename IDs, drop unplaced seqids, and sort the final GFF3
gffkit rename-sort \
  -i final.annotation.withUTR.pre_rename.gff3 \
  -o final.annotation.withUTR.gff3 \
  --prefix sample

Merge three or more annotations

Use repeated --add arguments. Files are merged in the order provided.

gffkit complement \
  --ref EviAnn.gff3 \
  --add ANNEVO.gff3 \
  --add Helixer.gff3 \
  --add PASA.gff3 \
  --output merged.multi.gff3 \
  -t 8

Extract the longest transcript

For each gene, get_longest_transcript selects the isoform with the longest concatenated CDS. If none of a gene's isoforms has CDS features, it uses the longest concatenated exon length. Ties retain the first isoform encountered in the input file.

Process one GFF/GFF3 file:

gffkit get_longest_transcript \
  -i annotation.gff3 \
  -o longest_gff

Process all .gff/.gff3 files directly inside a directory:

gffkit get_longest_transcript \
  -f input_gff_directory \
  -o longest_gff

Compressed .gff.gz and .gff3.gz inputs are also supported by this command. Output files are named *.longest.gff or *.longest.gff3. Existing output files are refused unless --force is supplied.

Calculate annotation statistics

Process one GFF/GFF3 file and write the default gffstat.txt:

gffkit stat -i annotation.gff3

Process all .gff/.gff3 files directly inside a directory and select an output path:

gffkit stat \
  -f input_gff_directory \
  -o annotation.gffstat.txt

The tab-separated output contains species/file name, gene count, average gene span, average transcript span, average CDS length per gene, average exon count per gene, average exon length, and average UTR length per gene. Gene and transcript lengths use their GFF genomic spans (end - start + 1); they are not spliced mature-transcript lengths.

Command overview

gffkit --help
gffkit detect-bridge --help
gffkit complement --help
gffkit add-utr --help
gffkit rename-sort --help
gffkit get_longest_transcript --help
gffkit stat --help
gffkit integrate --help

Parallel Processing

Use -t/--threads to select the number of worker processes. The option name is kept for command-line compatibility, but version 0.5.0 and later uses processes for CPU-bound work so that Python's GIL does not restrict execution to one core.

  • detect-bridge analyzes genes with a process pool and batches tasks to reduce process communication overhead.
  • complement parses the reference and supplementary files in parallel, then merges them in the original command-line order.
  • complement uses a dynamic chromosome interval index to compare each supplementary gene only with nearby overlapping reference genes.
  • integrate passes the worker count to the detect and complement steps. add-utr and rename-sort remain single-process.

Example:

gffkit integrate --annotation-a EviAnn.gff3 --annotation-b ANNEVO.gff3 -t 16

During detect-bridge, multiple worker processes should be visible in top or htop. CPU usage naturally drops during the single-process UTR and rename/sort steps. More workers also increase memory use; start with -t 4 or -t 8 for large annotations.

Annotation integration strategy

  • Annotation A, for example EviAnn/RNA-seq-supported GFF, is used as the global primary reference.
  • Annotation B, for example ANNEVO/deep-learning GFF, is used as the local primary reference only in suspicious merged-gene regions.
  • UTR features are reconstructed after merging using an exon-minus-CDS strategy.
  • Version 0.4.0 and later run rename-sort as the final integrate step. The final GFF3 keeps chromosome-mounted records, removes unplaced/scaffold/contig records, sorts features, rewrites ID/Parent, and writes an ID map next to the output.
  • Version 0.5.0 replaces CPU-bound threads with worker processes and adds a dynamic interval index for faster annotation overlap searches.
  • Version 0.6.0 adds the independent get_longest_transcript and stat commands. Neither command is included in integrate.
  • When multiple tools annotate the same gene locus, the GFF source column is combined with |, for example EviAnn|ANNEVO.

Rename and Sort

Run this step independently when you already have a merged GFF3:

gffkit rename-sort \
  -i merged.withUTR.gff3 \
  -o sample.renamed.sorted.gff3 \
  --prefix sample \
  --digits 5 \
  --keep-old-ids

This writes sample.renamed.sorted.gff3 and sample.renamed.sorted.gff3.id_map.tsv.

Version 0.6.0 changes

  • Added gffkit get_longest_transcript using the longest-CDS/longest-exon selection logic from get_longest_transcript.py.
  • Added mutually exclusive -i/--input single-file and -f/--folder directory input modes for longest-transcript extraction.
  • Added gffkit stat using the GFF statistics implementation from gffstat.py.
  • Added mutually exclusive -i/--input single-file and -f/--folder directory input modes for annotation statistics.
  • Kept both new commands independent from the four-step gffkit integrate workflow.
  • Updated the package version from 0.5.0 to 0.6.0.
  • Updated package author and credits to caijunhao.

Maintainer notes

When command-line options or behavior changes, update this README.md in the versioned package directory before building and uploading to PyPI.

License

MIT License.

Credits

Author: caijunhao

Metadata

Release files for gffkit 0.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for gffkit 0.6.0
File Size Uploaded
gffkit-0.6.0.tar.gz 42.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for gffkit 0.6.0
File Interpreter ABI Platform
gffkit-0.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 84.6 kB

Release files / gffkit-0.6.0.tar.gz

Download URL gffkit-0.6.0.tar.gz
Size 42.0 kB
Tags Source
SHA-256 checksum
How to use checksums
0db9fe429bf66511dd8264e59e014401dd4b1432807b0429f29b93ab128c0592
BLAKE2b-256 checksum
How to use checksums
fbd6001462e2d228cdb54682043f85467cf41d94fdb9915a979128fa64a66a9b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.7.12

Release files / gffkit-0.6.0-py3-none-any.whl

Download URL gffkit-0.6.0-py3-none-any.whl
Size 42.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1e701f3d415d0616f441e5b931a036a64c7d935afc9980f6ff2e57119b2f2eb9
BLAKE2b-256 checksum
How to use checksums
0c9aaf6cee4f096ffabdf4a1a2fbd229ec02f5314ff13b221d30132b4aef8497
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.7.12

Release history Release notifications | RSS feed

This release

0.6.0 This release

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.1

2 release files

0.3

2 release files

0.2

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page