Skip to main content

gffkit

gffkit is a lightweight toolkit for region-aware GFF/GTF annotation integration. It combines four utilities:

  1. detect-bridge: detect suspicious merged-gene artifacts caused by bridge transcripts.
  2. complement: complement/merge annotations, with optional region-swap mode.
  3. add-utr: reconstruct five_prime_UTR and three_prime_UTR features from exon/CDS coordinates.
  4. rename-sort: rename gene/transcript/child IDs with a prefix and sort the final GFF3.

Installation

pip install gffkit

Quick start

Full integration pipeline

gffkit integrate \
  --annotation-a EviAnn.gff3 \
  --annotation-b ANNEVO.gff3 \
  --outdir gffkit_out \
  --prefix sample \
  -t 8

Outputs:

  • gffkit_out/sample.suspicious.tsv
  • gffkit_out/sample.merged.gff3
  • gffkit_out/sample.final.withUTR.gff3.pre_rename.gff3
  • gffkit_out/sample.final.withUTR.gff3
  • gffkit_out/sample.final.withUTR.gff3.id_map.tsv

In integrate, --prefix sample is also used for final ID renaming. Final gene IDs are written like sample_C01g00001, transcript IDs like sample_C01g00001.t1, and child IDs like sample_C01g00001.t1.exon1.

Step-by-step usage

# 1. Detect suspicious merged genes in Annotation A
gffkit detect-bridge -i EviAnn.gff3 -o suspicious.tsv -t 8

# 2. Use A as the global reference, but switch to B in suspicious regions
gffkit complement \
  --ref EviAnn.gff3 \
  --add ANNEVO.gff3 \
  --swap_region_tsv suspicious.tsv \
  --swap_region_flank 100 \
  --output merged.gff3 \
  -t 8

# 3. Add UTR features
gffkit add-utr -i merged.gff3 -o final.annotation.withUTR.pre_rename.gff3

# 4. Rename IDs, drop unplaced seqids, and sort the final GFF3
gffkit rename-sort \
  -i final.annotation.withUTR.pre_rename.gff3 \
  -o final.annotation.withUTR.gff3 \
  --prefix sample

Merge three or more annotations

Use repeated --add arguments. Files are merged in the order provided.

gffkit complement \
  --ref EviAnn.gff3 \
  --add ANNEVO.gff3 \
  --add Helixer.gff3 \
  --add PASA.gff3 \
  --output merged.multi.gff3 \
  -t 8

Command overview

gffkit --help
gffkit detect-bridge --help
gffkit complement --help
gffkit add-utr --help
gffkit rename-sort --help
gffkit integrate --help

Parallel Processing

Use -t/--threads to select the number of worker processes. The option name is kept for command-line compatibility, but version 0.5.0 uses processes for CPU-bound work so that Python's GIL does not restrict execution to one core.

  • detect-bridge analyzes genes with a process pool and batches tasks to reduce process communication overhead.
  • complement parses the reference and supplementary files in parallel, then merges them in the original command-line order.
  • complement uses a dynamic chromosome interval index to compare each supplementary gene only with nearby overlapping reference genes.
  • integrate passes the worker count to the detect and complement steps. add-utr and rename-sort remain single-process.

Example:

gffkit integrate --annotation-a EviAnn.gff3 --annotation-b ANNEVO.gff3 -t 16

During detect-bridge, multiple worker processes should be visible in top or htop. CPU usage naturally drops during the single-process UTR and rename/sort steps. More workers also increase memory use; start with -t 4 or -t 8 for large annotations.

Annotation integration strategy

  • Annotation A, for example EviAnn/RNA-seq-supported GFF, is used as the global primary reference.
  • Annotation B, for example ANNEVO/deep-learning GFF, is used as the local primary reference only in suspicious merged-gene regions.
  • UTR features are reconstructed after merging using an exon-minus-CDS strategy.
  • Version 0.4.0 and later run rename-sort as the final integrate step. The final GFF3 keeps chromosome-mounted records, removes unplaced/scaffold/contig records, sorts features, rewrites ID/Parent, and writes an ID map next to the output.
  • Version 0.5.0 replaces CPU-bound threads with worker processes and adds a dynamic interval index for faster annotation overlap searches.
  • When multiple tools annotate the same gene locus, the GFF source column is combined with |, for example EviAnn|ANNEVO.

Rename and Sort

Run this step independently when you already have a merged GFF3:

gffkit rename-sort \
  -i merged.withUTR.gff3 \
  -o sample.renamed.sorted.gff3 \
  --prefix sample \
  --digits 5 \
  --keep-old-ids

This writes sample.renamed.sorted.gff3 and sample.renamed.sorted.gff3.id_map.tsv.

Maintainer notes

When command-line options or behavior changes, update this README.md in the versioned package directory before building and uploading to PyPI.

License

MIT License.

Metadata

Release files for gffkit 0.5.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for gffkit 0.5.0
File Size Uploaded
gffkit-0.5.0.tar.gz 34.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for gffkit 0.5.0
File Interpreter ABI Platform
gffkit-0.5.0-py3-none-any.whl Python 3 none any Details

Total release size: 69.9 kB

Release files / gffkit-0.5.0.tar.gz

Download URL gffkit-0.5.0.tar.gz
Size 34.8 kB
Tags Source
SHA-256 checksum
How to use checksums
4f20dbddfabf9cf0012ed321999584b9e9fb1635ba24ff1f67d9074460e9079b
BLAKE2b-256 checksum
How to use checksums
ccea0e656d501dd79d5f1ab81c26c595b1538b679a07717f8eae2a4403f748b8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.7.12

Release files / gffkit-0.5.0-py3-none-any.whl

Download URL gffkit-0.5.0-py3-none-any.whl
Size 35.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
449409d0bcfa2d453d254d50e4caec7e199f12ba1360599eee601aa9bbd15108
BLAKE2b-256 checksum
How to use checksums
8035867cd99850ddfa9af82021d5ede29c773e8b0bb9f7612685333c2f7137cd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.7.12

Release history Release notifications | RSS feed

0.6.0

2 release files

This release

0.5.0 This release

2 release files

0.4.0

2 release files

0.3.1

2 release files

0.3

2 release files

0.2

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page