gffkit
gffkit is a lightweight toolkit for region-aware GFF/GTF annotation integration.
It combines four utilities:
detect-bridge: detect suspicious merged-gene artifacts caused by bridge transcripts.complement: complement/merge annotations, with optional region-swap mode.add-utr: reconstructfive_prime_UTRandthree_prime_UTRfeatures from exon/CDS coordinates.rename-sort: rename gene/transcript/child IDs with a prefix and sort the final GFF3.
Installation
pip install gffkit
Quick start
Full integration pipeline
gffkit integrate \
--annotation-a EviAnn.gff3 \
--annotation-b ANNEVO.gff3 \
--outdir gffkit_out \
--prefix sample \
-t 8
Outputs:
gffkit_out/sample.suspicious.tsvgffkit_out/sample.merged.gff3gffkit_out/sample.final.withUTR.gff3.pre_rename.gff3gffkit_out/sample.final.withUTR.gff3gffkit_out/sample.final.withUTR.gff3.id_map.tsv
In integrate, --prefix sample is also used for final ID renaming. Final gene IDs are written like sample_C01g00001, transcript IDs like sample_C01g00001.t1, and child IDs like sample_C01g00001.t1.exon1.
Step-by-step usage
# 1. Detect suspicious merged genes in Annotation A
gffkit detect-bridge -i EviAnn.gff3 -o suspicious.tsv -t 8
# 2. Use A as the global reference, but switch to B in suspicious regions
gffkit complement \
--ref EviAnn.gff3 \
--add ANNEVO.gff3 \
--swap_region_tsv suspicious.tsv \
--swap_region_flank 100 \
--output merged.gff3 \
-t 8
# 3. Add UTR features
gffkit add-utr -i merged.gff3 -o final.annotation.withUTR.pre_rename.gff3
# 4. Rename IDs, drop unplaced seqids, and sort the final GFF3
gffkit rename-sort \
-i final.annotation.withUTR.pre_rename.gff3 \
-o final.annotation.withUTR.gff3 \
--prefix sample
Merge three or more annotations
Use repeated --add arguments. Files are merged in the order provided.
gffkit complement \
--ref EviAnn.gff3 \
--add ANNEVO.gff3 \
--add Helixer.gff3 \
--add PASA.gff3 \
--output merged.multi.gff3 \
-t 8
Command overview
gffkit --help
gffkit detect-bridge --help
gffkit complement --help
gffkit add-utr --help
gffkit rename-sort --help
gffkit integrate --help
Parallel Processing
Use -t/--threads to select the number of worker processes. The option name is
kept for command-line compatibility, but version 0.5.0 uses processes for
CPU-bound work so that Python's GIL does not restrict execution to one core.
detect-bridgeanalyzes genes with a process pool and batches tasks to reduce process communication overhead.complementparses the reference and supplementary files in parallel, then merges them in the original command-line order.complementuses a dynamic chromosome interval index to compare each supplementary gene only with nearby overlapping reference genes.integratepasses the worker count to the detect and complement steps.add-utrandrename-sortremain single-process.
Example:
gffkit integrate --annotation-a EviAnn.gff3 --annotation-b ANNEVO.gff3 -t 16
During detect-bridge, multiple worker processes should be visible in top or
htop. CPU usage naturally drops during the single-process UTR and rename/sort
steps. More workers also increase memory use; start with -t 4 or -t 8 for
large annotations.
Annotation integration strategy
- Annotation A, for example EviAnn/RNA-seq-supported GFF, is used as the global primary reference.
- Annotation B, for example ANNEVO/deep-learning GFF, is used as the local primary reference only in suspicious merged-gene regions.
- UTR features are reconstructed after merging using an exon-minus-CDS strategy.
- Version 0.4.0 and later run
rename-sortas the finalintegratestep. The final GFF3 keeps chromosome-mounted records, removes unplaced/scaffold/contig records, sorts features, rewritesID/Parent, and writes an ID map next to the output. - Version 0.5.0 replaces CPU-bound threads with worker processes and adds a dynamic interval index for faster annotation overlap searches.
- When multiple tools annotate the same gene locus, the GFF source column is combined with
|, for exampleEviAnn|ANNEVO.
Rename and Sort
Run this step independently when you already have a merged GFF3:
gffkit rename-sort \
-i merged.withUTR.gff3 \
-o sample.renamed.sorted.gff3 \
--prefix sample \
--digits 5 \
--keep-old-ids
This writes sample.renamed.sorted.gff3 and sample.renamed.sorted.gff3.id_map.tsv.
Maintainer notes
When command-line options or behavior changes, update this README.md in the versioned package directory before building and uploading to PyPI.
License
MIT License.
Metadata
Release files for gffkit 0.5.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| gffkit-0.5.0.tar.gz | 34.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| gffkit-0.5.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 69.9 kB
Release files / gffkit-0.5.0.tar.gz
| Download URL | gffkit-0.5.0.tar.gz |
|---|---|
| Size | 34.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4f20dbddfabf9cf0012ed321999584b9e9fb1635ba24ff1f67d9074460e9079b
|
|
BLAKE2b-256 checksum How to use checksums |
ccea0e656d501dd79d5f1ab81c26c595b1538b679a07717f8eae2a4403f748b8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.7.12
|
Release files / gffkit-0.5.0-py3-none-any.whl
| Download URL | gffkit-0.5.0-py3-none-any.whl |
|---|---|
| Size | 35.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
449409d0bcfa2d453d254d50e4caec7e199f12ba1360599eee601aa9bbd15108
|
|
BLAKE2b-256 checksum How to use checksums |
8035867cd99850ddfa9af82021d5ede29c773e8b0bb9f7612685333c2f7137cd
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.7.12
|