gffkit
gffkit is a lightweight toolkit for GFF/GTF annotation integration and analysis.
Version 0.6.0 provides six standalone utilities:
detect-bridge: detect suspicious merged-gene artifacts caused by bridge transcripts.complement: complement/merge annotations, with optional region-swap mode.add-utr: reconstructfive_prime_UTRandthree_prime_UTRfeatures from exon/CDS coordinates.rename-sort: rename gene/transcript/child IDs with a prefix and sort the final GFF3.get_longest_transcript: retain one longest transcript isoform per gene.stat: calculate basic gene-annotation statistics for one GFF or a directory of GFF files.
The integrate workflow continues to use only the original four integration
utilities. get_longest_transcript and stat are independent commands and are
not automatically run by integrate.
Installation
pip install gffkit
Quick start
Full integration pipeline
gffkit integrate \
--annotation-a EviAnn.gff3 \
--annotation-b ANNEVO.gff3 \
--outdir gffkit_out \
--prefix sample \
-t 8
Outputs:
gffkit_out/sample.suspicious.tsvgffkit_out/sample.merged.gff3gffkit_out/sample.final.withUTR.gff3.pre_rename.gff3gffkit_out/sample.final.withUTR.gff3gffkit_out/sample.final.withUTR.gff3.id_map.tsv
In integrate, --prefix sample is also used for final ID renaming. Final gene IDs are written like sample_C01g00001, transcript IDs like sample_C01g00001.t1, and child IDs like sample_C01g00001.t1.exon1.
Step-by-step usage
# 1. Detect suspicious merged genes in Annotation A
gffkit detect-bridge -i EviAnn.gff3 -o suspicious.tsv -t 8
# 2. Use A as the global reference, but switch to B in suspicious regions
gffkit complement \
--ref EviAnn.gff3 \
--add ANNEVO.gff3 \
--swap_region_tsv suspicious.tsv \
--swap_region_flank 100 \
--output merged.gff3 \
-t 8
# 3. Add UTR features
gffkit add-utr -i merged.gff3 -o final.annotation.withUTR.pre_rename.gff3
# 4. Rename IDs, drop unplaced seqids, and sort the final GFF3
gffkit rename-sort \
-i final.annotation.withUTR.pre_rename.gff3 \
-o final.annotation.withUTR.gff3 \
--prefix sample
Merge three or more annotations
Use repeated --add arguments. Files are merged in the order provided.
gffkit complement \
--ref EviAnn.gff3 \
--add ANNEVO.gff3 \
--add Helixer.gff3 \
--add PASA.gff3 \
--output merged.multi.gff3 \
-t 8
Extract the longest transcript
For each gene, get_longest_transcript selects the isoform with the longest
concatenated CDS. If none of a gene's isoforms has CDS features, it uses the
longest concatenated exon length. Ties retain the first isoform encountered in
the input file.
Process one GFF/GFF3 file:
gffkit get_longest_transcript \
-i annotation.gff3 \
-o longest_gff
Process all .gff/.gff3 files directly inside a directory:
gffkit get_longest_transcript \
-f input_gff_directory \
-o longest_gff
Compressed .gff.gz and .gff3.gz inputs are also supported by this command.
Output files are named *.longest.gff or *.longest.gff3. Existing output
files are refused unless --force is supplied.
Calculate annotation statistics
Process one GFF/GFF3 file and write the default gffstat.txt:
gffkit stat -i annotation.gff3
Process all .gff/.gff3 files directly inside a directory and select an
output path:
gffkit stat \
-f input_gff_directory \
-o annotation.gffstat.txt
The tab-separated output contains species/file name, gene count, average gene
span, average transcript span, average CDS length per gene, average exon count
per gene, average exon length, and average UTR length per gene. Gene and
transcript lengths use their GFF genomic spans (end - start + 1); they are not
spliced mature-transcript lengths.
Command overview
gffkit --help
gffkit detect-bridge --help
gffkit complement --help
gffkit add-utr --help
gffkit rename-sort --help
gffkit get_longest_transcript --help
gffkit stat --help
gffkit integrate --help
Parallel Processing
Use -t/--threads to select the number of worker processes. The option name is
kept for command-line compatibility, but version 0.5.0 and later uses processes for
CPU-bound work so that Python's GIL does not restrict execution to one core.
detect-bridgeanalyzes genes with a process pool and batches tasks to reduce process communication overhead.complementparses the reference and supplementary files in parallel, then merges them in the original command-line order.complementuses a dynamic chromosome interval index to compare each supplementary gene only with nearby overlapping reference genes.integratepasses the worker count to the detect and complement steps.add-utrandrename-sortremain single-process.
Example:
gffkit integrate --annotation-a EviAnn.gff3 --annotation-b ANNEVO.gff3 -t 16
During detect-bridge, multiple worker processes should be visible in top or
htop. CPU usage naturally drops during the single-process UTR and rename/sort
steps. More workers also increase memory use; start with -t 4 or -t 8 for
large annotations.
Annotation integration strategy
- Annotation A, for example EviAnn/RNA-seq-supported GFF, is used as the global primary reference.
- Annotation B, for example ANNEVO/deep-learning GFF, is used as the local primary reference only in suspicious merged-gene regions.
- UTR features are reconstructed after merging using an exon-minus-CDS strategy.
- Version 0.4.0 and later run
rename-sortas the finalintegratestep. The final GFF3 keeps chromosome-mounted records, removes unplaced/scaffold/contig records, sorts features, rewritesID/Parent, and writes an ID map next to the output. - Version 0.5.0 replaces CPU-bound threads with worker processes and adds a dynamic interval index for faster annotation overlap searches.
- Version 0.6.0 adds the independent
get_longest_transcriptandstatcommands. Neither command is included inintegrate. - When multiple tools annotate the same gene locus, the GFF source column is combined with
|, for exampleEviAnn|ANNEVO.
Rename and Sort
Run this step independently when you already have a merged GFF3:
gffkit rename-sort \
-i merged.withUTR.gff3 \
-o sample.renamed.sorted.gff3 \
--prefix sample \
--digits 5 \
--keep-old-ids
This writes sample.renamed.sorted.gff3 and sample.renamed.sorted.gff3.id_map.tsv.
Version 0.6.0 changes
- Added
gffkit get_longest_transcriptusing the longest-CDS/longest-exon selection logic fromget_longest_transcript.py. - Added mutually exclusive
-i/--inputsingle-file and-f/--folderdirectory input modes for longest-transcript extraction. - Added
gffkit statusing the GFF statistics implementation fromgffstat.py. - Added mutually exclusive
-i/--inputsingle-file and-f/--folderdirectory input modes for annotation statistics. - Kept both new commands independent from the four-step
gffkit integrateworkflow. - Updated the package version from 0.5.0 to 0.6.0.
- Updated package author and credits to
caijunhao.
Maintainer notes
When command-line options or behavior changes, update this README.md in the versioned package directory before building and uploading to PyPI.
License
MIT License.
Credits
Author: caijunhao
Metadata
Release files for gffkit 0.6.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| gffkit-0.6.0.tar.gz | 42.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| gffkit-0.6.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 84.6 kB
Release files / gffkit-0.6.0.tar.gz
| Download URL | gffkit-0.6.0.tar.gz |
|---|---|
| Size | 42.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
0db9fe429bf66511dd8264e59e014401dd4b1432807b0429f29b93ab128c0592
|
|
BLAKE2b-256 checksum How to use checksums |
fbd6001462e2d228cdb54682043f85467cf41d94fdb9915a979128fa64a66a9b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.7.12
|
Release files / gffkit-0.6.0-py3-none-any.whl
| Download URL | gffkit-0.6.0-py3-none-any.whl |
|---|---|
| Size | 42.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
1e701f3d415d0616f441e5b931a036a64c7d935afc9980f6ff2e57119b2f2eb9
|
|
BLAKE2b-256 checksum How to use checksums |
0c9aaf6cee4f096ffabdf4a1a2fbd229ec02f5314ff13b221d30132b4aef8497
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.7.12
|