PGGT
PGGT is a recoverable multi-genome workflow that converts pangenome graph
projections into allele, copy-number variation (CNV), and presence/absence
variation (PAV) components.
It runs six modules:
- project annotated genes onto GFA walks and build initial components;
- filter graph-supported components with DIAMOND and gene structure evidence;
- score every gene and mark low-score genes as noncoding;
- regroup singleton genes with reciprocal Miniprot and Pfam evidence;
- select representative genes and write allele/CNV/PAV component tables and W-only GFA files;
- write one filtered GFF3 per sample containing complete original annotation
models for genes classified as
PASS.
Module 3 uses component-specific score thresholds: 1.0 for singleton,
0.9 for one_to_one_partial, and 0.8 for one_to_one_allele and
one_to_many_supported. A gene is retained as PASS regardless of its score
when its CDS is complete, Pfam status is PASS, and total exon length is
strictly greater than 200 bp. Other genes below their component threshold
remain in the main component table with ng replacing the terminal gene
marker g. The original-to-output relationship and decision are recorded in
13.gene_id_mapping.tsv.
What's new in 1.3.0
- Module 3 applies component-specific decision thresholds and records the applied threshold, strong-evidence rescue flag, and decision reason.
- Complete-CDS genes with a Pfam hit and total exon length greater than 200 bp are retained by a strong coding-evidence rule.
- Original GFF biotype labels are not used as automatic PASS or NOCODING overrides.
- New module 6 creates
06.coding_gene_gff/with one<sample>.protein_coding.gff3file per sample. It selects genes from13.gene_id_mapping.tsv, preserves original IDs, and preserves source GFF order.
Installation
PGGT requires Linux and Python 3.9 or later.
pip install pggt
The workflow also calls external command-line programs that are not installed by pip:
bedtoolsdiamondsamtools- HMMER (
hmmscanandhmmpress) miniprot
For example, these programs can be prepared with Conda/Bioconda:
conda create -n pggt -c conda-forge -c bioconda \
python=3.10 bedtools diamond samtools hmmer miniprot
conda activate pggt
pip install pggt
Confirm that the command is available:
pggt --help
Quick start
Run input validation first. --dry-run creates no result files.
pggt \
--gfa-input pangenome.gfa \
--genome-dir genomes/ \
--gff-dir annotations/ \
--samples species_list.txt \
--expression-dir expression/ \
--pfam-hmm Pfam-A.hmm \
--output-dir pggt_out/ \
--threads 56 \
--parallel-jobs 4 \
--dry-run
After validation succeeds, run the same command without --dry-run.
To reuse Pfam work when rerunning the same species and annotations, add either an earlier PGGT result directory or its compact Pfam table:
pggt [all required arguments] \
--previous-pfam-results previous_pggt_out/
The previous and current runs must use a Pfam database with the same fingerprint. Reuse is sequence-based, so proteins absent from the previous table or with changed SHA256 values are scanned automatically.
PGGT 1.1 tables do not retain the statistics for every accepted hit. To avoid carrying the 1.1 best-hit direction bug forward, unambiguous no-domain and single-domain legacy rows are reusable; legacy multi-domain rows use an exact shared-cache record when available and otherwise are rescanned. Tables created by PGGT 1.2 or later retain the corrected best hit and are fully reusable.
The output directory must be new. Do not use a GraphTribe result directory as the PGGT output directory.
Input requirements
Sample list
species_list.txt must contain exactly one sample name per non-empty line
and at least two unique samples:
B73
EIL02
EIL31
EIL08
EIL37
Sample names may contain letters, numbers, underscores, dots, and hyphens. The
first sample is the reference sample used for representative-gene selection.
Every listed sample must occur in the second field of at least one GFA W
record.
The same sample names are used to resolve genome, GFF, and expression files. Each sample must resolve to exactly one file of each type.
GFA
--gfa-input accepts either one GFA file or a directory containing exactly
one *.gfa file. Use --gfa-name FILE when a directory contains multiple
GFA files.
GFA W coordinates are interpreted as 0-based, half-open intervals.
Genome FASTA files
Provide one uncompressed genome FASTA per sample. Recommended names are:
genomes/B73.fa
genomes/EIL02.fa
PGGT first checks <sample>.fa, <sample>.fasta, <sample>.fna, and
<sample>.fas. If none exists, it accepts one uniquely matching FASTA whose
name contains the sample name.
FASTA sequence names must match the chromosome or sequence names used by the corresponding GFF and GFA walks.
GFF/GFF3 files
Recommended names are:
annotations/B73.renamed.sorted.chr.gff3
annotations/EIL02.renamed.sorted.chr.gff3
PGGT also checks <sample>.gff3 and <sample>.gff, followed by one unique
matching GFF file whose name contains the sample name.
The annotations must provide gene, transcript, exon, and CDS relationships.
Gene IDs must be consistent with the expression tables and use the supported
terminal g plus digits pattern, such as B73_C05g06442; this allows PGGT to
mark a noncoding output ID as B73_C05ng06442 without ambiguity. Chromosome
names must match the corresponding genome FASTA.
Expression tables
Expression input is required for every sample. Each file must be a
tab-separated TSV table. The recommended file name is
<sample>.FPKM.tsv; <sample>.expression.tsv is also accepted.
Example:
Gene_ID leaf_FPKM root_FPKM
B73_C01g00001 12.50 3.20
B73_C01g00002 0 0.08
B73_C01g00003 0 0
Expression-table rules:
- The header must contain the exact, case-sensitive column name
Gene_ID. - At least one expression column must end exactly with
_FPKM. - Gene IDs must match gene IDs in the corresponding GFF/GFF3.
- FPKM cells must be numeric. Do not write
NA,N/A, or other text. - Empty FPKM cells are interpreted as zero.
- When multiple
*_FPKMcolumns are present, PGGT uses the maximum value for each gene. - Use one row per gene. If a gene ID occurs more than once, the later row replaces the earlier value.
- Genes absent from the expression table receive the
no_expressionscore.
Extra non-FPKM columns are allowed. PGGT first checks
<sample>.FPKM.tsv and <sample>.expression.tsv; otherwise, exactly one
*<sample>*.tsv file must exist in the expression directory. Ambiguous
matches stop the run.
Pfam database
--pfam-hmm must point to a non-empty Pfam-A.hmm file. PGGT prepares the
required HMMER index files in its working area when necessary.
Run control
# Resume a previously interrupted PGGT run
pggt [all original arguments] --resume
# Run only one module
pggt [all required arguments] --only-step 3
# Run a module range
pggt [all required arguments] --from-step 2 --to-step 4
Use the same inputs, output directory, thresholds, thread count, and job count when resuming. Completed modules are validated through their completion markers before being skipped.
Threads and shared caches
--threads is the total CPU budget. --parallel-jobs controls concurrent
per-sample work, Pfam shards, Miniprot mappings, and PAF parsing, and must not
exceed --threads.
For five samples on a sufficiently large 56-core node, --parallel-jobs 5
allows all five Miniprot target mappings to run concurrently. Values above the
sample count can still increase Pfam shard concurrency, but should be
benchmarked against available memory and storage bandwidth.
Shared caches are stored under ~/.cache/pggt:
~/.cache/pggt/pfam/pfam_cache.sqlite3caches Pfam results by HMM fingerprint and protein sequence SHA256;~/.cache/pggt/miniprot/stores Miniprot genome indexes.
Set the PGGT_CACHE_DIR environment variable to use a different cache root.
Main outputs
pggt_out/
├── 01.graph_mapping/
├── 02.component_filter/
│ └── 10.filtered_component_table.txt
├── 03.gene_score/
│ ├── 11.scored_gene_component_table.txt
│ ├── 12.nocoding_gene_scores.tsv
│ ├── 13.gene_id_mapping.tsv
│ └── score_tmp_file/
│ ├── 07.pfam_scores.tsv
│ └── 08.singleton_score.tsv
├── 04.component_cnv/
│ └── 04.orthogroup/
│ └── singleton_orthogroup_component_table.txt
├── 05.allele_CNV_PAV/
│ ├── allele_CNV_PAV_component_table.txt
│ └── W_gfa/
│ └── allele_CNV_PAV.W.gfa
├── 06.coding_gene_gff/
│ ├── coding_gene_gff_summary.tsv
│ ├── B73.protein_coding.gff3
│ └── EIL02.protein_coding.gff3
└── 00.run_info/
└── module_timings.tsv
Each completed module contains artifact_manifest.tsv. Large intermediate
DIAMOND, Pfam, and Miniprot files are not published into successful result
directories. Failed module workspaces remain under OUTPUT_DIR/.work/ for
diagnosis.
Standalone module 5
When a compatible module-4 component table already exists, run module 5 directly:
pggt-module5 \
--component-table singleton_orthogroup_component_table.txt \
--samples-file species_list.txt \
--gff-dir annotations/ \
--gene-id-map 13.gene_id_mapping.tsv \
--output-dir module5_out/ \
--representative-seed 1
--gene-id-map is required when the component table contains module-3 ng
IDs but the source GFF still contains the original g IDs.
Standalone module 6
Module 6 can be run directly from a module-3 mapping table:
pggt-module6 \
--gene-id-map 13.gene_id_mapping.tsv \
--samples-file species_list.txt \
--gff-dir annotations/ \
--output-dir coding_gene_gff/ \
--parallel-jobs 4
Only mapping rows with decision=PASS are selected. The module uses
original_gene_id for source-GFF lookup and retains the gene line plus all of
its original transcript, exon, CDS, UTR, and other descendant features.
Gene and transcript IDs are not renumbered. Source GFF row order is retained, so a coordinate-sorted source remains coordinate-sorted after filtering. Renumbering is deliberately avoided because it would break links to expression tables, proteins, component tables, and the module-3 ID mapping. Embedded FASTA sections are not copied into the filtered annotation files.
Command overview
pggt --help
pggt-module5 --help
pggt-module6 --help
Maintainer notes
Before a PyPI release, run the regression tests, build both the wheel and
source distribution, run twine check dist/*, and verify installation in a
clean environment. Update this README whenever command-line options or input
requirements change.
Major workflow changes must be developed in a new versioned package directory,
with matching updates to pyproject.toml, runtime __version__, this
README, PGGT流程原理.md, and CHANGELOG.md. Do not modify a previously
published version directory in place.
License
License information will be added before the first public release.
Metadata
Release files for pggt 1.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pggt-1.3.0.tar.gz | 121.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pggt-1.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 235.0 kB
Release files / pggt-1.3.0.tar.gz
| Download URL | pggt-1.3.0.tar.gz |
|---|---|
| Size | 121.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4c3826436b859aad67bf63aed60a2a88664c2f6f0a23c7a17b8b1eb1551c3802
|
|
BLAKE2b-256 checksum How to use checksums |
8e41c4748d2e460ca19a334c0069ac0542cab4c48a224006641a9acb10dd7259
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.7.12
|
Release files / pggt-1.3.0-py3-none-any.whl
| Download URL | pggt-1.3.0-py3-none-any.whl |
|---|---|
| Size | 113.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
9e7c0ccbc712039133644b441e5179d5204d7f89d117d700ed6267fdb5100663
|
|
BLAKE2b-256 checksum How to use checksums |
9cf8fe40e66e0582957258d70a4ef2c699afca4d66499312a041b122291de5d0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.7.12
|