Skip to main content

PGGT 1.4.0

PGGT is a recoverable multi-genome workflow that converts pangenome graph projections into allele, copy-number variation (CNV), and presence/absence variation (PAV) components.

It runs six modules:

  1. project annotated genes onto GFA walks and build initial components;
  2. filter graph-supported components with DIAMOND and gene structure evidence;
  3. score component genes using CDS, exon, Pfam and optional expression evidence;
  4. regroup singleton genes with reciprocal Miniprot and Pfam evidence;
  5. write member tables, an integer PAV/CNV matrix and two W-only GFA views;
  6. write both all-gene and protein-coding GFF3 models for each sample.

Module 3 uses component-specific score thresholds: 1.0 for singleton, 0.9 for one_to_one_partial, and 0.8 for one_to_one_allele and one_to_many_supported. A gene is retained as PASS regardless of its score when its CDS is complete, Pfam status is PASS, and total exon length is strictly greater than 200 bp. Other genes below their component threshold remain in the main component table with ng replacing the terminal gene marker g. The original-to-output relationship and decision are recorded in 13.gene_id_mapping.tsv.

What's new in 1.4.0

  • Expression input is optional. Omit -e / --expression-dir to skip only expression scoring; other scores, thresholds and strong-evidence rescue remain.
  • Each sample gets <sample>.all_genes.gff3 and <sample>.protein_coding.gff3. Noncoding IDs retain lowercase ng.
  • PAV_CNV.copy_number.tsv combines presence/absence and actual copy number: 0 absent, 1 one gene, 2, 3, 4, ... the number of gene copies.
  • W-only GFA is generated separately for all genes and PASS genes; repeated component visits retain distinct copies.
  • Exports validate coverage, gene hierarchy, matrix counts and W-path counts. Source GFF genes must all occur in the classification mapping and final components; unresolved differences stop export with a coverage report.

NOCODING is a computational classification, not an experimentally validated RNA biotype. Original biotypes do not automatically override classification. Source CDS and other features remain in all-gene exports for NOCODING genes.

Installation

PGGT requires Linux and Python 3.9 or later.

pip install pggt

The workflow also calls external command-line programs that are not installed by pip:

  • bedtools
  • diamond
  • samtools
  • HMMER (hmmscan and hmmpress)
  • miniprot

For example, these programs can be prepared with Conda/Bioconda:

conda create -n pggt -c conda-forge -c bioconda \
  python=3.10 bedtools diamond samtools hmmer miniprot
conda activate pggt
pip install pggt

Confirm that the command is available:

pggt --help

Quick start

Run input validation first. --dry-run creates no result files.

pggt \
  --gfa-input pangenome.gfa \
  --genome-dir genomes/ \
  --gff-dir annotations/ \
  --samples species_list.txt \
  --pfam-hmm Pfam-A.hmm \
  --output-dir pggt_out/ \
  --threads 56 \
  --parallel-jobs 4 \
  --dry-run

This example omits expression input. Add --expression-dir expression/ (or -e expression/) when expression tables are available. After validation succeeds, run the same command without --dry-run.

To reuse Pfam work when rerunning the same species and annotations, add either an earlier PGGT result directory or its compact Pfam table:

pggt [all required arguments] \
  --previous-pfam-results previous_pggt_out/

The previous and current runs must use a Pfam database with the same fingerprint. Reuse is sequence-based, so proteins absent from the previous table or with changed SHA256 values are scanned automatically.

PGGT 1.1 tables do not retain the statistics for every accepted hit. To avoid carrying the 1.1 best-hit direction bug forward, unambiguous no-domain and single-domain legacy rows are reusable; legacy multi-domain rows use an exact shared-cache record when available and otherwise are rescanned. Tables created by PGGT 1.2 or later retain the corrected best hit and are fully reusable.

The output directory must be new. Do not use a GraphTribe result directory as the PGGT output directory.

Input requirements

Sample list

species_list.txt must contain exactly one sample name per non-empty line and at least two unique samples:

B73
EIL02
EIL31
EIL08
EIL37

Sample names may contain letters, numbers, underscores, dots, and hyphens. The first sample is the reference sample used for representative-gene selection. Every listed sample must occur in the second field of at least one GFA W record.

Sample names resolve genome and GFF files, and expression files when enabled. Each enabled input type must resolve uniquely for every sample.

GFA

--gfa-input accepts either one GFA file or a directory containing exactly one *.gfa file. Use --gfa-name FILE when a directory contains multiple GFA files.

GFA W coordinates are interpreted as 0-based, half-open intervals.

Genome FASTA files

Provide one uncompressed genome FASTA per sample. Recommended names are:

genomes/B73.fa
genomes/EIL02.fa

PGGT first checks <sample>.fa, <sample>.fasta, <sample>.fna, and <sample>.fas. If none exists, it accepts one uniquely matching FASTA whose name contains the sample name.

FASTA sequence names must match the chromosome or sequence names used by the corresponding GFF and GFA walks.

GFF/GFF3 files

Recommended names are:

annotations/B73.renamed.sorted.chr.gff3
annotations/EIL02.renamed.sorted.chr.gff3

PGGT also checks <sample>.gff3 and <sample>.gff, followed by one unique matching GFF file whose name contains the sample name.

The annotations must provide gene, transcript, exon, and CDS relationships. Gene IDs must match expression tables when supplied and use the supported terminal g plus digits pattern, such as B73_C05g06442; this allows PGGT to mark a noncoding output ID as B73_C05ng06442 without ambiguity. Chromosome names must match the corresponding genome FASTA.

Expression tables

Expression input is optional for the run. If supplied, the directory must contain one file for every sample; invalid paths, missing files and invalid tables remain errors. Each file must be a tab-separated TSV table. The recommended file name is <sample>.FPKM.tsv; <sample>.expression.tsv is also accepted.

Example:

Gene_ID	leaf_FPKM	root_FPKM
B73_C01g00001	12.50	3.20
B73_C01g00002	0	0.08
B73_C01g00003	0	0

Expression-table rules:

  • The header must contain the exact, case-sensitive column name Gene_ID.
  • At least one expression column must end exactly with _FPKM.
  • Gene IDs must match gene IDs in the corresponding GFF/GFF3.
  • FPKM cells must be numeric. Do not write NA, N/A, or other text.
  • Empty FPKM cells are interpreted as zero.
  • When multiple *_FPKM columns are present, PGGT uses the maximum value for each gene.
  • Use one row per gene. If a gene ID occurs more than once, the later row replaces the earlier value.
  • Genes absent from a supplied table receive zero expression points with expression_status=GENE_NOT_FOUND.

Extra non-FPKM columns are allowed. PGGT first checks <sample>.FPKM.tsv and <sample>.expression.tsv; otherwise, exactly one *<sample>*.tsv file must exist in the expression directory. Ambiguous matches stop the run.

Without expression input, the total is CDS_score + exon_score + Pfam_score. With expression input, PGGT adds the existing expression_score. There is no renormalization, extra penalty or threshold change when expression is omitted. Other evidence can still produce PASS, including strong coding-evidence rescue.

The retained score_tmp_file/05.expression_scores.tsv contains audit rows even when scoring is skipped. Total-score rows also contain expression_status:

Situation max_FPKM expression_score expression_status
No expression directory NA 0.00 SKIPPED_NO_INPUT
Gene missing from supplied table NA 0.00 GENE_NOT_FOUND
Measured zero 0.000000 0.00 OBSERVED
Positive measurement Observed value Existing rule OBSERVED

The module-3 shell entry accepts optional -e / --expression; its Python entry accepts optional --expression-dir. Partial per-sample expression availability is not supported in this release.

Pfam database

--pfam-hmm must point to a non-empty Pfam-A.hmm file. PGGT prepares the required HMMER index files in its working area when necessary.

Run control

# Resume a previously interrupted PGGT run
pggt [all original arguments] --resume

# Run only one module
pggt [all required arguments] --only-step 3

# Run a module range
pggt [all required arguments] --from-step 2 --to-step 4

Use the same inputs, output directory, thresholds, thread count, and job count when resuming. Completed modules are validated through their completion markers before being skipped. Modules 5 and 6 additionally check the new exports and their counts, including valid empty coding files.

Use a new output directory when upgrading from 1.3.0, changing expression mode, or changing expression inputs. Resume checks the version, expression enablement, directory and content fingerprint and rejects mismatches. Compatible Pfam evidence can still be reused through --previous-pfam-results.

Threads and shared caches

--threads is the total CPU budget. --parallel-jobs controls concurrent per-sample work, Pfam shards, Miniprot mappings, and PAF parsing, and must not exceed --threads.

For five samples on a sufficiently large 56-core node, --parallel-jobs 5 allows all five Miniprot target mappings to run concurrently. Values above the sample count can still increase Pfam shard concurrency, but should be benchmarked against available memory and storage bandwidth.

Shared caches are stored under ~/.cache/pggt:

  • ~/.cache/pggt/pfam/pfam_cache.sqlite3 caches Pfam results by HMM fingerprint and protein sequence SHA256;
  • ~/.cache/pggt/miniprot/ stores Miniprot genome indexes.

Set the PGGT_CACHE_DIR environment variable to use a different cache root.

Main outputs

pggt_out/
├── 00.run_info/                         # parameters, logs, timings
├── 00.input_view/                       # expression links only if enabled
├── status/
├── 01.graph_mapping/
├── 02.component_filter/
│   └── 10.filtered_component_table.txt
├── 03.gene_score/
│   ├── 11.scored_gene_component_table.txt
│   ├── 12.nocoding_gene_scores.tsv
│   ├── 13.gene_id_mapping.tsv
│   └── score_tmp_file/
│       ├── 05.expression_scores.tsv
│       ├── 07.pfam_scores.tsv
│       ├── 08.singleton_score.tsv
│       └── 09.summary.txt
├── 04.component_cnv/
│   └── 04.orthogroup/
│       └── singleton_orthogroup_component_table.txt
├── 05.allele_CNV_PAV/
│   ├── allele_CNV_PAV_component_table.txt
│   ├── PAV_CNV.copy_number.tsv
│   ├── protein_coding_component_table.txt
│   ├── representative_gene_selection.tsv
│   ├── protein_coding_representative_selection.tsv
│   ├── gene_coverage.tsv
│   ├── module5_metadata.tsv
│   ├── artifact_manifest.tsv
│   └── W_gfa/
│       ├── <sample>.W.gfa               # compatibility copy: all genes
│       ├── allele_CNV_PAV.W.gfa         # compatibility copy: all genes
│       ├── all_genes/
│       │   ├── <sample>.all_genes.W.gfa
│       │   ├── all_genes.W.gfa
│       │   └── <view audit tables>
│       └── protein_coding/
│           ├── <sample>.protein_coding.W.gfa
│           ├── protein_coding.W.gfa
│           └── <view audit tables>
└── 06.coding_gene_gff/
    ├── <sample>.protein_coding.gff3
    ├── <sample>.all_genes.gff3
    ├── <sample>.gene_coverage.tsv
    ├── coding_gene_gff_summary.tsv
    ├── all_gene_gff_summary.tsv
    └── artifact_manifest.tsv

Each GFA view contains component_segment_map.tsv, gene_segment_map.tsv, chromosome_w_path_stats.tsv, missing_or_extra_genes.tsv and make_w_gfa.summary.tsv. No PASS genes produces an empty coding W file and status NO_CODING_GENES.

The integer matrix counts unique gene members including ng, not transcripts; sample columns follow the sample list. Member tables retain IDs for GFA export. Coding components keep the original component ID and class label while their membership, gene count and sample count are filtered. If a representative is noncoding, the coding view reselects a reproducible PASS representative. Join views by component ID; node names can differ. These are W-only gene-order paths, not complete sequence graphs with S/L records.

Each completed pipeline module contains artifact_manifest.tsv. Large intermediate DIAMOND, Pfam, and Miniprot files are not published into successful result directories. Failed module workspaces remain under OUTPUT_DIR/.work/ for diagnosis.

Standalone module 5

When a compatible module-4 component table already exists, run module 5 directly:

pggt-module5 \
  --component-table singleton_orthogroup_component_table.txt \
  --samples-file species_list.txt \
  --gff-dir annotations/ \
  --gene-id-map 13.gene_id_mapping.tsv \
  --output-dir module5_out/ \
  --representative-seed 1

--gene-id-map is now always required, including for all-PASS input, because both views require explicit classification. Required columns are sample, original_gene_id, output_gene_id and decision. Duplicate memberships or missing classifications are errors. gene_coverage.tsv reports unresolved GFF/mapping/component differences.

Standalone module 6

Module 6 can be run directly from a module-3 mapping table:

pggt-module6 \
  --gene-id-map 13.gene_id_mapping.tsv \
  --samples-file species_list.txt \
  --gff-dir annotations/ \
  --output-dir coding_gene_gff/ \
  --parallel-jobs 4

Both GFF views are written in one run. Protein-coding output selects decision=PASS and preserves original IDs. All-gene output includes PASS and NOCODING, uses output_gene_id (lowercase ng for NOCODING), and adds PGGT_decision and PGGT_original_gene_id attributes to gene rows.

Both views retain original transcript, exon, CDS and UTR models, coordinates, strand and phase. Child IDs remain unchanged; Parent and geneID references are updated for renamed or filtered genes. Multi-parent features and children before parents are supported; invalid hierarchies and rename collisions are rejected. Source row order is retained. Embedded FASTA is omitted.

The supported gene root feature is gene. Other gene-like roots such as pseudogene require upstream normalization and are rejected with a diagnostic. The original GFF and mapping must cover the same gene set. Per-sample coverage reports identify missing entries; this release does not invent classifications or singleton components for uncovered genes.

Command overview

pggt --help
pggt-module5 --help
pggt-module6 --help

Maintainer notes

Before a PyPI release, run the regression tests, build both the wheel and source distribution, run twine check dist/*, and verify installation in a clean environment. Update this README whenever command-line options or input requirements change.

Major workflow changes must be developed in a new versioned package directory, with matching updates to pyproject.toml, runtime __version__, this README, PGGT流程原理.md, and CHANGELOG.md. Do not modify a previously published version directory in place.

License

License information will be added before the first public release.

Metadata

Release files for pggt 1.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pggt 1.4.0
File Size Uploaded
pggt-1.4.0.tar.gz 141.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pggt 1.4.0
File Interpreter ABI Platform
pggt-1.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 264.7 kB

Release files / pggt-1.4.0.tar.gz

Download URL pggt-1.4.0.tar.gz
Size 141.9 kB
Tags Source
SHA-256 checksum
How to use checksums
39708afa959c3f8a8cf102d15ed6c2c58575680b00fc9794295e90f975bee253
BLAKE2b-256 checksum
How to use checksums
65e928249a41f91ea354391a4c5cd64d0fddd81622f6ec2a62ffed03194b88d8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.7.12

Release files / pggt-1.4.0-py3-none-any.whl

Download URL pggt-1.4.0-py3-none-any.whl
Size 122.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0384bc8b7d6c433a0e74141c93fed6d06a45b872df02877ebf4d77e88abc68f7
BLAKE2b-256 checksum
How to use checksums
ff2cb17361bd2cebb343b876b0cdc9318740a076025bcac72038261008bd6669
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.7.12

Release history Release notifications | RSS feed

This release

1.4.0 This release

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.1.0

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page