Skip to main content

PGGT

PGGT is a recoverable multi-genome workflow that converts pangenome graph projections into allele, copy-number variation (CNV), and presence/absence variation (PAV) components.

It runs five modules:

  1. project annotated genes onto GFA walks and build initial components;
  2. filter graph-supported components with DIAMOND and gene structure evidence;
  3. score singleton and eligible partial genes without removing low-score genes;
  4. regroup singleton genes with reciprocal Miniprot and Pfam evidence;
  5. select representative genes and write allele/CNV/PAV component tables and W-only GFA files.

Genes with a total module-3 score below 1.0 remain in the main component table and are additionally reported in 12.nocoding_gene_scores.tsv.

What's new in 1.1.0

  • Module 1 parses W records with a process pool and writes per-sample segment shards directly. It no longer creates or repeatedly scans the multi-gigabyte sample_segment_all_test.txt table.
  • Per-sample segment shards are sorted concurrently.
  • Module 2 uses FASTA .fai random access instead of loading complete genomes into each Python worker.
  • DIAMOND applies the existing weak-evidence lower bounds during search and streams hits directly into the component-pair filter. The full raw hit table is no longer written to disk.
  • Miniprot mappings use a rolling scheduler: a newly available job slot is filled immediately instead of waiting for an entire batch.
  • 00.run_info/module_timings.tsv records module wall-clock times.

Installation

PGGT requires Linux and Python 3.9 or later.

pip install pggt

The workflow also calls external command-line programs that are not installed by pip:

  • bedtools
  • diamond
  • samtools
  • HMMER (hmmscan and hmmpress)
  • miniprot

For example, these programs can be prepared with Conda/Bioconda:

conda create -n pggt -c conda-forge -c bioconda \
  python=3.10 bedtools diamond samtools hmmer miniprot
conda activate pggt
pip install pggt

Confirm that the command is available:

pggt --help

Quick start

Run input validation first. --dry-run creates no result files.

pggt \
  --gfa-input pangenome.gfa \
  --genome-dir genomes/ \
  --gff-dir annotations/ \
  --samples species_list.txt \
  --expression-dir expression/ \
  --pfam-hmm Pfam-A.hmm \
  --output-dir pggt_out/ \
  --threads 56 \
  --parallel-jobs 4 \
  --partial-na-ratio 0.4 \
  --dry-run

After validation succeeds, run the same command without --dry-run.

The output directory must be new. Do not use a GraphTribe result directory as the PGGT output directory.

Input requirements

Sample list

species_list.txt must contain exactly one sample name per non-empty line and at least two unique samples:

B73
EIL02
EIL31
EIL08
EIL37

Sample names may contain letters, numbers, underscores, dots, and hyphens. The first sample is the reference sample used for representative-gene selection. Every listed sample must occur in the second field of at least one GFA W record.

The same sample names are used to resolve genome, GFF, and expression files. Each sample must resolve to exactly one file of each type.

GFA

--gfa-input accepts either one GFA file or a directory containing exactly one *.gfa file. Use --gfa-name FILE when a directory contains multiple GFA files.

GFA W coordinates are interpreted as 0-based, half-open intervals.

Genome FASTA files

Provide one uncompressed genome FASTA per sample. Recommended names are:

genomes/B73.fa
genomes/EIL02.fa

PGGT first checks <sample>.fa, <sample>.fasta, <sample>.fna, and <sample>.fas. If none exists, it accepts one uniquely matching FASTA whose name contains the sample name.

FASTA sequence names must match the chromosome or sequence names used by the corresponding GFF and GFA walks.

GFF/GFF3 files

Recommended names are:

annotations/B73.renamed.sorted.chr.gff3
annotations/EIL02.renamed.sorted.chr.gff3

PGGT also checks <sample>.gff3 and <sample>.gff, followed by one unique matching GFF file whose name contains the sample name.

The annotations must provide gene, transcript, exon, and CDS relationships. Gene IDs must be consistent with the expression tables. Chromosome names must match the corresponding genome FASTA.

Expression tables

Expression input is required for every sample. Each file must be a tab-separated TSV table. The recommended file name is <sample>.FPKM.tsv; <sample>.expression.tsv is also accepted.

Example:

Gene_ID	leaf_FPKM	root_FPKM
B73_C01g00001	12.50	3.20
B73_C01g00002	0	0.08
B73_C01g00003	0	0

Expression-table rules:

  • The header must contain the exact, case-sensitive column name Gene_ID.
  • At least one expression column must end exactly with _FPKM.
  • Gene IDs must match gene IDs in the corresponding GFF/GFF3.
  • FPKM cells must be numeric. Do not write NA, N/A, or other text.
  • Empty FPKM cells are interpreted as zero.
  • When multiple *_FPKM columns are present, PGGT uses the maximum value for each gene.
  • Use one row per gene. If a gene ID occurs more than once, the later row replaces the earlier value.
  • Genes absent from the expression table receive the no_expression score.

Extra non-FPKM columns are allowed. PGGT first checks <sample>.FPKM.tsv and <sample>.expression.tsv; otherwise, exactly one *<sample>*.tsv file must exist in the expression directory. Ambiguous matches stop the run.

Pfam database

--pfam-hmm must point to a non-empty Pfam-A.hmm file. PGGT prepares the required HMMER index files in its working area when necessary.

Run control

# Resume a previously interrupted PGGT run
pggt [all original arguments] --resume

# Run only one module
pggt [all required arguments] --only-step 3

# Run a module range
pggt [all required arguments] --from-step 2 --to-step 4

Use the same inputs, output directory, thresholds, thread count, and job count when resuming. Completed modules are validated through their completion markers before being skipped.

Threads and shared caches

--threads is the total CPU budget. --parallel-jobs controls concurrent per-sample work, Pfam shards, Miniprot mappings, and PAF parsing, and must not exceed --threads.

For five samples on a sufficiently large 56-core node, --parallel-jobs 5 allows all five Miniprot target mappings to run concurrently. Values above the sample count can still increase Pfam shard concurrency, but should be benchmarked against available memory and storage bandwidth.

Shared caches are stored under ~/.cache/pggt:

  • ~/.cache/pggt/pfam/pfam_cache.sqlite3 caches Pfam results by HMM fingerprint and protein sequence SHA256;
  • ~/.cache/pggt/miniprot/ stores Miniprot genome indexes.

Set the PGGT_CACHE_DIR environment variable to use a different cache root.

Main outputs

pggt_out/
├── 01.graph_mapping/
├── 02.component_filter/
│   └── 10.filtered_component_table.txt
├── 03.singleton_score/
│   ├── 11.scored_singleton_component_table.txt
│   └── 12.nocoding_gene_scores.tsv
├── 04.singleton_cnv/
│   └── 04.orthogroup/
│       └── singleton_orthogroup_component_table.txt
├── 05.allele_CNV_PAV/
│   ├── allele_CNV_PAV_component_table.txt
│   └── W_gfa/
│       └── allele_CNV_PAV.W.gfa
└── 00.run_info/
    └── module_timings.tsv

Each completed module contains artifact_manifest.tsv. Large intermediate DIAMOND, Pfam, and Miniprot files are not published into successful result directories. Failed module workspaces remain under OUTPUT_DIR/.work/ for diagnosis.

Standalone module 5

When a compatible module-4 component table already exists, run module 5 directly:

pggt-module5 \
  --component-table singleton_orthogroup_component_table.txt \
  --samples-file species_list.txt \
  --gff-dir annotations/ \
  --output-dir module5_out/ \
  --representative-seed 1

Command overview

pggt --help
pggt-module5 --help

Maintainer notes

Before a PyPI release, run the regression tests, build both the wheel and source distribution, run twine check dist/*, and verify installation in a clean environment. Update this README whenever command-line options or input requirements change.

Major workflow changes must be developed in a new versioned package directory, with matching updates to pyproject.toml, runtime __version__, this README, PGGT流程原理.md, and CHANGELOG.md. Do not modify a previously published version directory in place.

License

License information will be added before the first public release.

Metadata

Release files for pggt 1.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pggt 1.1.0
File Size Uploaded
pggt-1.1.0.tar.gz 102.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pggt 1.1.0
File Interpreter ABI Platform
pggt-1.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 206.0 kB

Release files / pggt-1.1.0.tar.gz

Download URL pggt-1.1.0.tar.gz
Size 102.7 kB
Tags Source
SHA-256 checksum
How to use checksums
8577ac75ffb9f2f2370e092ebd9130b11f71e08df7bf1622dcac8f15c65bb946
BLAKE2b-256 checksum
How to use checksums
e6f19e4059c888c142fe7e6019058c7c22104133c077adf3c51ebb0f87333fe2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.7.12

Release files / pggt-1.1.0-py3-none-any.whl

Download URL pggt-1.1.0-py3-none-any.whl
Size 103.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
edf80ffa7ab5191e3ea89c66eff604a1c6f0fdc23b9311a2afea883168664c1a
BLAKE2b-256 checksum
How to use checksums
15f3dfd5755ebed018fa4ac1e83c9a1810e5eb5c8b636960d77052ce49e4dab0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.7.12

Release history Release notifications | RSS feed

1.4.0

2 release files

1.3.0

2 release files

1.2.0

2 release files

This release

1.1.0 This release

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page