Pangenomium
Scalable pangenomics toolkit for clustering genomes and proteins across large datasets.
Installation
mamba create -n pangenomium -c conda-forge -c bioconda skani mmseqs2 setuptools 'python>=3.9' -y
mamba activate pangenomium
pip install pangenomium
Dependencies
Citations
The methodology used for dereplicating genomes into pangenomes and proteins into orthologs was initially published in VEBA 2.0 which uses skani for pairwise ANI and MMseqs2 for protein clustering.
-
Espinoza JL, Phillips A, Prentice MB, Tan GS, Kamath PL, Lloyd KG, Dupont CL. Unveiling the microbial realm with VEBA 2.0: a modular bioinformatics suite for end-to-end genome-resolved prokaryotic, (micro)eukaryotic and viral multi-omics from either short- or long-read sequencing. Nucleic Acids Res. 2024 Aug 12;52(14):e63. doi: 10.1093/nar/gkae528.
-
Shaw J, Yu YW. Fast and robust metagenomic sequence comparison through sparse chaining with skani. Nat Methods. 2023 Nov;20(11):1661-1665. doi: 10.1038/s41592-023-02018-3.
-
Steinegger, M., Söding, J. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat Biotechnol 35, 1026–1028 (2017). https://doi.org/10.1038/nbt.3988
Quick Start
# End-to-end workflow
pangenomium end-to-end \
-i genomes_table.tsv \
--mode veba \
-o output_directory \
--n_threads_skani 8 \
--n_concurrent_mmseqs_tasks 4
# Or run individual steps
pangenomium cluster-genomes -i genome_list.txt -o genome_clustering
pangenomium cluster-proteins-from-pangenomes \
-g genomes_table.tsv \
-c genome_clustering/output/genomes_to_pangenomes.tsv.gz \
-o protein_clustering
Commands
cluster-genomes
Cluster genomes by ANI using skani.
pangenomium cluster-genomes \
-i INPUT \
-o OUTPUT_DIR \
--n_threads 8 \
--ani_threshold 95.0
Input formats:
- Simple list: One genome path per line
- Batch format: TSV with columns
[organism_type, id_genome, genome_filepath, protein_filepath] - VEBA format: TSV with columns
[organism_type, id_sample, id_genome, genome_filepath, protein_filepath, cds_filepath, gff_filepath]
Outputs:
genomes_to_pangenomes.tsv.gz- Genome to pangenome cluster assignmentsserialization/genome_clusters.graph.pkl.gz- NetworkX graph of genome relationshipsserialization/genome_clusters.dict.pkl.gz- Dictionary mapping genomes to clustersrepresentatives/genome_representatives.tsv.gz- Representative genome for each cluster
cluster-proteins
Cluster all proteins using MMseqs2.
pangenomium cluster-proteins \
-i proteins.faa \
-o OUTPUT_DIR \
--n_threads 8 \
--minimum_identity_threshold 50.0
Outputs:
protein_clusters.tsv.gz- Protein to orthogroup assignmentsserialization/protein_clusters.graph.pkl.gz- NetworkX graphserialization/protein_clusters.dict.pkl.gz- Dictionary mappingrepresentatives/protein_representatives.tsv.gz- Representative sequences
cluster-proteins-from-pangenomes
Cluster proteins within each pangenome independently using MMseqs2.
pangenomium cluster-proteins-from-pangenomes \
-g genomes_table.tsv \
-c genomes_to_pangenomes.tsv.gz \
-o OUTPUT_DIR \
--n_threads_per_task 2 \
--n_concurrent_tasks 4
Alternative input: Direct pangenome manifest with -i instead of -g and -c.
Outputs:
output/proteins_to_orthologs.tsv.gz- All protein to orthogroup assignmentsoutput/pangenome_tables/- Per-pangenome prevalence matrices (genome × orthogroup)
end-to-end
Complete workflow: genome clustering, protein clustering, and comprehensive output generation.
pangenomium end-to-end \
-i genomes_table.tsv \
--mode veba \
-o OUTPUT_DIR \
--n_threads_skani 8 \
--n_concurrent_mmseqs_tasks 4
Modes:
batch- Expects 4-column format:[organism_type, id_genome, genome_filepath, protein_filepath]veba- Expects 7-column VEBA format
Outputs:
Genome clustering results:
genome_clustering/output/genomes_to_pangenomes.tsv.gzgenome_clustering/output/serialization/- Graph and dictionary objectsgenome_clustering/output/representatives/- Representative genomes
Protein clustering results:
protein_clustering/output/proteins_to_orthologs.tsv.gzprotein_clustering/output/pangenome_tables/- Per-pangenome prevalence tables
Comprehensive outputs in output/:
genome_clusters.tsv.gz- Genome metadata with cluster assignmentsprotein_clusters.tsv.gz- Protein metadata with orthogroup assignmentsidentifiers.tsv.gz- All identifier mappingspangenome_tables/- Prevalence matrices for each pangenomerepresentatives.faa- Representative protein sequencescore_pangenomes/- Core pangenome sequences per clusterserialization/- Graph and dictionary objects
compile-pangenomes-table
Utility to create pangenome manifest from genome manifest and cluster assignments.
pangenomium compile-pangenomes-table \
-i genomes_table.tsv \
-c genomes_to_pangenomes.tsv.gz \
-o pangenomes_manifest.tsv
Input Format Examples
Simple genome list
/path/to/genome_001.fasta
/path/to/genome_002.fasta
/path/to/genome_003.fasta
Batch manifest (4 columns)
prokaryote genome_001 /path/to/genome_001.fasta /path/to/proteins_001.faa
prokaryote genome_002 /path/to/genome_002.fasta /path/to/proteins_002.faa
eukaryote genome_003 /path/to/genome_003.fasta /path/to/proteins_003.faa
VEBA manifest (7 columns)
prokaryote sample_A genome_001 /path/to/genome_001.fasta /path/to/proteins_001.faa /path/to/cds_001.ffn /path/to/annotation_001.gff
prokaryote sample_A genome_002 /path/to/genome_002.fasta /path/to/proteins_002.faa /path/to/cds_002.ffn /path/to/annotation_002.gff
prokaryote sample_B genome_003 /path/to/genome_003.fasta /path/to/proteins_003.faa /path/to/cds_003.ffn /path/to/annotation_003.gff
Pangenome manifest (3 columns)
genome_001 PanG-abc123 /path/to/proteins_001.faa
genome_002 PanG-abc123 /path/to/proteins_002.faa
genome_003 PanG-def456 /path/to/proteins_003.faa
Output File Descriptions
Cluster assignment files
- genomes_to_pangenomes.tsv.gz - Two columns:
[id_genome, id_pangenome] - proteins_to_orthologs.tsv.gz - Two columns:
[id_protein, id_orthogroup]
Comprehensive output tables
- genome_clusters.tsv.gz - Genome metadata: ID, organism type, cluster, paths, contig count, total length
- protein_clusters.tsv.gz - Protein metadata: ID, genome, contig, cluster, length, product
- identifiers.tsv.gz - All identifier mappings: genome, sample, pangenome, contig, protein, orthogroup
Prevalence tables
Per-pangenome matrices in pangenome_tables/:
- Rows: Genomes
- Columns: Orthogroups
- Values: Number of proteins from genome in orthogroup
Representative sequences
- representatives.faa - Representative protein for each orthogroup
- core_pangenomes/*.faa - Core proteins for each pangenome cluster
Serialization objects
- genome_clusters.graph.pkl.gz - NetworkX graph of genome ANI relationships
- genome_clusters.dict.pkl.gz - Dictionary: genome → pangenome cluster
- protein_clusters.graph.pkl.gz - NetworkX graph of protein similarity
- protein_clusters.dict.pkl.gz - Dictionary: protein → orthogroup
Algorithm Details
Genome clustering:
- Uses skani for fast ANI calculation
- Clusters with edgelist-to-clusters.py using relaxed mode (single-linkage)
- Default: 95% ANI, 50% alignment fraction
Protein clustering:
- Uses MMseqs2 easy-cluster
- Per-pangenome clustering prevents inflation of orthogroup sizes
- Default: 50% identity, 80% coverage (bidirectional)
Parallelization:
- Genome clustering: Multi-threaded skani
- Protein clustering: Multiple concurrent MMseqs2 jobs, each multi-threaded
Common Workflows
From VEBA output
pangenomium end-to-end \
-i veba_output/genomes_table.tsv \
--mode veba \
-o pangenome_analysis \
--n_threads_skani 16 \
--n_concurrent_mmseqs_tasks 8 \
--n_threads_mmseqs_per_task 4
Manual two-step workflow
# Step 1: Cluster genomes
pangenomium cluster-genomes \
-i genomes_table.tsv \
-o step1_genomes \
--n_threads 16
# Step 2: Cluster proteins per pangenome
pangenomium cluster-proteins-from-pangenomes \
-g genomes_table.tsv \
-c step1_genomes/output/genomes_to_pangenomes.tsv.gz \
-o step2_proteins \
--n_concurrent_tasks 8 \
--n_threads_per_task 4
Protein clustering without genome context
# Concatenate all proteins
cat proteins/*.faa > all_proteins.faa
# Cluster all proteins together
pangenomium cluster-proteins \
-i all_proteins.faa \
-o protein_clusters \
--n_threads 16
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
File details
Details for the file pangenomium-2026.5.19.tar.gz.
File metadata
- Download URL: pangenomium-2026.5.19.tar.gz
- Upload date:
- Size: 37.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.14.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1bff7f7e7e5ac79c5a130d8c3ecaf0256b53f7573ba17b5a3ef28f90f0966e55
|
|
| MD5 |
ca39b89bcbb673d4ce4a6c88aa9a2a5f
|
|
| BLAKE2b-256 |
425b2ab18998f710ebe7864fea60e36efa621d898cbef9ec09cf3d36de7ed861
|