Skip to main content

Progressive Genome Segment Enhancement (PGSE)

Overview

PGSE is an algorithm for predicting phenotypes from whole genome sequencing (WGS) data. It was intiially developed for the prediction of antimicrobial minimum inhibitory concentration (MIC) in bacterial strains. PGSE has higher accuracy, lower memory consumption, and shorter runtime compared to traditional $k$-mer based XGBoost models. PGSE is also able to run on distributed systems.

Contributors

Dr Yinzheng (William) Zhong, Univerisity of Liverpool (algorithm design & implementation)

Dr Alessandro Gerada, University of Liverpool (conceptualisation, R package, funding)

Prof William Hope, University of Liverpool (conceptualisation, funding, supervision)

Citation

@article{gerada2026prediction,
  title={Prediction of antimicrobial minimum inhibitory concentration from bacterial genomes using a scalable and interpretable machine learning approach},
  author={Gerada, Alessandro and Zhong, Yinzheng and Harper, Nicholas and Velluva, Anoop and Reza, Nada and Dubey, Vineet and Howard, Alex and Green, Peter L and Paterson, Steve and Hope, William},
  journal={npj Antimicrobials and Resistance},
  year={2026},
  publisher={Nature Publishing Group}
}

License

This project is licensed under the GNU Affero General Public License v3.0 (AGPL-3.0). See the LICENSE file for details.

Installation

PyPI

Make sure Python 3.10 or later is installed, then install pgse from PyPI with uv:

uv pip install pgse

or with pip:

pip install pgse

The published wheels bundle the compiled native counting kernel (a Rust extension) for Linux, macOS, and Windows. Segment counting runs through this kernel; if it is ever unavailable (e.g. a source install without a Rust toolchain) PGSE falls back to a slower pure-Python counter.

Conda

To use in a conda environment:

conda create -n pgse python=3.11
conda activate pgse
python -m pip install pgse

pgse is now available to import.

R

To use PGSE through R, install the package in an R session using:

install.packages("devtools")
devtools::install_github("yinzheng-zhong/PGSE", subdir = "R-package")

Usage

Training

Single node/machine

Import the pipeline from the package and run the pipeline like this. You can use your own argument parser or use the one provided by pgse. Also, you can instantiate the pipeline with a wrapper that provides the parameters directly.

# You can use your own argument parser or use the one provided by pgse.
# Or instantiate the pipeline with a wrapper that provides the parameters directly.
from pgse.environment.args import get_parser
from pgse import TrainingPipeline

if __name__ == "__main__":
  parser = get_parser()
  args = parser.parse_args()

  pipeline = TrainingPipeline(
    args.data_dir,
    args.label_file,
    args.pre_kfold_info_file,
    args.save_file,
    args.export_file,
    args.k,
    args.ext,
    args.target,
    args.features,
    args.folds,
    args.ea_min,
    args.ea_max,
    args.num_rounds,
    args.lr,
    args.dist,
    args.nodes,
    args.workers
  )

  pipeline.run()

Alternatively, to run PGSE as a standalone program on a local machine, install the package and use the following command as an example:

pgse-train \
        --label-file "../<path_to>/<you_labels>.csv" \
        --data-dir "../<you_data_dir>/" \
        --pre-kfold-info-file "../<k_fold_information>.json" \
        --save-file "../<saved progress>.save" \
        --export-file "../<exported files>" \
        --workers 8 \
        --features 10000 \
        --partition-size-target 5000 \
        --dist 0 \
        --k 6 \
        --target 70 \
        --ext 2 \
        --lr 0.001 \
        --num-rounds 6000 \
        --folds 5 \
        --ea-max 64 \
        --ea-min 0
  • --label-file (Required): path to the .csv label file

    Here the label file is a csv file with the following format:

    | labels | files     |
    | ------ | --------- |
    | 7      | file1.fna |
    | 7      | file2.fna |
    | 6      | file3.fna |
    

    The labels are the target values for the prediction task. The files are the file names (.fna files under --data-dir) containing the genome sequences.

  • --data-dir (Required): path to the data directory containing the .fna files. PGSE will be able to retrieve the genome sequences using this path and the file names in the label file.

  • --pre-kfold-info-file: path to the predefined k-fold info JSON file. This is not required but will be useful if you want to compare PGSE with other systems. Without this, PGSE will split the data into k folds randomly using a fixed seed. E.g.

    {
        "fold_0": [
            "Sample_208-MOLMIC_E33.scaffolds.fna",
            "Sample_726-MOLMIC_F29.scaffolds.fna",
            "Sample_474-MOLMIC_I14.scaffolds.fna",
            "Sample_111-MOLMIC_C61.scaffolds.fna",
            "Sample_087-MOLMIC_C25.scaffolds.fna",
            "Sample_467-MOLMIC_I6.scaffolds.fna",
            "..."
        ],
        "fold_1": [
            "Sample_208-MOLMIC_E33.scaffolds.fna",
            "Sample_726-MOLMIC_F29.scaffolds.fna",
            "Sample_474-MOLMIC_I14.scaffolds.fna",
            "Sample_111-MOLMIC_C61.scaffolds.fna",
            "Sample_087-MOLMIC_C25.scaffolds.fna",
            "Sample_467-MOLMIC_I6.scaffolds.fna",
            "..."
        ],
        "...": [
        "..."
        ]
    }
    
  • --save-file: file to save the progress. This is useful if you want to resume the training process.

  • --export-file: file to export the results. Normally without an extension. This name will be used to store the selected genome segments in an .txt file and the trained model in a .json file.

  • --workers: number of workers per node.

  • --features: Maximum number of features to keep after the feature importance calculation and ranking.

  • --partition-size-target: Target number of features in each XGBoost partition during feature selection (default 5000). Partitions are evenly sized, so the actual size lands between this and twice this. See Why do we perform feature partitioning? below. Larger partitions use more memory per worker and give each model a wider view of the features; 0 or less trains a single model over all features.

  • --dist: Using distributed computation or not. 0 for running on a single node/machine, 1 for running on multiple nodes.

  • --k: initial k-mer size.

  • --target: Maximum segment length to extend to.

  • --ext: Extension length in each round. Extension parameter p from the paper.

  • --lr: learning rate.

  • --num-rounds: Maximum rounds for the training process.

  • --folds: Number of folds for the k-fold cross-validation.

  • --metric: Validation metric reported for the held-out fold (default essential_agreement). See Validation metrics below.

  • --ea-max: Maximum number of censored essential agreement values. Don't need this unless you want to see more accurate EA information from the console output during the training. Only read by the essential_agreement metric.

  • --ea-min: Minimum number of censored essential agreement values. Similar to --ea-max.

  • --alphabet: the set of characters the input is made of, given as a single string. Defaults to DNA (atgc). See Alphabets below.

  • --case-sensitive: 0 (default) folds everything to lower case, 1 treats upper and lower case as distinct characters.

  • --complement: the complement used to canonicalise segments, one character per character of --alphabet. See Alphabets below.

  • --uint16: 0 (default) stores the segment-count matrix as float32; 1 stores it as uint16, halving the memory. Lossless for counts up to 65535 (larger counts are saturated). See Reducing memory usage below.

  • --sparse: 0 (default) stores the count matrix densely; 1 stores it as a sparse CSR matrix. For short, sparse inputs (e.g. SMILES strings) the matrix is almost entirely zeros, so this can save orders of magnitude. XGBoost reads the unstored zeros of a CSR matrix as missing values (not as 0), so the same value must be used at prediction time. See Reducing memory usage below.

Alphabets

PGSE is not limited to DNA. The alphabet is the set of characters the input is made of, and everything outside it is dropped while reading the input files. Any symbolic text can therefore be used, for example plain text:

pgse-train \
        --label-file "../<path_to>/<you_labels>.csv" \
        --data-dir "../<you_data_dir>/" \
        --alphabet "abcdefghijklmnopqrstuvwxyz " \
        --case-sensitive 0 \
        --k 3 \
        --target 12

The same options are available on the Python API:

from pgse import TrainingPipeline

pipeline = TrainingPipeline(
    data_dir='...',
    label_file='...',
    k=3,
    target=12,
    alphabet='abcdefghijklmnopqrstuvwxyz ',
    case_sensitive=False
)

Three things follow from the alphabet:

  • Case sensitivity. By default the alphabet is case-insensitive: the input is folded to lower case, so A and a are the same character. With --case-sensitive 1 nothing is folded and the alphabet has to list every character you want to keep, e.g. "abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ ".
  • Canonicalisation. DNA segments are counted on both strands, so a segment and its reverse complement are the same feature. This only makes sense when the alphabet has a complement. --complement defaults to reverse complementing for the DNA alphabet and to no canonicalisation for any other alphabet, which is almost always what you want. To set one explicitly, give one character per character of --alphabet (--alphabet augc --complement uacg for RNA), or pass an empty string to switch it off.
  • k-mer size. The number of initial k-mers is len(alphabet) ** k, so a larger alphabet needs a smaller --k. The 27-character alphabet above yields about 20k 3-mers, roughly the same as DNA with --k 7.

Inference has to use the same alphabet as training, since the exported segments are counted against it. Pass the same --alphabet, --case-sensitive and --complement values to pgse-predict; if the segments do not fit the alphabet, PGSE fails with an error rather than predicting on all-zero counts.

Validation metrics

The score reported for the held-out fold is chosen with --metric (or metric= on the Python API). Essential agreement is the default and keeps the behaviour of earlier versions, but it only makes sense for log2 MIC labels; anything else should pick a metric that matches its label.

--metric What it measures Best for
essential_agreement Fraction of predictions within one two-fold dilution of the label log2 MIC labels (default)
rmse Root mean squared error, in the units of the label Continuous labels, penalising large misses
mae Mean absolute error Continuous labels with outliers
mape Mean absolute percentage error over the non-zero labels Continuous labels compared in relative terms
r2 Fraction of the label's variance explained Continuous labels, e.g. growth rate
pearson Linear correlation between label and prediction Continuous labels, e.g. growth rate
spearman Rank correlation Continuous labels where only the ordering matters
accuracy Fraction of predictions that round to the exact label Integer labels, e.g. 0/1
mcc Matthews correlation coefficient, in [-1, 1] Imbalanced integer labels

For a continuous target such as growth rate, r2 is the usual headline number, with rmse for the error in the target's own units and spearman when only the ranking of the samples matters:

pgse-train \
        --label-file "../<path_to>/<your_labels>.csv" \
        --data-dir "../<your_data_dir>/" \
        --metric r2
from pgse import TrainingPipeline

pipeline = TrainingPipeline(data_dir='...', label_file='...', metric='r2')

PGSE always trains on squared error and always logs RMSE; --metric selects the extra score reported alongside it once the selected segments are trained.

PGSE trains a regressor (reg:squarederror) whatever the metric is — there is no classification objective. accuracy and mcc are there for labels that are whole numbers, such as a 0/1 target regressed and then read back as a class; both round the predictions to the nearest integer before scoring.

The metrics are the static methods of the Metric class in pgse/validation/metrics.py, with the array handling they share in pgse/validation/utils.py. Adding a metric is adding a method — its name becomes the --metric value and the first line of its docstring becomes its entry in --metric's help:

class Metric:
    ...

    @staticmethod
    def max_error(y_true, y_pred):
        """Largest absolute error over the samples.

        Args:
            y_true: True labels.
            y_pred: Predicted values.
        """
        return float(np.max(np.abs(as_array(y_true) - as_array(y_pred))))

A metric that needs extra parameters just declares them as keyword arguments. The pipeline passes all of its validation options to every metric and each receives only the ones its signature names, which is how essential_agreement gets ea_min and ea_max while the others ignore them.

Reducing memory usage

The segment-count matrix (one row per sample, one column per segment) is the largest object PGSE holds in memory. Two optional flags shrink it; they are independent and can be combined. Both default to off, so existing runs are unaffected.

  • --uint16 1 stores counts as 16-bit integers instead of 32-bit floats, halving the matrix. Counts are non-negative integers, so this is lossless up to 65535; any larger count is saturated to 65535.
  • --sparse 1 stores the matrix as a sparse CSR matrix. This is the big lever for short inputs where most segments are absent from most samples — for example SMILES strings, where each row has only tens of non-zero counts out of thousands or millions of columns, making the matrix >99% zeros. For long, dense inputs such as bacterial genomes the matrix is not sparse, so leave this off.
pgse-train \
        --label-file "../<path_to>/<your_labels>.csv" \
        --data-dir "../<your_data_dir>/" \
        --alphabet "..." \
        --sparse 1 \
        --uint16 1

The same options are available on the Python API as sparse=True and uint16=True:

from pgse import TrainingPipeline

pipeline = TrainingPipeline(data_dir='...', label_file='...', sparse=True, uint16=True)

Important: --sparse must match between training and prediction. In a sparse matrix, unstored zeros are read by XGBoost as missing values, whereas in a dense matrix a count of zero is an explicit 0. Training dense and predicting sparse (or vice versa) shifts the predictions. --uint16 has no such constraint. Pass the same values to pgse-predict (see below).

Distributed computation

To run PGSE on a distributed system, you need to use your environment specific setup. There are multiple examples about running PGSE using Slurm under the slurm-scripts directory.

  • job-pgse-array.sh: Run PGSE on a cluster using Slurm with multiple nodes for multiple antibiotics using array jobs. Here -dist is set to 0 as each task is running separately.
  • job-pgse-dist.sh: Run PGSE on a cluster using Slurm with multiple nodes for a single antibiotic. Here -dist is set to 1 as the task is running on different nodes.
  • job-pgse-single.sh: Run PGSE on a Slurm cluster with a single node for a single antibiotic. Here -dist is set to 0.

Inferencing

An example of how this can be done is provided in main-pgse-inf.py.

from pgse import InferencePipeline

MODEL_PATH = '../volatile/var/result-k6-CAZ-perf_fold_0.json'
SEGMENT_PATH = '../volatile/var/result-k6-CAZ-perf_fold_0.csv'

if __name__ == "__main__":
    # Instantiate the pipeline
    pipeline = InferencePipeline(MODEL_PATH, SEGMENT_PATH, workers=8)

    # files as a list of paths to the fasta files
    EG_1 = [
        '../volatile/cgr/Sample_002-MOLMIC_B2.scaffolds.fna',
        '../volatile/cgr/Sample_394-MOLMIC_H8.scaffolds.fna',
        '../volatile/cgr/Sample_385-MOLMIC_G79.scaffolds.fna',
        '../volatile/cgr/Sample_622-MOLMIC_K68.scaffolds.fna',
        '../volatile/cgr/Sample_252-MOLMIC_F2.scaffolds.fna',
        '../volatile/cgr/Sample_208-MOLMIC_E33.scaffolds.fna',
        '../volatile/cgr/Sample_443-MOLMIC_H62.scaffolds.fna',
        '../volatile/cgr/Sample_565-MOLMIC_J66.scaffolds.fna',
        '../volatile/cgr/Sample_339-MOLMIC_G29.scaffolds.fna',
        '../volatile/cgr/Sample_418-MOLMIC_H33.scaffolds.fna',
    ]

    result_1 = pipeline.run(EG_1)
    print(result_1)

    EG_2 = [
        '../volatile/cgr/Sample_394-MOLMIC_H8.scaffolds.fna',
        '../volatile/cgr/Sample_385-MOLMIC_G79.scaffolds.fna',
        '../volatile/cgr/Sample_622-MOLMIC_K68.scaffolds.fna',
        '../volatile/cgr/Sample_252-MOLMIC_F2.scaffolds.fna'
    ]

    result_2 = pipeline.run(EG_2)
    print(result_2)

To run the inference pipeline as a standalone program, install the package and use the following command as an example:

pgse-predict \
        --model-file "../<path_to_model>.json" \
        --segment-file "../<path_to_segment>.csv" \
        --data-dir "../<you_data_dir>/" \
        --workers 8

If the model was trained on a non-DNA alphabet, pass the same --alphabet, --case-sensitive and --complement values that were used for training. See Alphabets. If training used --sparse 1, prediction must use it too — see Reducing memory usage.

### R package

To use PGSE through the R package, consult the package
[documentation](https://github.com/yinzheng-zhong/PGSE/tree/main/R-package/).

## For Development

PGSE uses [uv](https://docs.astral.sh/uv/) for packaging and dependency
management. All project metadata and dependencies live in `pyproject.toml`.

**Prerequisite: a Rust toolchain.** Building from source compiles the native
counting kernel (the Rust crate under `native/`, built into the package as
`pgse._native`), so you need `cargo`/`rustc` on your `PATH`. Install them from
[rustup.rs](https://rustup.rs):
```bash
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh
# rustup only updates PATH for *new* shells, so load it into the current one
# (or open a new shell) before installing:
. "$HOME/.cargo/env"

The build fails loudly if the toolchain is missing rather than silently shipping an install without the extension, so every install has the fast path. (A pure-Python counter still exists as a runtime safety net, and logs a warning if it is ever used.)

Clone the repository and create a synced environment. This installs the runtime and dev dependencies and performs an editable install of pgse, compiling the Rust kernel as part of it:

uv sync

To run an editable install on its own:

uv pip install -e .

To build the distribution artifacts (the wheel bundles the compiled kernel):

uv build

Run the tests with:

uv run pytest

Releases to PyPI are automated by GitHub Actions: when the version in pyproject.toml changes on main, the workflow uses cibuildwheel to build one abi3 wheel per platform (Linux, macOS, Windows; each valid for CPython 3.10+), builds an sdist, and publishes them.

Acknowledgements

This work was funded, in part, by UKRI and the Wellcome trust.

This work was undertaken on Barkla, part of the High Performance Computing facilities at the Univeristy of Liverpool, UK.

Common Issues

XGBoost training is only using one core.

Some linux distributions need an environment variable OMP_NUM_THREADS=<num threads> to be set to allow XGBoost to use multiple cores.

Q & A

Why do we perform feature partitioning?

There are four reasons why feature partitioning is crucial in PGSE. First, feature partitioning is used as a memory reduction technique. The model is trained on a subset of the features at a time, therefore, the memory consumption is reduced while maintained a relatively stable RAM usage regardless of the number of total features. Second, feature partitioning helps to parallelise the training process. Each partition can be trained on a different worker across different nodes. This is particularly useful as XGBoost training consumes most of the time in the training process. Third, from the experiments we have conducted, we found that feature dimensionality affects the model's optimal hyperparameters. For example, higher feature dimensionality requires a shallower tree depth in general. PGSE is a dynamic system that and the total number of features can be different in each round. Therefore, partitioning the features into similarly-sized sub-features can help to minimise the impact of the feature dimensionality on the model's hyperparameters. Finally, feature partitioning helps to preserve the feature importance information from XGBoost. Likely due to the pruning process, more feature importance information will be lost (become 0) if the dimensionality increases.

The size of each partition is set with --partition-size-target (or partition_size_target= on TrainingPipeline), which defaults to 5000 features.

Why do we eliminate features?

If segment A is extended into segment B, A becomes a subsequences of B. For pairs like A and B, we only need to keep the ones with higher feature importance. Extension and elimination are two crucial parts of the PGSE system, which grows the genome segments longer and the elimination process guarantees that the growth will stop eventually. Additionally, elimination guarantees the convergence of the system as the feature dimensionality will start decreasing at some point till all features stop growing.

Release files for pgse 0.11.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pgse 0.11.0
File Size Uploaded
pgse-0.11.0.tar.gz 72.5 kB Details

Built distributions (wheels)

Table of built distributions (wheels) for pgse 0.11.0
File Interpreter ABI Platform
pgse-0.11.0-cp310-abi3-win_amd64.whl CPython 3.10 abi3 Windows x86-64 Details
pgse-0.11.0-cp310-abi3-manylinux_2_28_x86_64.whl CPython 3.10 abi3 Linux glibc 2.28+ x86-64 Details
pgse-0.11.0-cp310-abi3-macosx_11_0_arm64.whl CPython 3.10 abi3 macOS 11.0+ ARM64 Details

Total release size: 1.3 MB

Release files / pgse-0.11.0.tar.gz

Download URL pgse-0.11.0.tar.gz
Size 72.5 kB
Tags Source
SHA-256 checksum
How to use checksums
7f73763ef57b76f1f9f1ba230f121ce9bde7b8cbc29ebcd42fc0ed4e3089042a
BLAKE2b-256 checksum
How to use checksums
70fd11b897fd2cf55e7e5e7f1864a0f7c8b1949a9511130b9db4428ec9766c3d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.20

Release files / pgse-0.11.0-cp310-abi3-win_amd64.whl

Download URL pgse-0.11.0-cp310-abi3-win_amd64.whl
Size 342.1 kB
Tags CPython 3.10 Windows x86-64 abi3
SHA-256 checksum
How to use checksums
dcdb5b81c6eec12ec600f87fbe7642765f93456ba363191524addd53b5c3b600
BLAKE2b-256 checksum
How to use checksums
0e5905d876c9f17414f8b9a3861cf14e2320b63d93da411f754beb3fa67a06d8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.20

Release files / pgse-0.11.0-cp310-abi3-manylinux_2_28_x86_64.whl

Download URL pgse-0.11.0-cp310-abi3-manylinux_2_28_x86_64.whl
Size 472.4 kB
Tags CPython 3.10 Linux glibc 2.28+ x86-64 abi3
SHA-256 checksum
How to use checksums
1a8e3738cf99b33fc1cde1bb2a2e9134812ae4a7310aede84150413d402be08c
BLAKE2b-256 checksum
How to use checksums
a6acde5b25131c80100c711c39d6fede9fdc8afad8b6b3723f907c1086969587
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.20

Release files / pgse-0.11.0-cp310-abi3-macosx_11_0_arm64.whl

Download URL pgse-0.11.0-cp310-abi3-macosx_11_0_arm64.whl
Size 399.6 kB
Tags CPython 3.10 abi3 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
2ed1f0f7c982f05d3e08f8d198a06a9839a651900640a4a2e9cd90f021a88744
BLAKE2b-256 checksum
How to use checksums
425da2948e8176787034fc8969a6821c01468289985881a481dda1aa0bc600f8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.20

Release history Release notifications | RSS feed

This release

0.11.0 This release

4 release files

0.10.4

4 release files

0.9.1

1 release file

0.9.0

1 release file

0.8.5

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page