A pipeline for plasmid outbreak clustering and visualization.
Project description
Roundabout
A pipeline for plasmid outbreak clustering and visualization.
Tracking the spread of antimicrobial resistance (AMR) and virulence factors during outbreaks is a massive public health challenge. While traditional genomic epidemiology excels at tracing bacterial clones via chromosomal mutations (like SNP calling), it fundamentally struggles to track the highly mobile plasmids responsible for horizontal gene transfer (HGT). Because plasmids frequently recombine, fuse, and transfer across completely different bacterial species, they are notoriously difficult to properly cluster, align, and analyze.
Roundabout solves this problem by providing an automated, end-to-end pipeline tailored specifically for mobile genetic elements. It takes plasmid assemblies and annotates their features, clusters them based on structural similarity and replicon profiles, compares them to RefSeq plamid databases, and generates comparative visualizations.
Quickstart
Before running the full annotation pipeline for the first time, download and configure the required databases.
# Download and configure all default databases
roundabout --setup-db
Once the databases are configured, run the pipeline by pointing it to a directory containing your plasmid FASTA files (.fa, .fasta, or .fna).
# Run the full pipeline using 8 threads
roundabout -f path/to/fasta_directory -o roundabout_results -t 8
Cluster and visualize plasmids based solely on structural similarity:
# Run structural clustering and local visualizations only (No databases)
roundabout -f path/to/fasta_directory -o roundabout_results --skip-db
A full list of cli options can be accessed via standard help flags (-h, --help):
# A full list of cli options can be found with the "-h" flag
roundabout -h
Installation
Because Roundabout relies on heavily optimized, non-Python bioinformatics binaries (like BLAST, KMA, and Skani), using Conda (or Mamba) to manage the environment is strongly recommended.
Option 1: Bioconda (Recommended)
The easiest way to install Roundabout and all of its required external dependencies is through the Bioconda channel. This creates an isolated environment and installs everything in one step.
# Create a new environment and install roundabout
conda create -n roundabout-env -c conda-forge -c bioconda roundabout
# Activate the environment
conda activate roundabout-env
Option 2: Install from Source (Development)
Roundabout can be installed from source by cloning the repository.
# Clone the repository
git clone https://github.com/erinyoung/roundabout.git
cd roundabout
# Create the environment from the YAML file
conda env create -f environment.yml
# Activate the environment
conda activate roundabout-env
Post-Installation Database Setup
Regardless of how Roundabout was installed, download the necessary reference databases before running. These databases are for Bakta, Plasmidfinder, refseq-plasmid-dl, and AMRFinder can be very large, but also very useful.
roundabout --setup-db
Dependencies
graph TD
%% Styling
classDef input fill:#f9f,stroke:#333,stroke-width:2px;
classDef step fill:#bbf,stroke:#333,stroke-width:1px;
classDef database fill:#fcf,stroke:#333,stroke-width:1px,stroke-dasharray: 5 5;
classDef output fill:#bfb,stroke:#333,stroke-width:2px;
%% Nodes
A[Plasmid FASTAs]:::input
B[stage_and_split_fastas]:::step
C[staging_fastas/]:::output
subgraph "Parallel Annotation Loop"
D1[AMRFinderPlus]:::step
D2[PlasmidFinder]:::step
D3[Bakta]:::step
end
subgraph "External Reference Integration"
E1[refseq-plasmid-dl]:::step
E2[NCBI Master RefSeq Multi-FASTA]:::database
end
F[Skani Alignment Engine]:::step
subgraph "Clustering & Grouping Engine"
G1[define_groups_by_similarity]:::step
G2[define_groups_by_amr]:::step
G3[define_groups_by_pf]:::step
H[run_grouping_summary]:::step
end
subgraph "Cohort Isolated Visualizations"
I1[PyGenomeViz Synteny Plots]:::step
I2[MinkeMap Circular Alignment]:::step
I3[DaisyBlast Sequence Shattering]:::step
I4[Cohort Matrix Heatmaps]:::step
end
O[Final Results Directory]:::output
%% Edges
A --> B
B --> C
C --> D1
C --> D2
C --> D3
C --> F
E1 -->|Download & Filter| E2
E2 -->| extract_refseq_fastas | F
D1 -->|Profiles| G2
D2 -->|Profiles| G3
F -->|Raw ANI Matrix| G1
F -->|Global Hits DataFrame| H
G1 --> H
G2 --> H
G3 --> H
H -->|Deduplicated Cohorts| I1
H -->|Deduplicated Cohorts| I2
H -->|Deduplicated Cohorts| I3
H -->|Deduplicated Cohorts| I4
D3 -->|Generated .gbff Files| I1
C -->|Fasta Map Lookup| I2
C -->|Fasta Map Lookup| I3
F -->|Local Matrix DataFrame| I4
I1 --> O
I2 --> O
I3 --> O
I4 --> O
class A,C,O input;
Roundabout integrates several open-source bioinformatics tools to handle everything from annotation to alignment and visualization.
Core Annotation & Clustering
- Bakta: Rapid and standardized structural and functional annotation of plasmid sequences.
- AMRFinderPlus: Detection of antimicrobial resistance (AMR) genes, stress responses, and virulence factors.
- PlasmidFinder: Identification of plasmid replicon types (Incompatibility/Inc groups) to establish plasmid lineages.
- Skani: High-speed calculation of Average Nucleotide Identity (ANI) and alignment fractions for defining structural similarity cohorts.
- refseq-plasmid-dl: Downloading, filtering, and curating circular reference plasmids from NCBI RefSeq to provide global outbreak context.
Visualization Ecosystem
- pyGenomeViz: Generating linear synteny plots and comparing annotated genomic regions across plasmid cohorts.
- MinkeMap: Creating detailed circular alignment plots to visualize coverage and structural conservation against reference sequences.
- DaisyBlast: Performing self-BLAST sequence shattering, feature grouping, and generating structural dotplots.
Alignment Engines & Backend
- BLAST, MUMmer, MMseqs2, & progressiveMauve: The underlying computational alignment engines powering the structural synteny visualizations and sequence shattering.
- KMA: The high-speed k-mer alignment engine required for indexing and searching the PlasmidFinder database.
- Pandas & SciPy: The data manipulation and scientific computing backends used for matrix generation, parsing outputs, and mathematical threshold filtering.
Testing the Pipeline
Roundabout comes with a set of sample FASTA files to verify that the pipeline and its dependencies are functioning correctly. These files are located in the test/ directory.
Use the --skip-db flag to run a quick structural clustering test without downloading any databases:
roundabout -f test/ -o test_results/ --skip-db
To test the full annotation and global alignment capabilities, ensure the databases are set up (roundabout --setup-db), and run:
roundabout -f test/ -o test_results/
Once the run completes, explore the test_results/ directory to see the generated similarity matrices, annotation tables, and visual plots.
Example Visualizations
Roundabout automatically generates publication-ready figures for your plasmid cohorts.
1. PyGenomeViz Linear Synteny Plot
Visualizes structural rearrangements, inversions, and conserved gene blocks (like AMR genes) across a cohort.
2. MinkeMap Circular Plot
Maps sequence tracks against a reference backbone to highlight coverage and structural conservation.
3. Skani ANI Clustermap
Separates mixed FASTAs into distinct outbreak clusters based on Average Nucleotide Identity.
4. Skani Global Scatter Plot
Places local isolates into the context of the global NCBI RefSeq database.
Understanding the Output
Roundabout generates a structured output directory containing both the raw data and the final comparative visualizations. By default, these are saved in the results/ directory.
Here is a breakdown of the key output folders and what they contain:
1. Staging and Annotations
staging_fastas/: Contains the individual, cleaned sequence files. Multi-FASTA inputs are split here before processing.amrfinder_results/: Contains the raw TSV outputs from AMRFinderPlus for each sequence, alongside anamr_groups.jsonfile detailing the cohort assignments based on shared resistance profiles.plasmidfinder_results/: Contains the incompatibility/replicon typing results and theplasmidfinder_groups.jsonassignment file.bakta_results/: Contains the comprehensive structural and functional annotations for each sequence, including standard GFF3, GenBank (.gbff), and nucleotide/protein FASTA files.
2. Clustering and Similarity
skani_results/: The core clustering engine outputs.skani_matrix.tsv: The raw pairwise alignment metrics.local_ani_clustermap.png: A hierarchical heatmap showing how your local isolates group together.global_ani_scatter.png: A scatter plot placing your local isolates into context against the global NCBI RefSeq database.similarity_groups.json: The cohort assignments based on your defined ANI and alignment fraction thresholds.
3. Cohort-Specific Visualizations
Roundabout processes sequences that group together into isolated cohorts, generating dedicated visualizations for each outbreak cluster.
ani_heatmap_results/: Contains isolated, cohort-specific distance heatmaps.daisyblast_results/: Contains the output of the sequence shattering, structural feature grouping, and dotplots for each cohort.
Click to view daisyblast comparison examples
DaisyBlast syteny groups in separate visualizations
DaisyBlast mock blast results
DaisyBlast generates combined (shown) and separate dot plots
-
minkemap_results/: Contains circular sequence alignments showing how the cohort maps against a central reference sequence. -
pygenomeviz_results/: Contains syteny maps from fasta and gbff files created by
Click to view pygenomeviz comparison examples
BLAST Comparison (GBFF and FASTA)
MUMmer Comparison (GBFF and FASTA)
MMSEQS (GBFF only)
Progressive Mauve (FASTA only)
Background: From Nextflow to Native Python
Roundabout was originally developed and structured as a Nextflow workflow. While Nextflow is an excellent orchestration engine for executing large-scale genomic pipelines across high-performance computing clusters, it ultimately proved to be the wrong architectural fit for this specific project.
The vast majority of Roundabout's core processing, mathematical clustering, and visualization steps rely on native Python libraries (such as Pandas, SciPy, PyGenomeViz, and custom Python scripts). In a Nextflow environment, passing complex data structures—like Pandas DataFrames or clustering dictionaries—between disparate processes required constantly writing them to disk as intermediate text files and re-parsing them in the next step. This added unnecessary I/O overhead and made the codebase highly fragmented and difficult to maintain.
By converting Roundabout into a standalone, modular Python package, we achieved several major improvements:
- Simplified Maintenance: The codebase is now a single, cohesive Python application. Developers no longer have to manage Nextflow DSL alongside disconnected Python scripts.
- In-Memory Processing: Dataframes and dictionaries are passed directly between functions via the
engine.pyorchestrator, drastically reducing file I/O operations and speeding up execution. - Easier Distribution: Instead of requiring end-users to install Nextflow, configure execution profiles, and manage container runtimes, users can now install the entire toolchain with a single
conda installcommand. - Robust Testing: Moving to a native Python architecture allowed us to implement comprehensive unit and integration testing via
pytest, ensuring long-term stability and easier open-source contributions.
AI Collaboration Disclaimer
This project originally began as a Nextflow workflow and was entirely refactored, modernized, and packaged into a native Python library with the collaborative assistance of Gemini, an AI developed by Google.
AI was utilized as a peer-programming tool to accelerate the conversion of Nextflow processes into modular Python functions, implement structural code optimizations across the orchestrator modules, and design the automated pytest and GitHub Actions continuous integration suites. All automated scripts, integration tests, and architecture updates generated during this process were thoroughly reviewed, manually tested, and biologically validated to ensure long-term pipeline stability and accuracy for public health genomic epidemiology.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file roundabout-0.7.26160.tar.gz.
File metadata
- Download URL: roundabout-0.7.26160.tar.gz
- Upload date:
- Size: 72.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
48e7884a7ee354912a0aef8fd8e4610fcea422dd1caa5a6bcffe27e56717aa98
|
|
| MD5 |
98ecad00e13dc095deb0ea1c91b27678
|
|
| BLAKE2b-256 |
4a2d90248091cd10ed5b1452fd72a4b5dd17959f2acfae56aac901b4fee9a6d7
|
Provenance
The following attestation bundles were made for roundabout-0.7.26160.tar.gz:
Publisher:
publish_pypi.yml on erinyoung/roundabout
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
roundabout-0.7.26160.tar.gz -
Subject digest:
48e7884a7ee354912a0aef8fd8e4610fcea422dd1caa5a6bcffe27e56717aa98 - Sigstore transparency entry: 2171040695
- Sigstore integration time:
-
Permalink:
erinyoung/roundabout@2496ac6c91671d4b8aa4b8675feb41543ecc63d1 -
Branch / Tag:
refs/tags/0.7.26160 - Owner: https://github.com/erinyoung
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish_pypi.yml@2496ac6c91671d4b8aa4b8675feb41543ecc63d1 -
Trigger Event:
release
-
Statement type:
File details
Details for the file roundabout-0.7.26160-py3-none-any.whl.
File metadata
- Download URL: roundabout-0.7.26160-py3-none-any.whl
- Upload date:
- Size: 71.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
07a77496a5935af20b22b3ebd218d89cfe3c2948ef2a176109d7d73846b4d925
|
|
| MD5 |
193d570d43c10a8329089d2d43b25ce7
|
|
| BLAKE2b-256 |
28fdcaa71611c49c89f02ac943f775f93312b9a825ea305dd364a3bfb63ad632
|
Provenance
The following attestation bundles were made for roundabout-0.7.26160-py3-none-any.whl:
Publisher:
publish_pypi.yml on erinyoung/roundabout
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
roundabout-0.7.26160-py3-none-any.whl -
Subject digest:
07a77496a5935af20b22b3ebd218d89cfe3c2948ef2a176109d7d73846b4d925 - Sigstore transparency entry: 2171040733
- Sigstore integration time:
-
Permalink:
erinyoung/roundabout@2496ac6c91671d4b8aa4b8675feb41543ecc63d1 -
Branch / Tag:
refs/tags/0.7.26160 - Owner: https://github.com/erinyoung
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish_pypi.yml@2496ac6c91671d4b8aa4b8675feb41543ecc63d1 -
Trigger Event:
release
-
Statement type: