Skip to main content

Bygul: Amplicon & Metagenomics Read Simulator

Bygul is a Python 3 tool designed for simulating sequencing reads in wastewater surveillance and other metagenomic applications. It allows users to simulate complex multi-sample datasets with customizable proportions using industry-standard backends like wgsim and mason.


🏗 Installation

Bygul requires Python 3. Since it relies on external simulators (wgsim and mason), we recommend using Conda to manage dependencies.For more info on wgsim and mason simulator please check their documentations.

conda create -n bygul bioconda::bygul

Option 2: Via PyPI

pip install bygul

Note: Some binary dependencies (wgsim/mason) may need to be installed manually or built from source if using this method.

Option 3: Local Build from Source

git clone [https://github.com/andersen-lab/Bygul](https://github.com/andersen-lab/Bygul)
cd Bygul
pip install -e .

🧬 Usage: Amplicon Sequencing Mode

Use this mode when simulating specific genomic regions defined by a primer set.

Basic Command

bygul simulate-proportions --genomes [SAMPLE1.fasta,SAMPLE2.fasta] --primers [primer.bed] --reference [reference.fasta] --proportions [0.8,0.2] --outdir [output_dir]

Advanced Examples

  • Random Proportions & Mismatches: Simulate with random proportions and allow up to 2 SNPs in primer regions.
    bygul simulate-proportions --genomes sample1.fasta,sample2.fasta --primers primer.bed --reference reference.fasta --outdir results/ --maxmismatch 2
    
  • Switching Simulators: Use mason instead of the default wgsim.
    bygul simulate-proportions --genomes sample1.fasta,sample2.fasta --primers primer.bed --simulator mason
    
  • Custom Error Rates & Lengths: Pass simulator-specific parameters (e.g. indel fraction -R) directly.
    bygul simulate-proportions --genomes sample1.fasta,sample2.fasta --primers primer.bed -R 0.01
    
  • Using a csv file and all samples in a multi-fasta file:
    bygul simulate-proportions --csv samples.csv --multifasta samples.fasta
    

🌍 Usage: Metagenomics Mode

Simulate reads from entire samples without requiring a primer BED file or a reference sequence.

Basic Metagenomics Simulation

bygul simulate-proportions sample1.fasta,sample2.fasta --outdir results/ --simulation_mode metagenomics

Metagenomics with Specific Parameters

bygul simulate-proportions sample1.fasta,sample2.fasta --proportions 0.5,0.5 --outdir results/ --simulation_mode metagenomics --simulator mason --read_length 200

Metagenomics with csv and multifasta

bygul simulate-proportions --csv samples.csv --multifasta samples.fasta --outdir results/ --simulation_mode metagenomics

📝 Technical Notes

Parameter Handling

Bygul acts as a wrapper. While most flags are passed directly to the underlying simulators, the following are managed directly by Bygul for more realistic simulations:

  • --readcnt: Number of reads per amplicon.
  • --insert_size: Insert size mean value.
  • --insert_size_sd: Insert size standard deviation.
  • --read_length: Read length.
  • --wgsim_error_rate/--wgsim_amp_error_rate/art_seq_system: amplicon specific parameters.

To see all available backend flags, run:

wgsim --help
mason_simulator --help
art_illumina --help

Please note that some dependencies are not available through pypi. You need to install them using conda or build from source.

Reference file

Reference file is used only when the provided bed file does not have the sequence column. We strongly recommend for your file to have a sequence column as the program will extract the sequences from the reference sequence if not provided.

Using a CSV file for sample names and proportions

--csv and --multifasta are always provided together, the CSV file contains two columns sample_name and proportion. Samples with multiple contigs, must have their IDs as: sample_name|contig_name in the multifasta file.

Number of reads per amplicon

It is recommended to define the number of reads per amplicon to be greater than the number of contigs in your amplicon file. This is particularly important when your primers are designed for whole genome sequencing, where each amplicon may contain a substantial number of contigs. Setting too few reads per amplicon may result in empty read files for certain amplicons, leading to incomplete simulated reads.

Primer bed file

🧬 Input BED File Format

The pipeline expects a tab-delimited BED file where the first six columns represent standard genomic coordinates (chrom, chromStart, chromEnd, name, poolName, strand). Crucially, the fourth column (name) must follow a strict naming convention to prevent downstream parsing failures in variant-calling tools: [Scheme-Name]_[AmpliconNumber]_[Direction]_[OptionalSuffix] (e.g., SARS-CoV-2_3_LEFT or SARS-CoV-2_3_LEFT_alt). To ensure structural boundaries are parsed correctly, the prefix must not contain underscores, and any optional trailing modifiers must be restricted to standard alternative tags (_alt, _ALT1) or tracking indexes (_0, _1). Multi-level pool formatting, such as SARS-CoV-2_400_1_LEFT_1, is malformed and will fail validation.. The maximum number of mismatches allowed for each primer sequence is 1 SNP. To change this number, you may use the --maxmismatches flag.

Complete set of available parameters

To learn more about how to adjust other parameters for the simulator please read the documentation for wgsim and mason simulator. Users can pass any simulator parameter directly in their command. The only parameters set through bygul are --readcnt and --wgsim_insert_size,--wgsim_read_length and --wgsim_error_rate.

Simulated reads output

Simulated reads from all samples are located in provided_output_path/reads_1.fastq and provided_output_path/reads_2.fastq

Information about amplicon dropouts

In order to find out more about amplicon dropouts, please refer to provided_output_path/sample_name/amplicon_stats.csv file. Please note that primer_seq_x and primer_seq_y define the left and right primer sequence whereas left_match and right_match shows the actual sequence found in the sample for a better comparison of mismatching bases in the primer sequence. Additionally, if there are any ambiguous bases present in the matching sequence, the ambiguous_bases value returns true.

Metadata

Release files for bygul 4.0.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for bygul 4.0.2
File Size Uploaded
bygul-4.0.2.tar.gz 18.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for bygul 4.0.2
File Interpreter ABI Platform
bygul-4.0.2-py3-none-any.whl Python 3 none any Details

Total release size: 35.3 kB

Release files / bygul-4.0.2.tar.gz

Download URL bygul-4.0.2.tar.gz
Size 18.6 kB
Tags Source
SHA-256 checksum
How to use checksums
21d01521da77413e2f6d1e8dd258e005d2a2deb78d266e36a1443e53a578175b
BLAKE2b-256 checksum
How to use checksums
74c49cfdabcb209b0b35156cd00173b4c7c496d60a651af5ca28fcc9a96c6bed
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 28, 2026.

Transparency log

Release files / bygul-4.0.2-py3-none-any.whl

Download URL bygul-4.0.2-py3-none-any.whl
Size 16.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
db8bd1d341d95d875d9a2b45c540d1d8de19957a87fa017a2f83e59d4fd4fe5f
BLAKE2b-256 checksum
How to use checksums
1e10df5e85a2b0bfc1b0730e44ad753c2644dd23b93ae8e4cd64816bfb84b579
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 28, 2026.

Transparency log

Release history Release notifications | RSS feed

4.0.5

2 release files

4.0.3

2 release files

This release

4.0.2 This release

2 release files

4.0.1

2 release files

4.0.0

2 release files

3.2.0

2 release files

3.1.0

2 release files

3.0.1

2 release files

3.0.0

2 release files

2.0.0

2 release files

1.0.7

2 release files

1.0.6

2 release files

1.0.5

2 release files

1.0.3

2 release files

1.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page