Skip to main content
https://badge.fury.io/py/sequana-nanomerge.svg JOSS (journal of open source software) DOI https://github.com/sequana/nanomerge/actions/workflows/main.yml/badge.svg https://coveralls.io/repos/github/sequana/nanomerge/badge.svg?branch=main JOSS (journal of open source software) DOI Python 3.11 | 3.12

This is is the nanomerge pipeline from the Sequana project

Overview:

merge fastq files generated by Nanopore run and generates raw data QC.

Input:

individual fastq files generated by nanopore demultiplexing

Output:

merged fastq files for each barcode (or unique sample)

Status:

production

Citation:

Cokelaer et al, (2017), ‘Sequana’: a Set of Snakemake NGS pipelines, Journal of Open Source Software, 2(16), 352, JOSS DOI doi:10.21105/joss.00352

Installation

You can install the packages using pip:

pip install sequana_nanomerge --upgrade

An optional requirements is pycoQC, which can be install with conda/mamba using e.g.:

conda install pycoQC

you will also need graphviz installed.

Usage

sequana_nanomerge --help

If you data is barcoded, they are usually in sub-directories barcoded/barcodeXY so you will need to use a pattern (–input-pattern) such as */*.gz:

sequana_nanomerge --input-directory DATAPATH/barcoded --samplesheet samplesheet.csv
    --summary summary.txt --input-pattern '*/*fastq.gz'

otherwise all fastq files are in DATAPATH/ so the input pattern can just be *.fastq.gz:

sequana_nanomerge --input-directory DATAPATH --samplesheet samplesheet.csv
    --summary summary.txt --input-pattern '*fastq.gz'

The –summary is optional and takes as input the output of albacore/guppy demultiplexing. usually a file called sequencing_summary.txt

Note that the different between the two is the extra */ before the *.fastq.gz pattern since barcoded files are in individual subdirectories.

In both bases, the command creates a directory with the pipeline and configuration file. You will then need to execute the pipeline:

cd nanomerge
bash nanomerge.sh  # for a local run

This launches a snakemake pipeline.

Concerning the sample sheet, whether your data is barcoded or not, it should be a CSV file

barcode,project,sample
barcode01,main,A
barcode02,main,B
barcode03,main,C

For a non-barcoded run, you must provide a file where the barcode column can be set (empty):

barcode,project,sample
,main,A

or just removed:

project,sample
main,A

Usage with apptainer:

With apptainer, initiate the working directory as follows:

sequana_nanomerge --use-apptainer

Images are downloaded in the working directory but you can store then in a directory globally (e.g.):

sequana_nanomerge --use-apptainer --apptainer-prefix ~/.sequana/apptainers

and then:

cd nanomerge
sh nanomerge.sh

if you decide to use snakemake manually, do not forget to add apptainer options:

snakemake -s nanomerge.rules -c config.yaml --cores 4 --stats stats.txt --use-apptainer --apptainer-prefix ~/.sequana/apptainers --apptainer-args "-B /home:/home"

Requirements

This pipelines requires the following executable(s), which is optional:

  • pycoQC

  • dot

https://raw.githubusercontent.com/sequana/nanomerge/main/sequana_pipelines/nanomerge/dag.png

Details

This pipeline runs nanomerge in parallel on the input fastq files (paired or not). A brief sequana summary report is also produced.

Rules and configuration details

Here is the latest documented configuration file to be used with the pipeline. Each rule used in the pipeline may have a section in the configuration file.

Changelog

Version

Description

1.6.0

  • modernise packaging (poetry), drop click_completion, use importlib.metadata for version, refresh CI workflows.

1.5.0

  • refactoring to use Click

1.4.0

  • sub sampling was biased in v1.3.0. Using stratified sampling to correcly sample large file. Also set a –promethion option that auomatically sub sample 10% of the data

  • add summary table

1.3.0

  • handle large promethium run by using a sub sample of the sequencing summary file (–sample of pycoQC still loads the entire file in memory)

1.2.0

  • handle large promethium run by using find+cat instead of just cat to cope with very large number of input files.

1.1.0

  • add subsample option and set to 1,000,000 reads to handle large runs such as promethion

1.0.1

  • CSV can now handle sample or samplename column name in samplesheet.

  • Fix the pyco file paths, update requirements and doc

1.0.0

Stable release ready for production

0.0.1

First release.

Metadata

Release files for sequana-nanomerge 1.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sequana-nanomerge 1.6.0
File Size Uploaded
sequana_nanomerge-1.6.0.tar.gz 28.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sequana-nanomerge 1.6.0
File Interpreter ABI Platform
sequana_nanomerge-1.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 57.5 kB

Release files / sequana_nanomerge-1.6.0.tar.gz

Download URL sequana_nanomerge-1.6.0.tar.gz
Size 28.7 kB
Tags Source
SHA-256 checksum
How to use checksums
aa5af2efcdcb50e362a2b884657564363e6f73a03a80d52a2a9cfa7954abf968
BLAKE2b-256 checksum
How to use checksums
095acad5d869ce97038f55a787721bbf719522375884fcef293c6b0f13593ea6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.13.12

Release files / sequana_nanomerge-1.6.0-py3-none-any.whl

Download URL sequana_nanomerge-1.6.0-py3-none-any.whl
Size 28.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
dc44be5e67007ee51458d93e90f0d0524742a9cb87ce04a5b24aac5eba7f3488
BLAKE2b-256 checksum
How to use checksums
1dfc25b273bfb28327df6efd62888546e994dd5798eb5e597c8f74cdaef5e70c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.13.12

Release history Release notifications | RSS feed

This release

1.6.0 This release

2 release files

1.5.0

2 release files

1.4.0

1 release file

1.3.0

1 release file

1.2.0

1 release file

1.1.0

1 release file

1.0.0

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page