Skip to main content

GitHub PyPI GitHub Workflow Status SonarCloud Quality Gate SonarCloud Coverage
DOI fair-software.eu

ms2deepscore

ms2deepscore provides a Siamese neural network that is trained to predict molecular structural similarities (Tanimoto scores) from pairs of mass spectrometry spectra.

The library provides intuitive classes to prepare data, train a Siamese model, and compute similarities between pairs of spectra.

In addition to the prediction of a structural similarity, MS2DeepScore can also make use of an embedding evaluator predict the models accuracy for each spectrum.

Reference

If you use MS2DeepScore for your research, please cite the following:

"MS2DeepScore - a novel deep learning similarity measure to compare tandem mass spectra"
Florian Huber, Sven van der Burg, Justin J.J. van der Hooft, Lars Ridder, 13, Article number: 84 (2021), Journal of Cheminformatics, doi: https://doi.org/10.1186/s13321-021-00558-4

If you use MS2Deepscore 2.0 or higher please also cite:
Cross ionization mode chemical similarity prediction between tandem mass spectra in metabolomics
Niek de Jonge, Elena Chekmeneva, Robin Schmid, David Joas, Lem-Joe Truong, Justin J.J. van der Hooft, Florian Huber Nature Communications, 17, 2483 (2026). doi: https://doi.org/10.1101/2024.03.25.586580

Setup

Requirements

Python 3.11, 3.12, 3.13 (higher will likely work, but is not tested systematically).

Installation

Installation is expected to take 10-20 minutes.

Prepare environment

We recommend creating an Anaconda environment with

conda create --name ms2deepscore python=3.13
conda activate ms2deepscore
pip install ms2deepscore

# You can install hardware accelerated versions (for inference) with:
# for all x64/x86 and macOS aarch64:
pip install ms2deepscore[gpu]

# for Intel iGPU/Arc GPUs:
pip install ms2deepscore[intel]

Or, via conda:

conda create --name ms2deepscore python=3.13
conda activate ms2deepscore
conda install --channel bioconda --channel conda-forge matchms
pip install ms2deepscore

Alternatively, simply install in the environment of your choice by pip install ms2deepscore

Getting started: How to prepare data, train a model, and compute similarities.

We recommend to run the complete tutorial in notebooks/MS2DeepScore_tutorial.ipynb for a more extensive fully-working example on test data, including explanations on how to visualize the results. The expected run time on a laptop is less than 5 minutes, including automatic model and dummy data download. Alternatively, there are some example scripts below.

1) Compute spectral similarities

We provide a model which was trained on > 500,000 MS/MS combined spectra from GNPS, Mona, MassBank and MSnLib.

This model can be downloaded from from zenodo here. Only the ms2deepscore_model.pt is needed. The model works for spectra in both positive and negative ionization modes and even predictions across ionization modes can be made by this model.

To compute the similarities between spectra of your choice you can run the code below. There is a small example dataset available in the folder "./tests/resources/pesticides_processed.mgf". Alternatively you can of course use your own spectra, most common formats are supported, e.g. msp, mzml, mgf, mzxml, json, usi.

Pytorch version

from ms2deepscore.models import load_model
from matchms.Pipeline import Pipeline, create_workflow
from matchms.filtering.default_pipelines import DEFAULT_FILTERS
from ms2deepscore import MS2DeepScore

model_file_name = "ms2deepscore_model.pt"
spectrum_file_name = "pesticides.mgf"

# load in the ms2deepscore model
model = load_model(model_file_name, allow_legacy=True)

pipeline = Pipeline(create_workflow(query_filters=DEFAULT_FILTERS,
                                    score_computations=[[MS2DeepScore, {"model": model}]]))
report = pipeline.run(spectrum_file_name)
similarity_matrix = pipeline.scores.to_array()

The resulting similarity matrix, is a numpy array containing all the MS2DeepScore predictions between all spectra.

Hardware accelerated version (ONNX)

To use the hardware accelerated version with .onnx models (see Prepare environment):

from matchms.Pipeline import Pipeline, create_workflow
from matchms.filtering.default_pipelines import DEFAULT_FILTERS
from ms2deepscore import MS2DeepScoreONNX
from ms2deepscore.models import SiameseSpectralModelONNX

# Use the .onnx version here. See Model conversion on how to convert .pt models.
model_file_name = "ms2deepscore_model.onnx"
spectrum_file_name = "pesticides.mgf"

# load in the ms2deepscore model
model = SiameseSpectralModelONNX(model_file_name)

pipeline = Pipeline(create_workflow(query_filters=DEFAULT_FILTERS,
                                    score_computations=[[MS2DeepScoreONNX, {"model": model}]]))
report = pipeline.run(spectrum_file_name)
similarity_matrix = pipeline.scores.to_array()

2 Create embeddings

To calculate chemical similarity scores, MS2DeepScore first calculates an embedding (vector) representing each spectrum. This intermediate product can also be used to visualize spectra in "chemical space" by using a dimensionality reduction technique, like UMAP. You can either use the Pytorch version or the hardware accelerated version (ONNX) to create embeddings.

Pytorch version

from ms2deepscore import MS2DeepScore
from ms2deepscore.models import load_model

model = load_model("ms2deepscore_model.pt", allow_legacy=True)
cleaned_spectra = pipeline.spectra_queries

ms2ds_model = MS2DeepScore(model)
ms2ds_embeddings = ms2ds_model.get_embedding_array(cleaned_spectra)

Hardware accelerated version (ONNX)

from ms2deepscore import MS2DeepScoreONNX
from ms2deepscore.models import SiameseSpectralModelONNX

model = SiameseSpectralModelONNX("ms2deepscore_model.onnx")
cleaned_spectra = pipeline.spectra_queries

ms2ds_model = MS2DeepScoreONNX(model)
ms2ds_embeddings = ms2ds_model.get_embedding_array(cleaned_spectra)

The tutorial shows how to use these embeddings to create an interactive UMAP with overlaying smiles.

3) Train your own MS2DeepScore model

Training your own model is only recommended if you have some familiarity with machine learning. You can train a new model on a dataset of your choice. That, however, should contain a substantial amount of spectra to learn relevant features, say > 100,000 spectra of sufficiently diverse types. Alternatively you can add your in house spectra to an already available public library, for instance the data used for training the default MS2DeepScore model. We recommend checking the pair sampling tutorial, since the quality of the pair sampling has to be checked and potentially re-optimized for new datasets. Particularly for smaller training sets, the pair sampling can be suboptimal if not checked. To train your own model you can run the code below. Please first ensure cleaning your spectra. We recommend using the cleaning pipeline in matchms.

from ms2deepscore.SettingsMS2Deepscore import SettingsMS2Deepscore, SettingsEmbeddingEvaluator
from ms2deepscore.wrapper_functions.training_wrapper_functions import train_ms2deepscore_wrapper

spectrum_file = "./combined_libraries.mgf"
# The settins below use default training settings and use precursor mz and ionmode as additional metadata input. 
# Have a look in the SettingsMS2Deepscore class to check other hyperparameters.
settings = SettingsMS2Deepscore(
    spectrum_file_path=spectrum_file,
    additional_metadata=[("CategoricalToBinary", {"metadata_field": "ionmode",
                                                  "entries_becoming_one": "positive",
                                                  "entries_becoming_zero": "negative"}),
                         ("StandardScaler", {"metadata_field": "precursor_mz", 
                                             "mean": 0, "standard_deviation": 1000})], 
    ionisation_mode="both")

train_ms2deepscore_wrapper(
    settings, 
    SettingsEmbeddingEvaluator() # this results in also training the embedding evaluator. Leave as None if you don't want to train this.
)

Model conversion

In version 2.9 and earlier we used .pt models. But since ONNX models enable faster inference we changed to ONNX. If you have your own models that you would like to have converted to onnx, you can convert it using the code below.

from ms2deepscore.models import load_model

model = load_model("ms2deepscore_model.pt")
model.export_to_onnx("export_dir", "ms2deepscore_model.onnx")

Contributing

We welcome contributions to the development of ms2deepscore! Have a look at the contribution guidelines.

Developing new machine learning models

If you are a developer and you are interested in creating the next generation of machine learning models the MS2DeepScore code base might be a good starting point. In this tutorial we describe how MS2DeepScore training works, why this is important and where you can find this in the code base.

Metadata

Release files for ms2deepscore 2.10.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ms2deepscore 2.10.0
File Size Uploaded
ms2deepscore-2.10.0.tar.gz 102.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ms2deepscore 2.10.0
File Interpreter ABI Platform
ms2deepscore-2.10.0-py3-none-any.whl Python 3 none any Details

Total release size: 232.1 kB

Release files / ms2deepscore-2.10.0.tar.gz

Download URL ms2deepscore-2.10.0.tar.gz
Size 102.1 kB
Tags Source
SHA-256 checksum
How to use checksums
e40f8868725db78ed43b0d50f4bcb0c383c861e1427a2d0208fdf6bb26b5211a
BLAKE2b-256 checksum
How to use checksums
f38d752ce235f6b6d3e6bd49985d4b30746eef2363dcd2ee4a4300cf8bcfdca5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / ms2deepscore-2.10.0-py3-none-any.whl

Download URL ms2deepscore-2.10.0-py3-none-any.whl
Size 130.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
349e739f5f6ca8aae213db1e88e2049a9143d8f81b87fdde0b838ac53f8ec2ba
BLAKE2b-256 checksum
How to use checksums
3ec81dfdfd0469adae857ba53b4fd3ceba9f4a6cc5b0a82e496839ce88167bae
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

2.10.0 This release

2 release files

2.9.0

2 release files

2.8.0

2 release files

2.7.2

2 release files

2.7.1

2 release files

2.7.0

2 release files

2.6.0

2 release files

2.5.5

2 release files

2.5.4

2 release files

2.5.3

2 release files

2.5.2

2 release files

2.5.1

2 release files

2.5.0

2 release files

2.4.0

2 release files

2.3.0

2 release files

2.2.0

2 release files

2.1.0

2 release files

2.0.0

2 release files

1.0.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.1

2 release files

0.3.0

1 release file

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page