Skip to main content

BMXP - The Metabolomics Platform at the Broad Institute

pip install bmxp

Please cite: https://www.biorxiv.org/content/10.1101/2023.06.09.544417v1.full

This is a collection of tools for processing our data, which powers our cloud processing workflow. Each tool is meant to be a standalone module that performs a step in our processing pipeline. They are written in Python and C, and designed to be perfomant and cloud-compatible.

  • Eclipse - Align two or more same-method nontargeted LCMS datasets
  • Gravity - Cluster redundant LCMS features based on RT and Correlation (And someday, XIC shape)
  • Blueshift - Drift Correction via pooled technical replicates and internal standards
  • Formation - Formatting and Final QC
  • Chroma - Read .raw and .mzml files

We expect users to be familiar with Python and already have an understanding of LCMS Metabolomics data processing and the specific steps they wish to accomplish.

While the tools are and always will be standalone, we are working on linking them closer together with a shared schema, and eventually may have a pipeline ability to run all steps, given a set of parameters.

We are open to feedback and suggestions, with a focus on performance and application in pipelines.

Shared Schema

All BMXP modules use a shared schema and file formats with our prefered columns headers. These files are (along with their labels):

  • Feature Metadata bmxp.FMDATA - Describes the feature. Index default is Compound_ID
  • Injection Metadata bmxp.IMDATA - Describes the Injection. Index default is injection_id
  • Sample Metadata bmxp.SMDATA - Describes the biospecimen from which the Injection is derived. Index default is broad_id
  • Feature Abundances - Pivot table of Feature x Injection (Compound_ID x injection_id) containing the abundances.

Some modules (Blueshift, Eclipse) require merging Feature Metadata + Feature Abundances.

These can be changed globally so that all packages will use the same terminology. To update the schema, modify the dictionary objects in the module directly prior to running code. For example:

import bmxp
from bxmp.eclipse import MSAligner
from bxmp.blueshift import DriftCorrection
from bmxp.gravity import cluster
bmxp.FMDATA['Compound_ID'] = 'Feature_ID'
bmxp.IMDATA['injection_id'] = 'Filename'

# continue with work...

With those changes above, Eclipse, Blushift and Gravity will use "Feature_ID" and "Filename" as column headers instead of "Compound_ID" and "injection_id".

Feature Metadata - bmxp.FMDATA

Feature Metadata describes the LCMS feature. This is a mixture of fundamental nontargeted feature information, annotation info, and anything else.

Feature Specific

  • Compound_ID - Index, Project-unique feature ID (a bit of a misnomer)
  • RT - Unitless retention time, may or may not be scaled
  • MZ - Unsigned mass-to-charge ratio
  • Intensity - Average feature intensity
  • Method - Human Readable name of LCMS method used
  • __extraction_method - Name of extraction method/software used. Used to denote mixed Targeted/Nontargeted

Annotation

  • Annotation_ID - Method-unique annotation label
  • Adduct - Adduct form of the annotation
  • __annotation_id - Globally unique annotation identifier
  • Metabolite - Preferred display/reporting name of metabolite
  • Non_Quant - Boolean denoting that a feature is not quanitifiable

Generated by Gravity

  • Cluster_Num - Cluster number assigned during Gravity clustering
  • Cluster_Size - Number of members in the cluster

Generated by Blueshift

  • Batches Skipped - Batches that were skipped due to lack of PREFs

Injection Metadata - bmxp.IMDATA

  • injection_id - Index, Injection name, usually filename without the extension
  • broad_id - Assigned biospeciemn label
  • program_id - Biospecimen label as received (inherited from Sample Metadata)
  • injection_type - Type of injection ("sample", "prefa", "prefb", "blank", "other-", "not_used-")
  • comments - Comments about the injection
  • column_number - Column number, in multi-column studies
  • injection_order - Injection number, not skipping blanks or non-samples
  • batches - Denotes batches ('batch_start' or 'batch_end')

Generated by Blueshift

  • QCRole - Role in drift correction ("QC-drift_correction", "QC-pooled_ref", "QC-not_used", "sample")

Sample Metadata - bmxp.SMDATA

  • broad_id - Assigned biospecimen label
  • Arbitrary Metadata Columns - Any column label except labels in Injection Metadata

Metadata

Release files for bmxp 0.5.6

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for bmxp 0.5.6
File Size Uploaded
bmxp-0.5.6.tar.gz 2.7 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for bmxp 0.5.6
File Interpreter ABI Platform
bmxp-0.5.6-py3-none-any.whl Python 3 none any Details

Total release size: 3.9 MB

Release files / bmxp-0.5.6.tar.gz

Download URL bmxp-0.5.6.tar.gz
Size 2.7 MB
Tags Source
SHA-256 checksum
How to use checksums
59fcb52e4c7c15189a8946c265eb8340e411cfdf9ac72a28c97fe51fd8e35ca5
BLAKE2b-256 checksum
How to use checksums
f58b98e431dcf84343589b2cc8917e86c4060d3fa003a9bf10676b4c9a53da5a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.13

Release files / bmxp-0.5.6-py3-none-any.whl

Download URL bmxp-0.5.6-py3-none-any.whl
Size 1.2 MB
Tags Python 3
SHA-256 checksum
How to use checksums
0c2d5895010736f9af89e182d4c76aa85dbb1e69a853890da18ce3e7890c70a4
BLAKE2b-256 checksum
How to use checksums
234d74d0a12b2e40339b5d29d1066292296771671016ec507295482d68a9c55c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.13

Release history Release notifications | RSS feed

0.5.7

2 release files

This release

0.5.6 This release

2 release files

0.5.5

1 release file

0.5.4

1 release file

0.5.3

2 release files

0.5.2

2 release files

0.5.1

1 release file

0.5.0

2 release files

0.4.8

2 release files

0.4.7

2 release files

0.4.6

2 release files

0.4.5

2 release files

0.4.4

1 release file

0.4.3

1 release file

0.4.2

1 release file

0.4.1

1 release file

0.4.0

1 release file

0.3.18

1 release file

0.3.17

1 release file

0.3.16

1 release file

0.3.15

2 release files

0.3.13

1 release file

0.3.12

1 release file

0.3.11

2 release files

0.3.10

1 release file

0.3.9

2 release files

0.3.8

2 release files

0.3.7

2 release files

0.3.6

2 release files

0.3.5

1 release file

0.3.4

2 release files

0.3.3

1 release file

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.0

2 release files

0.1.14

2 release files

0.1.13

2 release files

0.1.11

2 release files

0.1.10

1 release file

0.1.9

1 release file

0.1.8

1 release file

0.1.7

1 release file

0.1.6

1 release file

0.1.5

1 release file

0.1.4

1 release file

0.1.3

1 release file

0.1.2

1 release file

0.1.1

1 release file

0.1.0

1 release file

0.0.18

1 release file

0.0.17

1 release file

0.0.15

1 release file

0.0.14

1 release file

0.0.13

1 release file

0.0.11

1 release file

0.0.10

1 release file

0.0.9

1 release file

0.0.8

1 release file

0.0.7

1 release file

0.0.6

1 release file

0.0.5

1 release file

0.0.4

1 release file

0.0.3

1 release file

0.0.2

1 release file

0.0.1

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page