Skip to main content
MZMLpy Logo

A lightweight Python library for parsing mzML mass spectrometry files. Implements a type-safe, lazy-loading API with direct support for modern mzML structures (>= 1.1.0).

Python package codecov PyPI version DOI Python 3.12+ License: MIT

Installation

pip install mzmlpy

Optional extras:

pip install mzmlpy[numpress]   # MS-Numpress decoding
pip install mzmlpy[zstd]       # Zstandard compression
pip install mzmlpy[rapidgzip]  # Parallel gzip decompression (recommended for .gz files)

MCP integration

Version 0.9.0 preserves the stored numeric dtype when decoding arrays, including exact int64 values. Numpress retains its float64 reconstruction. Code requiring float64 can use .astype(np.float64). See the numeric type migration guide.

An optional local MCP server exposes file discovery, full acquisition metadata, validation, metadata inventories and comparisons, and bounded array access to AI clients. Background jobs, reference resources, workflow prompts, and optional lossless JSONL exports support larger tasks. Spectrum processing stays in Spectacular, and plotting stays in companion visualization tools. Install with pip install "mzmlpy[mcp]" and launch with python -m mzmlpy mcp --root /absolute/path/to/data. See the MCP guide.

Quick Start

from mzmlpy import Mzml

with Mzml("path/to/file.mzML") as reader:
    print(f"File: {reader.file_name}  |  Spectra: {len(reader.spectra)}")

    for spectrum in reader.spectra:
        mz = spectrum.mz
        intensity = spectrum.intensity
        print(f"  {spectrum.id} MS{spectrum.ms_level}{len(mz)} peaks")

Both .mzML and .mzML.gz files are supported. Metadata is parsed eagerly; binary data is decoded on demand.

Reading Gzipped Files

When opening .mzML.gz files, the gzip_mode parameter controls how the file is accessed:

"auto" is the default. It selects an embedded index first, then a current extracted cache, then complete rapidgzip sidecars, and finally creates an extracted cache. The selected route is available through reader.access_strategy.

Self-indexed gzip files created by write_indexed_gzip are detected automatically when in_memory=False. Their index lives inside the gzip header, so random access needs no extracted copy, sidecar index, or optional dependency.

from mzmlpy import Mzml, write_indexed_gzip

write_indexed_gzip("data.mzML", "data.indexed.mzML.gz")

with Mzml("data.indexed.mzML.gz", in_memory=False) as reader:
    spec = reader.spectra["controllerType=0 controllerNumber=1 scan=1234"]

The writer also accepts an ordinary .mzML.gz input. It preserves the decompressed mzML bytes exactly and writes the destination atomically. The embedded layout is compatible with pyMZML's FU version 1 indexed gzip reader.

Mode Description
"auto" (default) Reuse the fastest valid representation already available, otherwise extract into the central cache.
"extract" Decompress to <tmpdir>/mzmlpy/ and cache across sessions. First open pays decompression cost. Subsequent opens reuse the cache instantly. The OS clears tmp on reboot.
"indexed" Seekable access to the compressed file using rapidgzip. No decompression to disk. Requires pip install mzmlpy[rapidgzip].
"stream" Stream sequentially. Lowest startup cost but no efficient random access.

For most use cases, "extract" or "indexed" is recommended:

# Automatic selection with observable behavior
with Mzml("data.mzML.gz", in_memory=False) as reader:
    print(reader.access_strategy)
    spec = reader.spectra[0]

# Indexed — no extraction, seekable access (requires rapidgzip)
with Mzml("data.mzML.gz", gzip_mode="indexed", in_memory=False) as reader:
    spec = reader.spectra[0]

To reclaim disk space before the OS clears tmp on reboot:

from mzmlpy import clear_cache
clear_cache()

Performance

"extract" pays a one-time decompression cost then matches plain .mzML speed on later opens (the extracted copy is cached). "indexed" pays a one-time index-build cost for seekable access with no disk copy. "stream" has the lowest startup cost but random access re-scans from the start, so it's sequential-only in practice. See benchmarks/ for a reproducible harness with real numbers on real files, including a head-to-head against pyteomics and pymzml.

mzmlpy vs pymzml

Compared against pymzml 2.6.0 on a Bruker timsTOF file with ion mobility (10 spectra, 6.7 MB):

Benchmark mzmlpy pymzml Ratio
Startup 0.012s 0.092s 8.0x faster
Iterate (decode) 0.039s 0.228s 5.8x faster
Random access 0.012s 0.110s 9.2x faster

Both libraries produce identical m/z and intensity arrays. The gap narrows on smaller files (~1.1--1.3x) and widens on larger, more complex files. See the full results in the Benchmarks page or run benchmarks/bench_vs_pymzml.py yourself.

For full usage examples see the Getting Started guide and API Reference.

Using an AI coding assistant? Point it at llms.txt — a compact, accurate API guide for generating correct mzmlpy code.

Validation and filtering

from mzmlpy import Mzml, validate

report = validate("data.mzML", decode_binary=True)
print(report.valid, report.issues)

with Mzml("data.mzML", in_memory=False) as reader:
    for spectrum in reader.spectra.filter(ms_level=2, retention_time=(60, 180)):
        print(spectrum.id)

Validation reports structural and decoding problems without repairing the input. Filtering uses metadata and inclusive retention-time bounds in seconds, without decoding arrays. See the guide for check scope, precursor filters, cache behavior, and python -m mzmlpy CLI commands.

Citation

Citation metadata are provided in CITATION.cff. The archived v0.6.0 release is available from Zenodo at doi:10.5281/zenodo.21960080.

Benchmarks

benchmarks/ contains a reproducible harness comparing mzmlpy against pyteomics and pymzml on compression-format support, throughput, and gzip handling. See benchmarks/README.md for how to run it and current results.

Development

just lint     # ruff check
just format   # ruff isort + format
just ty       # ty type checker
just test     # pytest

# or all at once:
just check

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mzmlpy-0.9.0.tar.gz (89.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mzmlpy-0.9.0-py3-none-any.whl (99.6 kB view details)

Uploaded Python 3

File details

Details for the file mzmlpy-0.9.0.tar.gz.

File metadata

  • Download URL: mzmlpy-0.9.0.tar.gz
  • Upload date:
  • Size: 89.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for mzmlpy-0.9.0.tar.gz
Algorithm Hash digest
SHA256 67026ff9817e000ed1b4ad18e611af1fc3b6ee78e79b415c341f130040268d04
MD5 c70a6a6231a316904000f35df84daeb7
BLAKE2b-256 e0996f653cd282d5a3af805769a9279e90f7c3647b2a0885dee0c36fa6bf06fd

See more details on using hashes here.

File details

Details for the file mzmlpy-0.9.0-py3-none-any.whl.

File metadata

  • Download URL: mzmlpy-0.9.0-py3-none-any.whl
  • Upload date:
  • Size: 99.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for mzmlpy-0.9.0-py3-none-any.whl
Algorithm Hash digest
SHA256 c7462d2648ee186556f04253279c48fb9a6681c0e9f4494e30506d303631f18e
MD5 33e3c23d6319685f1cd35200ab29fc16
BLAKE2b-256 44ad9ac11b745e7aae7d3c683246aebdbdd3a0eee449b16770ecaa1b545eb3bc

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

0.9.0 This release

2 files

0.8.0

2 files

0.7.0

2 files

0.6.0

2 files

0.5.0

2 files

0.4.0

2 files

0.3.1

2 files

0.3.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page