Skip to content

Repository files navigation

MZMLpy Logo

Python package codecov PyPI version DOI Python 3.12+ License: MIT

mzmlpy is a Python library for reading mzML mass spectrometry files. It's built for people writing proteomics or metabolomics pipelines who need a reader that's fast on large files, tells them exactly what's wrong with a malformed file, and doesn't force a full decode just to look at a spectrum's metadata.

Why mzmlpy?

  • Lazy by design — metadata is parsed up front; binary m/z and intensity arrays are only decoded when you actually touch them.
  • Fast random access — opens and indexes a 48 MB Orbitrap file in about 0.05 s (pyteomics: about 1 s); full decoding is on par with pyteomics and pymzml (see below).
  • Type-safe — dataclass-based models with full type annotations, not loosely-typed XML trees.
  • Handles gzip well — reads .mzML.gz directly, with a self-indexed gzip format for random access without decompressing the whole file.
  • Common compressions — zlib out of the box; zstd and MS-Numpress through optional extras.
  • Validates, not just parses — a validate() function reports structural and decoding problems instead of silently producing bad data.

Install

pip install mzmlpy

Optional extras:

pip install mzmlpy[numpress]   # MS-Numpress decoding
pip install mzmlpy[zstd]       # Zstandard compression
pip install mzmlpy[rapidgzip]  # Parallel gzip decompression (recommended for .gz files)
pip install mzmlpy[mcp]        # MCP server for AI coding assistants

Quick example

from mzmlpy import Mzml

with Mzml("path/to/file.mzML") as reader:
    print(f"File: {reader.file_name}  |  Spectra: {len(reader.spectra)}")

    for spectrum in reader.spectra:
        mz = spectrum.mz
        intensity = spectrum.intensity
        print(f"  {spectrum.id} MS{spectrum.ms_level} — {len(mz)} peaks")

Both .mzML and .mzML.gz files are supported. Metadata is parsed eagerly; binary data is decoded on demand.

What else it can do

from mzmlpy import Mzml, validate

# Structural/decoding validation, no repair attempted
report = validate("data.mzML", decode_binary=True)
print(report.valid, report.issues)

# Filter by metadata without decoding any arrays
with Mzml("data.mzML") as reader:
    for spectrum in reader.spectra.filter(ms_level=2, rt_range=(60, 180)):
        print(spectrum.id)

Gzipped files get the same lazy, indexable access as plain mzML — gzip_mode="auto" uses an embedded index, then rapidgzip if installed, then decompression into memory, and never writes files next to yours. For fast re-opens, convert once with write_indexed_gzip (random access with no extra files) or open once with gzip_mode="indexed" to save reusable sidecar indexes. Ion mobility data (e.g. Bruker timsTOF PASEF) is exposed on the spectrum whether it's stored as a binary array or a scan-level parameter.

Feature Notes
.mzML / .mzML.gz Transparent gzip handling, including self-indexed files
Validation validate() reports issues without altering the file
Filtering By MS level, retention time, and precursor, without decoding arrays
Ion mobility Detects both array-based and scan-level IM data
MCP server pip install mzmlpy[mcp] — file discovery, metadata, and bounded array access for AI clients
CLI python -m mzmlpy for validation and inspection from the shell

See the Getting Started guide and API Reference for the full picture, including gzip mode details, the CLI, and the MCP server.

Using an AI coding assistant? Point it at llms.txt, a short index of the package and its docs, or at llms-full.txt for the full API guide with signatures and examples.

In the tacular-omics family

mzmlpy reads mzML; tdfpy reads the Bruker timsTOF .d format the same way. Both feed spectra into spxtacular, the shared spectrum-processing layer for deisotoping, deconvolution, and downstream analysis.

Links

Citation

Citation metadata are provided in CITATION.cff. All archived releases are available from Zenodo at doi:10.5281/zenodo.21960079.

Benchmarks

benchmarks/ contains a reproducible harness comparing mzmlpy against pyteomics and pymzml on compression-format support, throughput, and gzip handling.

Measured 2026-09-24 on a 47.9 MB, 3,392-spectrum Q Exactive HF file (PRIDE PXD015669, QEHF1_09771_JB), pyteomics 5.0.1 and pymzml 2.7.0, range over two runs on a shared workstation (indicative only):

Benchmark mzmlpy pyteomics pymzml
Open + build index 0.047–0.056 s 0.94–1.08 s —
Decode every spectrum 2.2–2.7 s 3.4–4.8 s 2.8–3.3 s
Open + 8 scattered reads 0.053–0.058 s 1.1–2.1 s error on this file

Full decoding is of the same order in all three; the large differences are index construction, random access and encoding coverage (pymzml returns no peaks for Numpress-then-zlib and zstd arrays; pyteomics cannot read zstd). See benchmarks/README.md for how to run it yourself, and the full results on the Benchmarks page.

License

MIT — see LICENSE.

About

Fast, lazy mzML reader for Python: typed models, gzip random access, validation, Numpress/zstd decoding

Topics

Resources

Contributing

Security policy

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages