A Python package for extracting data from Bruker timsTOF data files (.tdf and .tdf_bin). Includes a Numba-accelerated centroiding algorithm for efficient extraction of ion mobility data.
tdfpy reads Bruker timsTOF .d acquisitions straight from analysis.tdf and analysis.tdf_bin — no Bruker native library required. It gives you familiar Python objects for DDA, DIA, and PRM runs, plus a tunable, Numba-accelerated centroiding pipeline for pulling clean, ion-mobility-resolved peaks out of raw PASEF frames.
It's for proteomics and mass spec developers who want to script against timsTOF data without hand-rolling SQLite queries or reverse-engineering the binary frame format.
- Pure Python, no native dependency —
analysis.tdf_binis decoded directly, so it runs on Linux, macOS, and Windows, x86-64 and ARM - One API for DDA, DIA, and PRM — frames, precursors, isolation windows, targets, and transitions are all typed Python objects
- Composable peak pipeline — chain region exclusion, smoothing, and noise filters before centroiding, or use short-hand defaults
- Two centroiders — a Numba-JIT'd greedy merge in float m/z space, and a watershed region-grower in integer TOF-index space, swappable without touching surrounding code
- Lazy spectral access — frame metadata loads upfront; raw peak data is only decoded when you call
.merged_peaks(),.scan_peaks(),.raw_peaks(), or.centroid() - Query by m/z and RT, not just row index
pip install tdfpyRequires Python 3.12+. On Python 3.12/3.13 the zstandard package is installed automatically; Python 3.14+ uses the standard library's zstd module.
Optional extras:
pip install "tdfpy[viz]" # matplotlib-based plotting helpers
pip install "tdfpy[mcp]" # MCP server for AI-agent access to acquisitionsfrom tdfpy import DDA
with DDA("sample.d") as dda:
# Iterate over MS1 frames
for frame in dda.ms1:
print(f"Frame {frame.frame_id} at RT {frame.rt:.1f}s")
peaks = frame.centroid() # shape (N, 3): [m/z, intensity, 1/K0]
print(f" {len(peaks)} centroided peaks")
break
# Iterate over precursors (MS2)
for precursor in dda.precursors:
print(f"Precursor {precursor.precursor_id}: {precursor.largest_peak_mz:.4f} m/z")
peaks = precursor.merged_peaks() # MS2 centroided by tdfpy (mobility collapse + merge)
breakDIA and PRM acquisitions work the same way with DIA(...) and PRM(...); see the getting started guide for both.
| Feature | Example |
|---|---|
| Lookups & queries | dda.precursors.query(precursor_mz=1292.63, mz_tolerance=20.0, rt=2400.0, rt_tolerance=30.0) — by ID or by m/z/RT window |
| Custom peak pipelines | frame.centroid(exclude=ChargeStateRegion(), smooth=Smooth(...), noise=[MadThreshold(k=3), ...], centroid=WatershedCentroider(...)) |
| Noise filter shorthand | frame.centroid(noise="mad") or frame.centroid(noise=500.0) for common cases |
| CLI validation | tdfpy validate sample.d --full checks every binary frame without modifying the acquisition |
| MCP server | tdfpy-mcp exposes acquisition inspection and spectrum extraction as tools for AI agents |
Full pipeline options (region exclusion, smoothing, the two centroiders, and the noise-filter chain) are covered in the analysis guide and API reference.
tdfpy is the timsTOF reader in the tacular-omics family. mzmlpy reads mzML files the same way, and spxtacular builds spectrum-processing pipelines on top of either.
Full documentation: tacular-omics.github.io/tdfpy
If you use tdfpy in published work, please cite it — see CITATION.cff or the DOI record.