This release is a pre-release and may not be stable for production use.
tadamz
tadamz is a specialized Python package for targeted LC-MS(/MS) (Liquid Chromatography-Mass Spectrometry) data analysis, built on top of the emzed3 mass spectrometry framework.
Developed at ETH Zurich (Institute of Microbiology), tadamz provides a robust, modular, and highly configurable pipeline to extract, classify, normalize, calibrate, and quantify targeted compounds from MS1, MS/MS, or SRM chromatograms.
🚀 Key Features
- 📊 Peak Extraction & Integration: High-performance chromatography integration, baseline subtraction (via
pybaselines), and extraction of chromatogram peaks based on targeted precursor $m/z$ and retention times. - 🤖 Peak Classification: Integrates a Random Forest classifier (
scikit-learn) to score peak quality (peak_quality_score), filtering out noise or false positives based on peak metrics (symmetry, FWHM, peak-to-noise ratio, etc.). - 🔄 RT Adaptation & Co-elution Alignment: Corrects and adapts target retention times dynamically across batches and alignments using reference co-eluting compound peaks.
- ⚖️ Calibration & Absolute Quantification: Fits weighted (e.g., $1/x$, $1/x^2$, $1/s^2$) linear or quadratic calibration models to calibrant tables to achieve accurate absolute concentration calculations.
- 📉 Normalization Options: Includes sample-wise normalization, standard-free normalization, and natural abundance/isotopologue overlay correction for stable isotope labeling experiments.
- 📝 Automated Report Generation: Automatically generates comprehensive PDF reports featuring global abundance heatmaps and compound-wise acquisition-order trends (for retention times, FWHM, and abundance).
- 💻 Interactive emzed GUI Integration: Built-in tools for interactive table exploration, manual peak scoring, and data visualization via
emzed-spyder.
📦 Installation
tadamz requires Python >= 3.11 and the core emzed framework.
You can install tadamz directly via pip from the repository:
pip install git+https://gitlab.com/emzed3_extensions/targeted.git
This will automatically install dependencies, including:
emzed >= 3.0.2pyyaml,uncertainties,pybaselines == 1.1.0numpy >= 2,seaborn,sklearn-migrator >= 0.22.1
🛠️ Getting Started
1. Configuration (YAML)
tadamz workflows are driven by configuration dictionaries, typically stored as YAML files:
extract_peaks:
integration_algorithm: linear
ms_data_type: MS_Chromatogram
mz_tol_abs: 0.3
peak_search_window_size: 60
subtract_baseline: true
classify_peaks:
scoring_model: random_forest_classification
scoring_model_params:
classifier_name: srm_peak_classifier
normalize_peaks:
sample_wise: true
correct_int_std_for_nat_abundance: false
processing_steps:
- extract_peaks
- classify_peaks
postprocessings:
- postprocessing1
- postprocessing2
postprocessing1:
- classify_peaks
- coeluting_peaks
- normalize_peaks
2. Workflow Execution
Load targets, samples, and config, and execute the pipeline:
from tadamz import run_workflow, postprocess_result_table, load_config
from tadamz.in_out import load_targets_table, load_samples_from_folder
# Load inputs
config = load_config("path/to/config.yaml")
targets_table = load_targets_table("path/to/targets_table.xlsx")
samples = load_samples_from_folder("path/to/raw_data_folder")
# Run main workflow
result = run_workflow(targets_table, samples, config)
# Post-process (e.g., RT coelution alignment & peak normalization)
result = postprocess_result_table(result, config, postprocess_id=0)
🤖 Training a custom Peak Classifier
Train custom peak classifiers for your targeted assays using manually scored peaks:
from tadamz import generate_peak_classifier
generate_peak_classifier(
classifier_name="my_custom_srm_classifier",
path_to_folder="path/to/save/classifier",
path_to_table="path/to/measured_peaks.table",
ms_data_type="MS_Chromatogram",
inspect=True, # Opens interactive GUI for manual scoring
)
📊 Generating Reports
tadamz can create comprehensive PDF reports of your targeted analysis:
from tadamz.generate_report import generate_report
# Generate global and compound-wise report figures
global_plots, compound_plots = generate_report(result_table, pdf_folder="path/to/output_folder")
📂 Project Structure
tadamz/
├── src/
│ └── tadamz/ # Main package source code
│ ├── calibration/ # Regression models & absolute calibration curve fitting
│ ├── data/ # Default configuration files & sample datasets
│ ├── scoring/ # Peak quality metrics & Random Forest classifiers
│ ├── in_out.py # IO helpers for tables, YAML, JSON, and PKL files
│ ├── workflow.py # Main workflow definition and runner functions
│ └── generate_report.py # PDF report generator using matplotlib and seaborn
├── tests/ # Unit tests & regression tests (run via pytest)
├── setup.cfg # Metadata, classifiers, and project dependencies
└── pyproject.toml # Build system configuration
📄 License & Credits
- Author: Patrick Kiefer (pkiefer@ethz.ch)
- Organization: Institute of Microbiology, ETH Zurich
- License: MIT License
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file tadamz-1.0.0a4.tar.gz.
File metadata
- Download URL: tadamz-1.0.0a4.tar.gz
- Upload date:
- Size: 395.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.12.10
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
21a60a1f85a793cc08fe5ede311a962b495c5d293f7df16210baf5edb98f5dec
|
|
| MD5 |
6fd81115e10e09c3cf92fc587cfc131b
|
|
| BLAKE2b-256 |
02be96526b0b6e706b0ebb41b220fb45a5df9384a45c6acc454a89b14fb112ae
|
File details
Details for the file tadamz-1.0.0a4-py3-none-any.whl.
File metadata
- Download URL: tadamz-1.0.0a4-py3-none-any.whl
- Upload date:
- Size: 371.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.12.10
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c129ef950c4f1d3e8e6ae782a50e6c7fc8c0dc0786a0aec3b5152e44880f174a
|
|
| MD5 |
1035e002da3e145d9d94704ff451d8fb
|
|
| BLAKE2b-256 |
98076c0596061499c78462885f8c5deb46b75dbd847cc1507c3f32a3003f3036
|