Skip to main content

Python bindings for mf4-rs - A minimal library for working with ASAM MDF 4 measurement files

Project description

mf4-rs

mf4-rs is a minimal Rust library for working with ASAM MDF 4 (Measurement Data Format) files. It supports parsing existing files as well as writing new ones through a safe API, implementing a subset of the MDF 4.1 specification sufficient for simple data logging and inspection tasks.

The library ships for three ecosystems, all released together from this repository:

Ecosystem Package Install
Rust mf4-rs on crates.io cargo add mf4-rs
Python (PyO3 bindings) mf4-rs on PyPI pip install mf4-rs
TypeScript / JavaScript (WebAssembly bindings) mf4-rs on npm npm install mf4-rs

Architecture

High-Level Structure

The codebase is organized into distinct layers:

1. API Layer (src/api/)

  • High-level user-facing API for working with MDF files
  • MDF - Main entry point for parsing files from disk
  • ChannelGroup - Wrapper providing ergonomic access to channel group metadata
  • Channel - High-level channel representation with value decoding

2. Writer Module (src/writer/)

  • MdfWriter - Core writer for creating MDF 4.1-compliant files
  • Guarantees little-endian encoding, 8-byte alignment, and zero-padding
  • Handles block linking and manages open data blocks during writing
  • Supports both single record writing (write_record) and batch operations (write_records)

3. Block Layer (src/blocks/)

  • Low-level MDF block implementations matching the specification
  • Each block type (HeaderBlock, ChannelBlock, ChannelGroupBlock, etc.) has parsing and serialization
  • Conversion system supporting various data transformations (linear, formula, lookup tables)
  • Common utilities for block headers and data type handling

4. Parsing Layer (src/parsing/)

  • File parsing and memory management using memory-mapped files
  • Raw block parsers that maintain references to memory-mapped data
  • Channel value decoder supporting multiple data types
  • Lazy evaluation - channels and values are decoded on demand

5. Utilities (src/)

  • cut.rs - Time-based file cutting functionality
  • merge.rs - File merging utilities
  • error.rs - Centralized error handling
  • index.rs - MDF file indexing system for fast metadata-based access
  • signal.rs - Signal struct pairing channel values with the group's master/time axis

6. Bindings

Key Design Patterns

Memory-Mapped File Access: The parser uses memmap2 to avoid loading entire files into memory, enabling efficient handling of large measurement files.

Lazy Evaluation: Channel groups, channels, and values are created as lightweight wrappers that decode data only when accessed.

Builder Pattern: The writer uses closure-based configuration for channels and channel groups, allowing flexible setup while maintaining type safety.

Block Linking: The MDF format uses address-based linking between blocks. The writer maintains a position map to update links after blocks are written.

Usage

Building and Testing

# Build the project
cargo build

# Run all tests
cargo test

# Run specific test file
cargo test --test api

Examples

The project includes simplified examples in the examples/ directory:

  • write_file.rs - Comprehensive example of writing MDF files with multiple channels
  • read_file.rs - Demonstrates parsing and inspecting MDF files
  • index_operations.rs - Shows advanced indexing, byte-range reading, and conversion resolution
  • merge_files.rs - Merging multiple MF4 files
  • cut_file.rs - Time-based file cutting
  • visualize_layout.rs - Dumping the block layout of a file
  • wasm-smoke/ - Browser smoke test for the WebAssembly bindings

Run them with:

cargo run --example write_file
cargo run --example read_file
cargo run --example index_operations

Working with MDF Files

Basic File Creation Pattern:

use mf4_rs::writer::MdfWriter;
use mf4_rs::blocks::common::DataType;
use mf4_rs::parsing::decoder::DecodedValue;

let mut writer = MdfWriter::new("output.mf4")?;
writer.init_mdf_file()?;
let cg = writer.add_channel_group(None, |_| {})?;

// Create master channel (usually time)
let time_ch_id = writer.add_channel(&cg, None, |ch| {
    ch.data_type = DataType::FloatLE;
    ch.name = Some("Time".to_string());
    ch.bit_count = 64;
})?;
writer.set_time_channel(&time_ch_id)?; // Mark as master channel

// Add data channels with master as parent
writer.add_channel(&cg, Some(&time_ch_id), |ch| {
    ch.data_type = DataType::UnsignedIntegerLE;
    ch.name = Some("DataChannel".to_string());
    ch.bit_count = 32;
})?;

writer.start_data_block_for_cg(&cg, 0)?;
writer.write_record(&cg, &[
    DecodedValue::Float(1.0),              // Time
    DecodedValue::UnsignedInteger(42),     // Data
])?;
writer.finish_data_block(&cg)?;
writer.finalize()?;

Basic File Parsing Pattern:

use mf4_rs::api::mdf::MDF;

let mdf = MDF::from_file("input.mf4")?;
for group in mdf.channel_groups() {
    println!("channels: {}", group.channels().len());
    for channel in group.channels() {
        let values = channel.values()?;
        // Process values...
    }
}

MDF Indexing System

The library includes a powerful indexing system that allows you to:

  1. Create lightweight JSON indexes of MDF files containing all metadata needed for data access
  2. Read channel data without full file parsing using only the index and targeted file I/O
  3. Serialize/deserialize indexes for persistent storage and sharing
  4. Support multiple data sources through the ByteRangeReader trait (local files, HTTP, S3, etc.)

The public index API is name-based: channels and groups are addressed by name, not by positional indices.

Basic Indexing Workflow:

use mf4_rs::index::MdfIndex;

// Create an index from an MDF file (never reads sample data — only metadata)
let index = MdfIndex::from_file("data.mf4")?;

// Save index to JSON for later use
index.save_to_file("data_index.json")?;

// Later: load index and read specific channel data. The data source is not
// serialized with the index, so re-attach it after loading:
let mut index = MdfIndex::load_from_file("data_index.json")?;
index.set_file("data.mf4"); // or index.set_url("https://cdn.example.com/data.mf4")

// Lazy read: fetches only the byte ranges the channel needs and returns a
// Signal (values paired with the group's master/time axis)
let signal = index.read("Temperature")?;
println!("{:?} {:?}", signal.timestamps, signal.values);

// Disambiguate duplicate channel names by group
let engine_temp = index.read_in("Engine", "Temperature")?;

// Or build the index directly over HTTP (requires the `http` feature)
let remote = MdfIndex::from_url("https://cdn.example.com/data.mf4")?;

// Power users: compute raw byte ranges (e.g. for HTTP Range requests)
let ranges = index.byte_ranges("Temperature")?;

Custom data sources (S3, caches, …) plug in through the ByteRangeReader trait: index.open(reader) returns an MdfReader with values(name) / signal(name) methods.

Python Bindings

mf4-rs includes high-performance Python bindings generated using pyo3. This allows you to use the library's features directly from Python with minimal overhead. Wheels for Linux, macOS, and Windows are published to PyPI on every release.

Installation

pip install mf4-rs
# or
uv pip install mf4-rs

For development against a local checkout, use maturin (requires a Rust compiler):

pip install maturin
maturin develop --release

Python Examples

Check the python_examples/ directory for complete scripts:

  • write_file.py - Creating MDF files
  • read_file.py - Reading and inspecting files
  • index_operations.py - Using the indexing system
  • pandas_example.py - Pandas Series / DatetimeIndex integration
  • visualize_layout.py - Inspecting the block layout of a file

Basic Usage

All navigation in the Python API is by name. read() returns a pandas.Series (indexed by the group's master channel as a DatetimeIndex); values() returns a plain numpy float64 array (invalid samples are NaN).

import mf4_rs

# Writing a file
writer = mf4_rs.MdfWriter("output.mf4")
writer.init_mdf_file()
group = writer.add_channel_group("MyGroup")

# Add channels (the time channel is automatically the master)
time_ch = writer.add_time_channel(group, "Time")
data_ch = writer.add_float_channel(group, "Data")

# Write data
writer.start_data_block(group)
writer.write_record(group, [
    mf4_rs.create_float_value(0.1),  # Time
    mf4_rs.create_float_value(42.0)  # Data
])
writer.finish_data_block(group)
writer.finalize()

# Reading a file
mdf = mf4_rs.Mdf("output.mf4")
for group in mdf.groups:
    print(f"Group: {group.name}, Channels: {[c.name for c in group.channels]}")

series = mdf.read("Data")     # pandas Series with DatetimeIndex
array = mdf.values("Data")    # numpy float64 array (no pandas needed)
same = mdf["Data"]            # __getitem__ is read()

# Lazy remote/indexed reads
index = mf4_rs.MdfIndex.from_file("output.mf4")
index.save("output_index.json")

index = mf4_rs.MdfIndex.load("output_index.json")
index.source = "output.mf4"   # file path or http(s):// URL
series = index.read("Data")   # fetches only the byte ranges it needs

TypeScript / JavaScript Bindings (WebAssembly)

mf4-rs also ships WebAssembly bindings (src/wasm.rs, built with wasm-bindgen) wrapped in a typed TypeScript package under js/, published to npm as mf4-rs on every release. It mirrors the name-based API of the Python bindings: an in-memory Mdf reader, an MdfWriter producing a Uint8Array, and a JSON-serialisable MdfIndex for lazy, partial reads of large or remote files over any RangeSource (HTTP fetch, local file, in-memory bytes, or your own).

import { Mdf } from "mf4-rs";

const mdf = await Mdf.fromFile("./measurement.mf4");
console.log(mdf.channelNames());
const speed = mdf.values("VehicleSpeed");   // number[] (conversions applied, invalid -> NaN)
const signal = mdf.read("VehicleSpeed");    // { timestamps, values }
mdf.dispose();

See js/README.md for the full API, the writer, remote reads via MdfIndex + FetchRangeSource, and browser-target notes.

Performance

mf4-rs is designed for high performance:

  • Use write_records for batch operations instead of multiple write_record calls
  • Data blocks automatically split when they exceed 4MB to maintain performance
  • Memory-mapped file access minimizes memory usage for large files
  • Channel values are decoded lazily only when accessed
  • Use indexing for repeated access to the same files to avoid re-parsing overhead

Note: Previous benchmarks have been removed as they are being updated.

Dependencies

  • nom - Binary parsing combinators
  • byteorder - Endianness handling
  • memmap2 - Memory-mapped file I/O
  • meval - Mathematical expression evaluation for formula conversions
  • thiserror - Error handling derive macros
  • serde / serde_json - JSON serialization for the index system

Optional feature flags:

  • http - HTTP range reader (ureq + vendored TLS) for MdfIndex::from_url and URL sources
  • pyo3 - Python bindings (pyo3, numpy, pyo3-stub-gen; implies http)
  • wasm - WebAssembly bindings (wasm-bindgen, js-sys, serde-wasm-bindgen)

Releases

Releases are fully automated. Every merge to main runs .github/workflows/release.yml, which:

  1. Reads commit messages since the last v* tag.
  2. Computes a SemVer bump from Conventional Commits prefixes (feat: → minor, fix:/perf: → patch, feat!: / BREAKING CHANGE: → major; other types do not produce a release).
  3. Updates Cargo.toml, pyproject.toml, Cargo.lock, js/package.json, js/package-lock.json, and CHANGELOG.md, then tags vX.Y.Z.
  4. Builds wheels for Linux (x86_64, aarch64), macOS (x86_64, aarch64), and Windows (x64), plus an sdist, and publishes them to PyPI via OIDC trusted publishing.
  5. Publishes the Rust crate to crates.io.
  6. Builds the WebAssembly bindings with wasm-pack, compiles and tests the TypeScript wrapper, and publishes the mf4-rs npm package.

PR titles are linted against the convention by .github/workflows/pr-title.yml. Because PRs are squash-merged, the PR title is the source of truth for both the changelog and the version bump. See CHANGELOG.md for release history.

How PR titles map to version bumps

PR title prefix (or footer) Bump Example
feat!: / <type>!: / body contains BREAKING CHANGE: major feat!: drop Python 3.8 support
feat: minor feat: add ##DZ block decompression for index reads
fix: / perf: patch fix: handle empty CHANGELOG.md on first release
chore: / docs: / refactor: / test: / ci: / build: / style: no release docs: clarify VLSD record format

If multiple commits land between releases, the highest bump wins. Runs with no qualifying commits are a clean no-op — no tag is created and no artifacts are published, which is why the chore(release): vX.Y.Z commit pushed back by the workflow does not trigger another release.

What gets published

Each successful release run publishes, in parallel, after the tag is pushed:

  • PyPI: mf4-rs wheels for Linux (x86_64, aarch64), macOS (x86_64 on 13, aarch64 on 14), and Windows (x64), plus an sdist. Authentication uses PyPI Trusted Publishing (OIDC) — no long-lived API token is stored. The pypi GitHub Actions environment gates this job.
  • crates.io: the mf4-rs Rust crate, gated by a cargo publish --dry-run step. Authentication uses the CARGO_REGISTRY_TOKEN repository secret.
  • npm: the mf4-rs TypeScript/WebAssembly package from js/. The publish runs the package's prepublishOnly script (wasm-pack build, TypeScript compile, full js test suite), so a failing build or test aborts the publish. Authentication uses npm Trusted Publishing (OIDC) — no long-lived token is stored, and npm attaches a provenance attestation automatically.
  • GitHub Releases: a release named vX.Y.Z whose body is the same changelog section appended to CHANGELOG.md.

Manual triggers and recovery

  • Re-run a release: the Release workflow accepts workflow_dispatch, so a maintainer can re-run it from the Actions tab. If the tag already exists, bump-and-tag will fail; in that case re-run only the publish-pypi / publish-crates / publish-npm jobs from the original run.
  • Skip a release intentionally: use a non-releasing prefix (chore:, docs:, refactor:, test:, ci:, build:, style:) for the PR title. The workflow will run, compute bump=none, and exit cleanly without creating a tag.
  • First-time setup: a PyPI Trusted Publisher entry for mf4-rs pointing at dmagyar-0/mf4-rs workflow release.yml, environment pypi; a CARGO_REGISTRY_TOKEN secret; an npm Trusted Publisher entry on the mf4-rs npm package (package Settings → Trusted Publisher: GitHub Actions, organization dmagyar-0, repository mf4-rs, workflow release.yml); and — if branch protection blocks GITHUB_TOKEN pushes to main — a RELEASE_PAT secret with contents:write. All are complete for this repo. Note that npm's Trusted Publisher settings live on the package page, so the package must exist first: bootstrap it once by publishing manually from a logged-in machine (cd js && npm run build:wasm && npm run build && npm publish --access public), then configure the Trusted Publisher — every subsequent release publishes tokenlessly from CI.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mf4_rs-3.3.0.tar.gz (317.1 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

mf4_rs-3.3.0-cp38-abi3-win_amd64.whl (889.3 kB view details)

Uploaded CPython 3.8+Windows x86-64

mf4_rs-3.3.0-cp38-abi3-manylinux_2_28_x86_64.whl (3.4 MB view details)

Uploaded CPython 3.8+manylinux: glibc 2.28+ x86-64

mf4_rs-3.3.0-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (3.5 MB view details)

Uploaded CPython 3.8+manylinux: glibc 2.17+ ARM64

mf4_rs-3.3.0-cp38-abi3-macosx_11_0_arm64.whl (1.0 MB view details)

Uploaded CPython 3.8+macOS 11.0+ ARM64

File details

Details for the file mf4_rs-3.3.0.tar.gz.

File metadata

  • Download URL: mf4_rs-3.3.0.tar.gz
  • Upload date:
  • Size: 317.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for mf4_rs-3.3.0.tar.gz
Algorithm Hash digest
SHA256 7f43514ddc313fa8a14ec07eaea16f9cd1f0e3c191dc53a80fc1115a007bb1f4
MD5 423c279598457b5b720d491f86ec2dfa
BLAKE2b-256 a7f0c71c9c7224cd5928e4531742295a29fee1d32c109042bd61e23c934614b2

See more details on using hashes here.

Provenance

The following attestation bundles were made for mf4_rs-3.3.0.tar.gz:

Publisher: release.yml on dmagyar-0/mf4-rs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mf4_rs-3.3.0-cp38-abi3-win_amd64.whl.

File metadata

  • Download URL: mf4_rs-3.3.0-cp38-abi3-win_amd64.whl
  • Upload date:
  • Size: 889.3 kB
  • Tags: CPython 3.8+, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for mf4_rs-3.3.0-cp38-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 28182b841fa4893cfe4ab47af7a0c5fc2eb0dd4f0dfbcb0c8352c647a54f135a
MD5 08925bcf6698a55b28b2762a6ffb67d0
BLAKE2b-256 ea791d7f1943d8dae9bac2b7d819ca65537ffa7d5b209f39b29ba361cf5e6cc0

See more details on using hashes here.

Provenance

The following attestation bundles were made for mf4_rs-3.3.0-cp38-abi3-win_amd64.whl:

Publisher: release.yml on dmagyar-0/mf4-rs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mf4_rs-3.3.0-cp38-abi3-manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for mf4_rs-3.3.0-cp38-abi3-manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 deb40c5c2a04441dad1fe410e4dc81677a866e28500c9dab40d93458bd69b03b
MD5 1a6a2eaf23106312d95d87332040fda4
BLAKE2b-256 b0a4b602d13394b829718d8b5a385ff8cf1e6f22647f329e4fd9a382acfba733

See more details on using hashes here.

Provenance

The following attestation bundles were made for mf4_rs-3.3.0-cp38-abi3-manylinux_2_28_x86_64.whl:

Publisher: release.yml on dmagyar-0/mf4-rs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mf4_rs-3.3.0-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for mf4_rs-3.3.0-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 9f9a4ad41fd76db3a23a5200ad40f2bfe057e2333a2f2d1a75f862b79955f58f
MD5 72dd5ee3113095dda2a9d7a6982f9b6a
BLAKE2b-256 892bb47880ed1c343534f9836810155e870e4b02fe40c6cf6a8967f5420ad303

See more details on using hashes here.

Provenance

The following attestation bundles were made for mf4_rs-3.3.0-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:

Publisher: release.yml on dmagyar-0/mf4-rs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mf4_rs-3.3.0-cp38-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for mf4_rs-3.3.0-cp38-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 442fba443239ae1c12de41d66404bdd022c3a9ef2a3adcd8dfba31d3eb07566c
MD5 5de86e165d19f1a7e972d9efbc3f5da0
BLAKE2b-256 c5516c7d30e00f6bb864d955b0b316e114b876ca83ce72bf8a59a17150b4b94e

See more details on using hashes here.

Provenance

The following attestation bundles were made for mf4_rs-3.3.0-cp38-abi3-macosx_11_0_arm64.whl:

Publisher: release.yml on dmagyar-0/mf4-rs

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page