🧩 Module: light_mlt.py
Lightweight, reproducible preprocessing pipeline for tabular datasets — integrating categorical encoding via Modular Linear Tokenization (MLT) and continuous feature scaling.
Designed for efficient fit/transform workflows with full reversibility and schema persistence.
📘 Reference
Schmitz, T. (2025). light-mlt: Modular Linear Tokenization for Scalable Categorical Encoding.
DOI: 10.5281/zenodo.17467914
🔖 Key Features
- Deterministic preprocessing for categorical and continuous data
- Append-only vocabulary management (
cats_map.pkl) - Fully reversible categorical encoding via MLT
- Continuous feature scaling with
StandardScaler - Persistent schema for consistent transformations
- Minimal dependencies:
numpy,pandas,scikit-learn
📦 Artifacts
Each fit() operation generates or updates the following artifacts (default directory: light_mlt_artifacts/):
| File | Description |
|---|---|
schema.pkl |
Schema metadata (columns, types, MLT config) |
scaler.pkl |
Trained StandardScaler for continuous columns |
cats_map.pkl |
Append-only {label → id} mapping for categorical features |
mlt_params.pkl |
MLT parameters (p, n, M, Minv) per column |
preprocessed.csv |
Optional transformed dataset export |
🧪 Detailed Examples
Below are complete examples demonstrating how to use light_mlt.py for preprocessing, transformation, and reversibility.
1️⃣ Basic Usage
import pandas as pd
from light_mlt import fit_transform, inverse_transform
# Example dataset
df = pd.DataFrame({
"city": ["São Paulo", "Curitiba", "São Paulo", "Florianópolis"],z
"vehicle": ["Truck", "Car", "Car", "Bus"],
"mileage": [12.4, 25.8, 31.5, 18.7],
})
print("=== Original Data ===")
print(df)
# --- Step 1: Fit + Transform ---
df_t, path, token_cols, report = fit_transform(
df,
categorical_cols=["city", "vehicle"],
continuous_cols=["mileage"],
dir="light_mlt_artifacts/"
)
print("\n=== Transformed Data ===")
print(df_t.head())
print("\nGenerated token columns:", token_cols)
print("Report:", report)
# --- Step 2: Inverse Transform ---
df_rec = inverse_transform(df_t, dir="light_mlt_artifacts/")
print("\n=== Reconstructed Data ===")
print(df_rec)
Metadata
Release files for light-mlt 0.1.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| light_mlt-0.1.3.tar.gz | 11.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| light_mlt-0.1.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 20.6 kB
Release files / light_mlt-0.1.3.tar.gz
| Download URL | light_mlt-0.1.3.tar.gz |
|---|---|
| Size | 11.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1829b56f5e9a1c0ec22cb19ed298402d4d5328d0f417133edc7c291939eda18a
|
|
BLAKE2b-256 checksum How to use checksums |
90e57905d6fde550e0ebc904611ebda7122d39242bd612b2aa2fcc81961ae43e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.0.1 CPython/3.10.12
|
Release files / light_mlt-0.1.3-py3-none-any.whl
| Download URL | light_mlt-0.1.3-py3-none-any.whl |
|---|---|
| Size | 9.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
50c36fb5b22481d55f2f74313c844436ea528f73ecb8cbb70d72eefd032a1644
|
|
BLAKE2b-256 checksum How to use checksums |
997b9eb68ee74635eb66f03908b8717e21e4c553b3d90d856ba7d144e944c33b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.0.1 CPython/3.10.12
|