Skip to main content

Transformer of tabular data into Integer, Float and Categorical (IFC) variables with missing data imputation.

Project description

ifcfill

ifcfill is a Python library for preparing tabular data for synthetic-data generation. It transforms real tabular data into Integer, Float, and Categorical (IFC) variables with fast NumPy-powered missing data imputation, then applies the same learned inverse transformation to generated synthetic data so it can be mapped back to the original table structure.

PyPI version Python versions License Documentation


Purpose

ifcfill is designed to sit around tabular synthetic-data generators:

  1. Fit IFCTransformer on real data.
  2. Transform the real data into generator-friendly IFC variables.
  3. Train or run any tabular generator on the transformed data.
  4. Apply inverse_transform() to the synthetic output using the mappings learned from the real data.

If inverse transformation happens later or on another machine, save the fitted transformation state with save() and load it back with IFCTransformer.load().

The package is intentionally unsupervised: it does not require or model a target variable.


Features

  • Flexible input — accepts a pandas.DataFrame or a path to a CSV file
  • Automatic type inference — detects integer, float, categorical, and datetime columns automatically; user overrides supported per column
  • Configurable imputation — choose fill strategy independently for each type:
    • Integer: mean, median, mode, zero
    • Float: mean, median, mode, zero
    • Categorical: constant (default), mode
  • Categorical missingness as a category — missing categorical values are transformed into a learnable category and converted back to missing values during inverse_transform()
  • Namespaced missing sentinel — the default categorical missing category is __ifcfill_missing__ to reduce collisions with real values
  • Optional categorical label encoding — fill categorical values first, then encode categories as integer codes through a separate label-encoding layer with inverse mappings
  • Datetime → integer conversion — converts date/time columns to integers relative to a configurable anchor date and time unit (days, seconds, ms, …)
  • Constant column removal — automatically drops true constant columns while preserving categorical missing categories when they are learnable
  • Missing value tracking — records the count and fraction of missing values per column at fit time, accessible via missing_report_
  • Full transformation bookkeepinginverse_transform() restores dropped constants, original column order, and optionally re-introduces missing values at the original rate
  • Portable fitted state — save all learned transformations to JSON and load them later for consistent transform/inverse-transform workflows on another machine
  • Synthetic-data workflow support — apply one inverse transformation consistently to both transformed real data and generated synthetic data

Installation

pip install ifcfill

Quick Start

import pandas as pd
from ifcfill import IFCTransformer

df = pd.DataFrame({
    "age":    [25, 30, None, 40],
    "salary": [50_000.5, None, 75_000.0, 90_000.25],
    "city":   ["London", None, "Paris", "London"],
    "joined": pd.to_datetime(["2020-01-01", "2021-06-15", None, "2023-03-10"]),
    "flag":   ["yes", "yes", "yes", "yes"],   # constant → will be dropped
})

tf = IFCTransformer(
    int_fill="median",
    float_fill="mean",
    datetime_anchor="1970-01-01",
    datetime_unit="D",
)

transformed = tf.fit_transform(df)
print(transformed)

# Inspect missing-value distribution captured at fit time
print(tf.missing_report_)

# Restore original structure.
# Categorical missing categories are converted back to missing values.
restored = tf.inverse_transform(transformed, restore_missing=True, random_state=42)
print(restored)

# Save everything needed to transform or inverse-transform later.
tf.save("ifcfill-state.json")

# On another machine or in another process:
loaded_tf = IFCTransformer.load("ifcfill-state.json")
restored_again = loaded_tf.inverse_transform(transformed)

From a CSV file

transformed = IFCTransformer().fit_transform("data.csv")

Parallel column processing

# Use all available CPUs for per-column fit/transform work
tf = IFCTransformer(n_jobs=-1)
transformed = tf.fit_transform(df)

Override column types

tf = IFCTransformer(col_types={"age": "categorical", "score": "float"})
transformed = tf.fit_transform(df)

Label encoding for generators

tf = IFCTransformer(cat_encoding="label")
transformed = tf.fit_transform(df)

# Safe copies of learned category/code dictionaries
print(tf.get_category_mappings())

restored = tf.inverse_transform(transformed)

Documentation

Full documentation including the API reference is available at: https://eulerlettersai.github.io/ifcfill


License

Apache 2.0 — see LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

ifcfill-0.3.2.tar.gz (22.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

ifcfill-0.3.2-py3-none-any.whl (18.5 kB view details)

Uploaded Python 3

File details

Details for the file ifcfill-0.3.2.tar.gz.

File metadata

  • Download URL: ifcfill-0.3.2.tar.gz
  • Upload date:
  • Size: 22.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.6

File hashes

Hashes for ifcfill-0.3.2.tar.gz
Algorithm Hash digest
SHA256 7cee0234f588f48b8a6744f444944db929fca5cfa9f20096672b21b93fb87209
MD5 e9dff95dda842034fb6df98c335fea10
BLAKE2b-256 a9a9b5122db096e99ade1ef6bd9bf88fb6ac258a35ef7509c0f567c70012fd27

See more details on using hashes here.

File details

Details for the file ifcfill-0.3.2-py3-none-any.whl.

File metadata

  • Download URL: ifcfill-0.3.2-py3-none-any.whl
  • Upload date:
  • Size: 18.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.6

File hashes

Hashes for ifcfill-0.3.2-py3-none-any.whl
Algorithm Hash digest
SHA256 fd51ac0bfc881cd69306fdadda473f8ceb410f6178aa9ffa1fe9fd969562bd13
MD5 94c43b65d75381750e927248f3817dc6
BLAKE2b-256 2a537cd6638fc0238c9f0fc4fc70982f1b7db2037957ffae6ba459cbc2e5348c

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page