Skip to main content

Transformer of tabular data into Integer, Float and Categorical (IFC) variables with missing data imputation.

Project description

ifcfill

ifcfill is a Python library for preparing tabular data for synthetic-data generation. It transforms real tabular data into Integer, Float, and Categorical (IFC) variables with fast NumPy-powered missing data imputation, then applies the same learned inverse transformation to generated synthetic data so it can be mapped back to the original table structure.

PyPI version Python versions License Documentation


Purpose

ifcfill is designed to sit around tabular synthetic-data generators:

  1. Fit IFCTransformer on real data.
  2. Transform the real data into generator-friendly IFC variables.
  3. Train or run any tabular generator on the transformed data.
  4. Apply inverse_transform() to the synthetic output using the mappings learned from the real data.

If inverse transformation happens later or on another machine, save the fitted transformation state with save() and load it back with IFCTransformer.load().

The package is intentionally unsupervised: it does not require or model a target variable.


Features

  • Flexible input — accepts a pandas.DataFrame or a path to a CSV file
  • Automatic type inference — detects integer, float, categorical, and datetime columns automatically; user overrides supported per column
  • Configurable imputation — choose fill strategy independently for each type:
    • Integer: mean, median, mode, zero
    • Float: mean, median, mode, zero
    • Categorical: constant (default), mode
  • Categorical missingness as a category — missing categorical values are transformed into a learnable category and converted back to missing values during inverse_transform()
  • Namespaced missing sentinel — the default categorical missing category is __ifcfill_missing__ to reduce collisions with real values
  • Optional categorical label encoding — fill categorical values first, then encode categories as integer codes through a separate label-encoding layer with inverse mappings
  • Datetime → integer conversion — converts date/time columns to integers relative to a configurable anchor date and time unit (days, seconds, ms, …)
  • Constant column removal — automatically drops true constant columns while preserving categorical missing categories when they are learnable
  • Missing value tracking — records the count and fraction of missing values per column at fit time, accessible via missing_report_
  • Full transformation bookkeepinginverse_transform() restores dropped constants, original column order, and optionally re-introduces missing values at the original rate
  • Portable fitted state — save all learned transformations to JSON and load them later for consistent transform/inverse-transform workflows on another machine
  • Synthetic-data workflow support — apply one inverse transformation consistently to both transformed real data and generated synthetic data

Installation

pip install ifcfill

Quick Start

import pandas as pd
from ifcfill import IFCTransformer

df = pd.DataFrame({
    "age":    [25, 30, None, 40],
    "salary": [50_000.5, None, 75_000.0, 90_000.25],
    "city":   ["London", None, "Paris", "London"],
    "joined": pd.to_datetime(["2020-01-01", "2021-06-15", None, "2023-03-10"]),
    "flag":   ["yes", "yes", "yes", "yes"],   # constant → will be dropped
})

tf = IFCTransformer(
    int_fill="median",
    float_fill="mean",
    datetime_anchor="1970-01-01",
    datetime_unit="D",
)

transformed = tf.fit_transform(df)
print(transformed)

# Inspect missing-value distribution captured at fit time
print(tf.missing_report_)

# Restore original structure.
# Categorical missing categories are converted back to missing values.
restored = tf.inverse_transform(transformed, restore_missing=True, random_state=42)
print(restored)

# Save everything needed to transform or inverse-transform later.
tf.save("ifcfill-state.json")

# On another machine or in another process:
loaded_tf = IFCTransformer.load("ifcfill-state.json")
restored_again = loaded_tf.inverse_transform(transformed)

From a CSV file

transformed = IFCTransformer().fit_transform("data.csv")

Parallel column processing

# Use all available CPUs for per-column fit/transform work
tf = IFCTransformer(n_jobs=-1)
transformed = tf.fit_transform(df)

Override column types

tf = IFCTransformer(col_types={"age": "categorical", "score": "float"})
transformed = tf.fit_transform(df)

Label encoding for generators

tf = IFCTransformer(cat_encoding="label")
transformed = tf.fit_transform(df)

# Safe copies of learned category/code dictionaries
print(tf.get_category_mappings())

restored = tf.inverse_transform(transformed)

Documentation

Full documentation including the API reference is available at: https://eulerlettersai.github.io/ifcfill


License

Apache 2.0 — see LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

ifcfill-0.3.4.tar.gz (23.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

ifcfill-0.3.4-py3-none-any.whl (18.8 kB view details)

Uploaded Python 3

File details

Details for the file ifcfill-0.3.4.tar.gz.

File metadata

  • Download URL: ifcfill-0.3.4.tar.gz
  • Upload date:
  • Size: 23.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.15

File hashes

Hashes for ifcfill-0.3.4.tar.gz
Algorithm Hash digest
SHA256 bb7af23f077adc245a4437f1b31b26c37911882c8374f612fc963f0afe400127
MD5 7cb21b97b0acad1eba4bca80c0bb6705
BLAKE2b-256 a95cc4a67f83f963de010d6c2f7afbf85853bcf893df97e97e423b009a4cdd50

See more details on using hashes here.

File details

Details for the file ifcfill-0.3.4-py3-none-any.whl.

File metadata

  • Download URL: ifcfill-0.3.4-py3-none-any.whl
  • Upload date:
  • Size: 18.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.15

File hashes

Hashes for ifcfill-0.3.4-py3-none-any.whl
Algorithm Hash digest
SHA256 968cd6e065ff01f2a62cc0d37d61fa42100c3a986a0dbbd65a9797f2e1de13a8
MD5 57714914d067c7c8a51ef321ca37fcf2
BLAKE2b-256 a6428e1d4a8ead86a0f8f3e9730feaba28d13fd70a00e8c8dc5a6cc3cb0326fe

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page