Skip to main content

Transformer of tabular data into Integer, Float and Categorical (IFC) variables with missing data imputation.

Project description

ifcfill

ifcfill is a Python library for preparing tabular data for synthetic-data generation. It transforms real tabular data into Integer, Float, and Categorical (IFC) variables with fast NumPy-powered missing data imputation, then applies the same learned inverse transformation to generated synthetic data so it can be mapped back to the original table structure.

PyPI version Python versions License Documentation


Purpose

ifcfill is designed to sit around tabular synthetic-data generators:

  1. Fit IFCTransformer on real data.
  2. Transform the real data into generator-friendly IFC variables.
  3. Train or run any tabular generator on the transformed data.
  4. Apply inverse_transform() to the synthetic output using the mappings learned from the real data.

If inverse transformation happens later or on another machine, save the fitted transformation state with save() and load it back with IFCTransformer.load().

The package is intentionally unsupervised: it does not require or model a target variable.


Features

  • Flexible input — accepts a pandas.DataFrame or a path to a CSV file
  • Automatic type inference — detects integer, float, categorical, and datetime columns automatically; user overrides supported per column
  • Configurable imputation — choose fill strategy independently for each type:
    • Integer: mean, median, mode, zero
    • Float: mean, median, mode, zero
    • Categorical: constant (default), mode
  • Categorical missingness as a category — missing categorical values are transformed into a learnable category and converted back to missing values during inverse_transform()
  • Namespaced missing sentinel — the default categorical missing category is __ifcfill_missing__ to reduce collisions with real values
  • Optional categorical label encoding — fill categorical values first, then encode categories as integer codes through a separate label-encoding layer with inverse mappings
  • Datetime → integer conversion — converts date/time columns to integers relative to a configurable anchor date and time unit (days, seconds, ms, …)
  • Constant column removal — automatically drops true constant columns while preserving categorical missing categories when they are learnable
  • Missing value tracking — records the count and fraction of missing values per column at fit time, accessible via missing_report_
  • Full transformation bookkeepinginverse_transform() restores dropped constants, original column order, and optionally re-introduces missing values at the original rate
  • Portable fitted state — save all learned transformations to JSON and load them later for consistent transform/inverse-transform workflows on another machine
  • Synthetic-data workflow support — apply one inverse transformation consistently to both transformed real data and generated synthetic data

Installation

pip install ifcfill

Quick Start

import pandas as pd
from ifcfill import IFCTransformer

df = pd.DataFrame({
    "age":    [25, 30, None, 40],
    "salary": [50_000.5, None, 75_000.0, 90_000.25],
    "city":   ["London", None, "Paris", "London"],
    "joined": pd.to_datetime(["2020-01-01", "2021-06-15", None, "2023-03-10"]),
    "flag":   ["yes", "yes", "yes", "yes"],   # constant → will be dropped
})

tf = IFCTransformer(
    int_fill="median",
    float_fill="mean",
    datetime_anchor="1970-01-01",
    datetime_unit="D",
)

transformed = tf.fit_transform(df)
print(transformed)

# Inspect missing-value distribution captured at fit time
print(tf.missing_report_)

# Restore original structure.
# Categorical missing categories are converted back to missing values.
restored = tf.inverse_transform(transformed, restore_missing=True, random_state=42)
print(restored)

# Save everything needed to transform or inverse-transform later.
tf.save("ifcfill-state.json")

# On another machine or in another process:
loaded_tf = IFCTransformer.load("ifcfill-state.json")
restored_again = loaded_tf.inverse_transform(transformed)

From a CSV file

transformed = IFCTransformer().fit_transform("data.csv")

Parallel column processing

# Use all available CPUs for per-column fit/transform work
tf = IFCTransformer(n_jobs=-1)
transformed = tf.fit_transform(df)

Override column types

tf = IFCTransformer(col_types={"age": "categorical", "score": "float"})
transformed = tf.fit_transform(df)

Label encoding for generators

tf = IFCTransformer(cat_encoding="label")
transformed = tf.fit_transform(df)

# Safe copies of learned category/code dictionaries
print(tf.get_category_mappings())

restored = tf.inverse_transform(transformed)

Documentation

Full documentation including the API reference is available at: https://eulerlettersai.github.io/ifcfill


License

Apache 2.0 — see LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

ifcfill-0.3.3.tar.gz (22.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

ifcfill-0.3.3-py3-none-any.whl (18.8 kB view details)

Uploaded Python 3

File details

Details for the file ifcfill-0.3.3.tar.gz.

File metadata

  • Download URL: ifcfill-0.3.3.tar.gz
  • Upload date:
  • Size: 22.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.6

File hashes

Hashes for ifcfill-0.3.3.tar.gz
Algorithm Hash digest
SHA256 404ebdcad5b90bebd463c53322f8f9dd1b03e66ec398706486e889a00eb59a18
MD5 6f8adffafcd1a638347ec4d70cddb94f
BLAKE2b-256 61cdb75798747c109d24a30d5b3236d5915e2862b057be7ce044a390aa2fb909

See more details on using hashes here.

File details

Details for the file ifcfill-0.3.3-py3-none-any.whl.

File metadata

  • Download URL: ifcfill-0.3.3-py3-none-any.whl
  • Upload date:
  • Size: 18.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.6

File hashes

Hashes for ifcfill-0.3.3-py3-none-any.whl
Algorithm Hash digest
SHA256 e42e09bf69b6fedb9f289cf5fc20753a0e42e27153761aa295d701a9fbc0bd55
MD5 dcddf1a076d9f9422e8d2a1b7cf13f70
BLAKE2b-256 bb9e482fe7f5485d0ee9f9013971c924369ac70ad143ca2fa2502b45bdf124f0

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page