Skip to main content

Transformer of tabular data into Integer, Float and Categorical (IFC) variables with missing data imputation.

Project description

ifcfill

ifcfill is a Python library for preparing tabular data for synthetic-data generation. It transforms real tabular data into Integer, Float, and Categorical (IFC) variables with fast NumPy-powered missing data imputation, then applies the same learned inverse transformation to generated synthetic data so it can be mapped back to the original table structure.

PyPI version Python versions License Documentation


Purpose

ifcfill is designed to sit around tabular synthetic-data generators:

  1. Fit IFCTransformer on real data.
  2. Transform the real data into generator-friendly IFC variables.
  3. Train or run any tabular generator on the transformed data.
  4. Apply inverse_transform() to the synthetic output using the mappings learned from the real data.

If inverse transformation happens later or on another machine, save the fitted transformation state with save() and load it back with IFCTransformer.load().

The package is intentionally unsupervised: it does not require or model a target variable.


Features

  • Flexible input — accepts a pandas.DataFrame or a path to a CSV file
  • Automatic type inference — detects integer, float, categorical, and datetime columns automatically; user overrides supported per column
  • Configurable imputation — choose fill strategy independently for each type:
    • Integer: mean, median, mode, zero
    • Float: mean, median, mode, zero
    • Categorical: constant (default), mode
  • Categorical missingness as a category — missing categorical values are transformed into a learnable category and converted back to missing values during inverse_transform()
  • Namespaced missing sentinel — the default categorical missing category is __ifcfill_missing__ to reduce collisions with real values
  • Optional categorical label encoding — fill categorical values first, then encode categories as integer codes through a separate label-encoding layer with inverse mappings
  • Datetime → integer conversion — converts date/time columns to integers relative to a configurable anchor date and time unit (days, seconds, ms, …)
  • Constant column removal — automatically drops true constant columns while preserving categorical missing categories when they are learnable
  • Missing value tracking — records the count and fraction of missing values per column at fit time, accessible via missing_report_
  • Full transformation bookkeepinginverse_transform() restores dropped constants, original column order, and optionally re-introduces missing values at the original rate
  • Portable fitted state — save all learned transformations to JSON and load them later for consistent transform/inverse-transform workflows on another machine
  • Synthetic-data workflow support — apply one inverse transformation consistently to both transformed real data and generated synthetic data

Installation

pip install ifcfill

Quick Start

import pandas as pd
from ifcfill import IFCTransformer

df = pd.DataFrame({
    "age":    [25, 30, None, 40],
    "salary": [50_000.5, None, 75_000.0, 90_000.25],
    "city":   ["London", None, "Paris", "London"],
    "joined": pd.to_datetime(["2020-01-01", "2021-06-15", None, "2023-03-10"]),
    "flag":   ["yes", "yes", "yes", "yes"],   # constant → will be dropped
})

tf = IFCTransformer(
    int_fill="median",
    float_fill="mean",
    datetime_anchor="1970-01-01",
    datetime_unit="D",
)

transformed = tf.fit_transform(df)
print(transformed)

# Inspect missing-value distribution captured at fit time
print(tf.missing_report_)

# Restore original structure.
# Categorical missing categories are converted back to missing values.
restored = tf.inverse_transform(transformed, restore_missing=True, random_state=42)
print(restored)

# Save everything needed to transform or inverse-transform later.
tf.save("ifcfill-state.json")

# On another machine or in another process:
loaded_tf = IFCTransformer.load("ifcfill-state.json")
restored_again = loaded_tf.inverse_transform(transformed)

From a CSV file

transformed = IFCTransformer().fit_transform("data.csv")

Parallel column processing

# Use all available CPUs for per-column fit/transform work
tf = IFCTransformer(n_jobs=-1)
transformed = tf.fit_transform(df)

Override column types

tf = IFCTransformer(col_types={"age": "categorical", "score": "float"})
transformed = tf.fit_transform(df)

Label encoding for generators

tf = IFCTransformer(cat_encoding="label")
transformed = tf.fit_transform(df)

# Safe copies of learned category/code dictionaries
print(tf.get_category_mappings())

restored = tf.inverse_transform(transformed)

Documentation

Full documentation including the API reference is available at: https://eulerlettersai.github.io/ifcfill


License

Apache 2.0 — see LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

ifcfill-0.3.6.tar.gz (22.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

ifcfill-0.3.6-py3-none-any.whl (18.8 kB view details)

Uploaded Python 3

File details

Details for the file ifcfill-0.3.6.tar.gz.

File metadata

  • Download URL: ifcfill-0.3.6.tar.gz
  • Upload date:
  • Size: 22.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.15

File hashes

Hashes for ifcfill-0.3.6.tar.gz
Algorithm Hash digest
SHA256 37be53bd31b44f9c8aec81e3c90f7dd36d58864d1143105d24054bb7dff58dd3
MD5 fb003090eade8ef92510d6cb704a0e8e
BLAKE2b-256 8bcd1d6ad71a427dc63a86b761a4f4f07a153ad3ea953d06aec4c1dbb21491df

See more details on using hashes here.

File details

Details for the file ifcfill-0.3.6-py3-none-any.whl.

File metadata

  • Download URL: ifcfill-0.3.6-py3-none-any.whl
  • Upload date:
  • Size: 18.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.15

File hashes

Hashes for ifcfill-0.3.6-py3-none-any.whl
Algorithm Hash digest
SHA256 1c0215cd3ed184f9f371c0f7050350a5b1075ae97bb3a4ef1d8c0e20eabece40
MD5 ca370f6004afa1c7bee3526fcbe7cef8
BLAKE2b-256 fdfde65b1b4bafe417ad40c7298319966d0b7abde61d3293b866613216c39f35

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page