Skip to main content

Imputify

A Python library for evaluating and performing missing data imputation. It measures imputation quality across three dimensions: reconstruction (how close are imputed values to the truth?), distribution preservation (are statistical properties maintained?), and predictive utility (can downstream models still perform well?).

The library is fully compatible with scikit-learn's fit/transform API and provides ready-to-use imputers: KNN, statistical baselines, autoencoders (DAE, VAE), GAIN, and a decoder-only LLM fine-tuned for tabular imputation.

This library is part of my master's research proposal, so apart from scikit-learn compatibility, expect breaking changes. The API will stabilize as the research progresses.


Why missingness matters

Not all missing data is created equal. The mechanism behind missingness determines which imputation methods will work. MCAR (Missing Completely at Random) is the easy case, values disappear randomly with no pattern, like a sensor failing at random times. MAR (Missing at Random) is trickier, missingness depends on other observed values, like high earners being more likely to skip income questions. MNAR (Missing Not at Random) is the hardest, missingness depends on the missing value itself, like very sick patients being unable to complete health surveys.

Most imputation methods assume MCAR or MAR. MNAR breaks these assumptions because the data you're trying to recover is exactly what's causing it to be missing. This is where I think LLMs might help, they can learn complex conditional distributions from the observed data and extrapolate patterns that simpler methods miss.

Evaluation

A good imputation isn't just "close to the true value". Imputify measures quality from three complementary perspectives:

Reconstruction, point-wise accuracy, with MAE, RMSE, NRMSE for numerical features and accuracy for categorical features.

Distribution, statistical properties, as Wasserstein distance, KS statistic, KL divergence (how much distributions shifted), as well as Correlation shift (did we break relationships between variables?)

Predictive utility, downstream impact, by training a model on original vs imputed data and compare the performance gap.

Predictive metrics:

Classification: accuracy, precision, recall, F1

Regression: R², MAE, RMSE

The overall score combines these into a single number in [0, 1]. Reconstruction and distribution are normalized as 1 / (1 + error), predictive as 1 - |Δmetrics|. The final score is the mean of all three.

Imputers

Imputer Category Description
StatisticalImputer Baseline Mean/median for numerical, mode for categorical
KNNImputer Baseline k-nearest neighbors
MICEImputer Baseline Multiple Imputation by Chained Equations
MissForestImputer Baseline Random Forest-based iterative imputation
XGBoostImputer Baseline XGBoost-based iterative imputation
DAEImputer Deep Learning Denoising AutoEncoder with swap noise
VAEImputer Deep Learning Variational AutoEncoder (probabilistic latent space)
GAINImputer Deep Learning Generative Adversarial Imputation Nets
DecoderOnlyImputer LLM Fine-tuned decoder-only transformer via structured JSON serialization

Example

from imputify.imputer import DAEImputer
from imputify.missing import introduce_missing, PatternConfig
from imputify.metrics import evaluate

# Create realistic missing data (MNAR pattern)
pattern = PatternConfig(incomplete_vars=['income'], mechanism='MNAR')
X_missing, mask = introduce_missing(X, proportion=0.3, patterns=[pattern])

# Impute
imputer = DAEImputer(hidden_dim=128, epochs=100)
X_imputed = imputer.fit_transform(X_missing)

# Evaluate across all dimensions
results = evaluate(X_original, X_imputed, mask, y=y)
print(f"Overall score: {results.overall_score:.3f}")

Installation

If you don't have uv installed, do yourself a favor and:

# Linux & macOS
curl -LsSf https://astral.sh/uv/install.sh | sh

# macOS (Homebrew)
brew install uv

# Windows
# Well, check their installation page: https://docs.astral.sh/uv/getting-started/installation/

Once installed, simply clone the repo and run uv sync to install dependencies:

git clone https://github.com/gabfssilva/imputify
cd imputify
uv sync

Requires Python 3.12+.

Open the project using your favorite IDE and that's it.

License

MIT

Metadata

Release files for imputify 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for imputify 0.1.0
File Size Uploaded
imputify-0.1.0.tar.gz 4.6 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for imputify 0.1.0
File Interpreter ABI Platform
imputify-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 4.6 MB

Release files / imputify-0.1.0.tar.gz

Download URL imputify-0.1.0.tar.gz
Size 4.6 MB
Tags Source
SHA-256 checksum
How to use checksums
7d0f5c483952a4fe6b9075a4d00e0ea88b5ba72acd671aa056a3e674ee4f08d5
BLAKE2b-256 checksum
How to use checksums
f79a83c249d63502db38e8dfe84ee9c605b3b3b0447d369a8b64ede5924b1963
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Feb 10, 2026.

Transparency log

Release files / imputify-0.1.0-py3-none-any.whl

Download URL imputify-0.1.0-py3-none-any.whl
Size 55.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2ce2e751f5ce3e9c1d407a9ceae8616e5693a6e3e548d5cbba9ac575ae1518e8
BLAKE2b-256 checksum
How to use checksums
de9ebdfb6b97cceb75734f7a61c64179709c3a5b838f1ab3a5fae434ab6466d3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Feb 10, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page