Skip to main content
nanuq

A reusable Machine Learning toolkit for the full data science lifecycle. From raw data to a deployed model — loading, exploration, cleaning, encoding, modeling, interpretation, and production persistence, all in one coherent library.

Python 3.11+ License MIT Tests Code style: ruff


Why nanuq?

Most data science projects re-implement the same plumbing over and over: loading files, checking for missing values, encoding categories, comparing a few models, tuning a decision threshold, then saving the result for production. That boilerplate is rarely the interesting part, and rewriting it each time invites subtle bugs and inconsistencies.

nanuq packages those recurring steps into a single, well-tested library you can install and reuse across projects. It is intentionally not a black-box AutoML tool: every function does one clear thing, and when several approaches exist for a task (imputation, scaling, handling class imbalance, encoding), nanuq offers both the atomic functions and a comparator that runs them side by side so you can make an informed choice.

Key features

  • Full lifecycle coverage — 16 modules spanning loading, EDA, cleaning, encoding, scaling, transformations, feature engineering, balancing, modeling, evaluation, interpretation, plotting, reporting, and persistence.
  • Comparator pattern — for tasks with multiple valid strategies, compare them on your data in a single call instead of guessing.
  • Production-ready persistence — save a model and its metadata as separate, versionable artifacts, with SHA256 integrity checks and custom-threshold inference.
  • No surprises — no in-place mutation of your DataFrames, no hidden global state, explicit return values.
  • Documented and tested — systematic docstrings, runnable examples, 161 tests.

Installation

pip install nanuq

Or with uv:

uv pip install nanuq

Requires Python 3.11 or newer.

Quick start

import nanuq

# 1. Load and explore
df = nanuq.load_csv("data.csv")
nanuq.dataset_overview(df)

# 2. Rank features by their association with the target
ranking = nanuq.rank_features_by_target_association(df, target="churn")

# 3. Train and evaluate a model
from sklearn.ensemble import GradientBoostingClassifier
model = GradientBoostingClassifier().fit(X_train, y_train)
nanuq.evaluate_classifier(model, X_train, y_train, X_test, y_test)

# 4. Save for production (model + metadata, kept separate)
meta = nanuq.build_metadata(model, X_train.columns.tolist(), threshold=0.29)
nanuq.save_artifacts(model, meta, "artifacts/")

# 5. Reload and predict with a custom decision threshold
model, meta = nanuq.load_artifacts("artifacts/")
preds = nanuq.predict_with_threshold(
    model, X_new,
    threshold=meta["threshold"],
    feature_order=meta["feature_order"],
)

Philosophy

  • Reusable. Designed to be installed as a dependency and shared across projects, not copy-pasted between notebooks.
  • Comparative. When several methods exist for the same task, nanuq exposes the individual functions and a comparator that puts them side by side, so the choice is data-driven rather than habitual.
  • Educational. Readable code, systematic docstrings, and worked examples — the library is meant to be learned from, not just called.
  • No magic. No side effects, no in-place DataFrame mutation, predictable outputs you can reason about.

Modules

Module Purpose
loading Read CSV / Excel / JSON / Parquet / SQL into DataFrames
exploration Dataset overview, column typing, quick diagnostics
cleaning Missing-value imputation (mean, median, mode) with reports
outliers Outlier detection (Z-score, IQR) and a comparator
encoding One-Hot, Label, Ordinal, binary encoders and a comparator
scaling MinMax, Standard scaling and a comparator
transformations Log, sqrt, Box-Cox, quantile, power transforms
feature_engineering Polynomial features, interactions, discretization
balancing class_weight, SMOTE, undersampling for imbalanced data
bivariate Chi², Mann-Whitney, Pearson/Spearman correlations
models Ready-to-use linear, tree, boosted, and clustering pipelines
evaluation Metrics, stratified CV, precision-recall curve, optimal threshold
interpretation SHAP, permutation importance, native feature importance
persistence Save/load models and metadata for production
plots Boxplots, confusion matrices, importance charts
reporting Automated EDA report generation

The comparator pattern

Rather than committing to a single model or transformation up front, nanuq lets you put options side by side and decide from the results.

# Benchmark several models on the same train/test split
results = nanuq.benchmark_models(
    {
        "logistic": LogisticRegression(),
        "random_forest": RandomForestClassifier(),
        "gradient_boosting": GradientBoostingClassifier(),
    },
    X_train, y_train, X_test, y_test,
    task="classification",
)
#   -> a DataFrame ranking each model by its metrics

# Compare a feature's distribution before and after a transformation
nanuq.compare_distribution(df["income"], nanuq.log_transform(df["income"]))

# Compare a feature's mean across target classes (quick signal check)
nanuq.compare_means_by_class(df, feature="overtime_hours", target="churn")

For atomic preprocessing you also get clear, single-purpose functions: scale_minmax, scale_standard, one_hot_encode, label_encode_column, detect_outliers_iqr, and more — no hidden defaults.

Production with persistence

The persistence module follows a deliberate convention: keep the binary model and its human-readable metadata in separate files.

from nanuq import build_metadata, save_artifacts, load_artifacts, predict_with_threshold

# Save: two separate files
meta = build_metadata(
    model, X.columns.tolist(), threshold=0.286,
    metrics={"auc": 0.83, "f1": 0.54},
)
save_artifacts(model, meta, "artifacts/")
#   artifacts/model.joblib   -> the sklearn estimator (binary)
#   artifacts/metadata.json  -> threshold, feature order, metrics (readable)

# Load + predict with the stored threshold (hash is verified automatically)
model, meta = load_artifacts("artifacts/")
preds = predict_with_threshold(
    model, X_new,
    threshold=meta["threshold"],
    feature_order=meta["feature_order"],
)

Why separate the two? Because it lets you:

  • inspect metrics and configuration without loading scikit-learn (plain JSON),
  • version the metadata in Git while the model stays binary,
  • adjust the decision threshold without retraining or touching the model,
  • detect a corrupted or swapped model file via its SHA256 hash.

Threshold optimization

On imbalanced datasets, the default 0.5 decision threshold is rarely optimal. nanuq's evaluate_classifier reports the full set of metrics (including the precision-recall trade-off) so you can pick the threshold that best fits your objective — often dramatically improving recall on the minority class.

# Full evaluation on train AND test, with the metrics you need to choose a threshold
metrics = nanuq.evaluate_classifier(model, X_train, y_train, X_test, y_test)

# Once chosen, the threshold travels with the model in metadata.json,
# and inference applies it consistently:
preds = nanuq.predict_with_threshold(
    model, X_new,
    threshold=meta["threshold"],   # e.g. 0.286 instead of 0.5
    feature_order=meta["feature_order"],
)

Storing the threshold alongside the model keeps inference consistent with the way the model was tuned.

Development

# Install in editable mode with the dev tools
uv pip install -e ".[dev]"

# Run the test suite
pytest

# Lint
ruff check src/ tests/

License

MIT — © 2026 Olivier Gruwe


nanuq means "polar bear" in Inuktitut. 🐻‍❄️

Release files for nanuq 0.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for nanuq 0.2.1
File Size Uploaded
nanuq-0.2.1.tar.gz 124.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for nanuq 0.2.1
File Interpreter ABI Platform
nanuq-0.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 191.2 kB

Release files / nanuq-0.2.1.tar.gz

Download URL nanuq-0.2.1.tar.gz
Size 124.0 kB
Tags Source
SHA-256 checksum
How to use checksums
8272c8d1ca7ee95839f20491bdee699e21e0ac27aa7ea9fa0ebb8501dc21d10b
BLAKE2b-256 checksum
How to use checksums
24b8a49a338db0859077a70faae367db931f07b2e0d799d272b5d66dc2375aaf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.11.19 {"installer":{"name":"uv","version":"0.11.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / nanuq-0.2.1-py3-none-any.whl

Download URL nanuq-0.2.1-py3-none-any.whl
Size 67.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
309d6b1c08a2030f53e18aba992bc8d65c3a1f57d3e26ea9651790b6072ea487
BLAKE2b-256 checksum
How to use checksums
2364ebdace138f4268f4df8c900f97beda5eeef217726a4571dc2160851a1d7a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.11.19 {"installer":{"name":"uv","version":"0.11.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

0.2.1 This release

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page