Skip to main content

AutoCarver Logo

PyPI Python License SPEC 0 Docs Tests Coverage

AutoCarver in one loop: discretize, rank groupings, carve

AutoCarver automates supervised feature discretization (binning) to maximize statistical association with your target — using Tschuprow's T or Cramér's V — and validates the chosen bins against a held-out dev set. It supports binary classification, multiclass classification, and regression, and is widely used for credit scoring, fraud detection, and risk modeling.

🆕 What's New

📊 Cross-validated robustness. fit now accepts a cv argument for extra held-out robustness views on top of (or instead of) a dev set: carver.fit(X, y, cv=5). Accepts an int, any scikit-learn splitter, or explicit index pairs, resolved via sklearn.model_selection.check_cv — folds veto over-fit combinations but never reorder them (ranks stay anchored to the full train set). See Cross-validation folds.

🤖 LLM & MCP integration. AutoCarver now ships a local Model Context Protocol server: point an MCP-aware assistant (VS Code Copilot, Claude Desktop, Cursor, …) at a data file and let it qualify the columns and carve them against your target through tool calls. The server runs fully on your machine — your dataset is never sent to AutoCarver or any external service (only your own LLM provider sees what the assistant shares). Carving quality depends on the LLM, so have a human confirm the feature definitions before production use. See the LLM & MCP guide.

pip install "autocarver[mcp]"

Install

pip install autocarver

Quick Start

Binary classification on the Titanic dataset:

from pathlib import Path

import pandas as pd
from sklearn.model_selection import train_test_split

from AutoCarver import BinaryCarver, Features

# 1. Load data
url = "https://web.stanford.edu/class/archive/cs/cs109/cs109.1166/stuff/titanic.csv"
data = pd.read_csv(url)
target = "Survived"

# 2. Train / dev split, stratified on the target
train, dev = train_test_split(data, test_size=0.33, random_state=42, stratify=data[target])

# 3. Declare features by type
features = Features(
    categoricals=["Sex"],
    numericals=["Age", "Fare", "Siblings/Spouses Aboard", "Parents/Children Aboard"],
    ordinals={"Pclass": ["1", "2", "3"]},
)

# 4. Fit the carver (dev set drives the robustness checks)
carver = BinaryCarver(features=features, min_freq=0.05, max_n_mod=5)
train_processed = carver.fit_transform(train, train[target], X_dev=dev, y_dev=dev[target])
dev_processed = carver.transform(dev)

# 5. Inspect the carved buckets, target rate, and association
print(carver.summary)

# 6. Persist for later use
carver.save(Path("titanic_carver.json"))
# carver = BinaryCarver.load(Path("titanic_carver.json"))

For multiclass classification use MulticlassCarver (one binning per feature, against the full K-class target) — or OneVsRestCarver for a separate binning per class; for regression use ContinuousCarver — the API is identical. To pre-select features by target association and inter-feature redundancy, pipe the carved output through ClassificationSelector or RegressionSelector.

Why AutoCarver?

  • Optimal supervised binning — exhaustive search over admissible bin combinations maximizes Tschuprow's T (default) or Cramér's V. For fixed min_freq, max_n_mod and metric, no other combination scores higher.
  • Robust to data drift — every candidate bin combination is validated on a dev set, rejecting any whose target rates flip or whose buckets fall below min_freq.
  • First-class ordinal featuresOrdinalDiscretizer enforces your declared modality order, so under-represented levels are merged with their nearest neighbour instead of being collapsed by frequency.
  • Inspect what was carvedfeatures.summary and features.history give you the bin definitions, per-bin target rate / frequency, and the full carving trace right off the fitted carver.
  • Interpretable buckets — human-readable boundaries you can audit, document, and ship to a scorecard.
  • Dimensionality reduction — groups under-represented modalities and caps bins per feature (max_n_mod), which is especially useful before one-hot encoding.
  • Feature pre-selectionClassificationSelector / RegressionSelector rank features by target association and filter on inter-feature correlation.

How does it compare?

AutoCarver optbinning sklearn KBinsDiscretizer
Supervised (uses y) yes yes no
Algorithm exhaustive search over admissible combinations mixed-integer program (CBC) quantile / uniform / k-means
Optimality for given min_freq / max_n_mod / metric guaranteed — best of every admissible combination provably optimal under MIP constraints n/a — no target objective
Target types binary, multiclass, continuous binary, multiclass, continuous n/a
Numeric and categorical and ordinal in one fit yes one binner per feature numeric only
Ordinal features with enforced order yes — OrdinalDiscretizer preserves your declared order via user_splits workaround (loses ordering) no
NaN handled as its own modality yes yes no (raises)
Held-out dev-set robustness check yes — dev set + optional k-fold CV, built into fit no (script CV yourself) no
Per-bin stats + carving history after fit features.summary, features.history binning_table no
JSON round-trip persistence yes (carver.save("...json")) via pickle via pickle
sklearn Pipeline compatible yes yes yes
Feature pre-selection helpers ClassificationSelector, RegressionSelector no no

Side-by-side runnable snippets and a "when to pick which" guide live on the comparison page.

Documentation

Full reference, tutorials, and end-to-end notebook examples on ReadTheDocs.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

autocarver-7.5.2.tar.gz (152.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

autocarver-7.5.2-py3-none-any.whl (204.8 kB view details)

Uploaded Python 3

File details

Details for the file autocarver-7.5.2.tar.gz.

File metadata

  • Download URL: autocarver-7.5.2.tar.gz
  • Upload date:
  • Size: 152.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for autocarver-7.5.2.tar.gz
Algorithm Hash digest
SHA256 8660c8208e2b96a433cecc165a368af1b2f3bedbf749553890a93d6243c86036
MD5 d92e3daa7d9ca60aae12e2d4e99826e7
BLAKE2b-256 3b014722486f7bb8253204cf3898254ef9a2727a0bdade335c6ca24fe9a543f0

See more details on using hashes here.

Provenance

The following attestation bundles were made for autocarver-7.5.2.tar.gz:

Publisher: release.yml on mdefrance/AutoCarver

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file autocarver-7.5.2-py3-none-any.whl.

File metadata

  • Download URL: autocarver-7.5.2-py3-none-any.whl
  • Upload date:
  • Size: 204.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for autocarver-7.5.2-py3-none-any.whl
Algorithm Hash digest
SHA256 ba39e4aba30999770fc47597137a4d1b5ecd26f24afd1412e1ff7a5e64550222
MD5 8e521f31144c790a7e79271a883440ee
BLAKE2b-256 7a64e1d8fe39940d050165bc973e122b5399f2d5eb0c51a704a3c5db9d108a05

See more details on using hashes here.

Provenance

The following attestation bundles were made for autocarver-7.5.2-py3-none-any.whl:

Publisher: release.yml on mdefrance/AutoCarver

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

7.7.0

2 files

7.6.3

2 files

7.6.2

2 files

7.6.1

2 files

7.6.0

2 files

7.5.5

2 files

7.5.4

2 files

7.5.3

2 files

This release

7.5.2 This release

2 files

7.5.1

2 files

7.5.0

2 files

7.4.0

2 files

7.3.9

2 files

7.3.8

2 files

7.3.7

2 files

7.3.6

2 files

7.3.5

2 files

7.3.4

2 files

7.3.3

2 files

7.3.2

2 files

7.3.1

2 files

7.3.0

2 files

7.2.9

2 files

7.2.8

2 files

7.2.7

2 files

7.2.6

2 files

7.2.5

2 files

7.2.2

2 files

7.2.1

2 files

7.2.0

2 files

7.1.11

2 files

7.1.10

2 files

7.1.9

2 files

7.1.8

2 files

7.1.7

2 files

7.1.6

2 files

7.1.5

2 files

7.1.4

2 files

7.1.3

2 files

7.1.2

2 files

7.1.1

2 files

7.1.0

2 files

7.0.14

2 files

7.0.13

2 files

7.0.12

2 files

7.0.10

2 files

7.0.9

2 files

7.0.8

2 files

7.0.7

2 files

7.0.6

2 files

7.0.5

2 files

7.0.4

2 files

7.0.3

2 files

7.0.2

2 files

7.0.1

2 files

7.0.0

2 files

6.0.5

2 files

6.0.4

2 files

6.0.3

2 files

6.0.2

2 files

5.4.9

2 files

5.4.8

2 files

5.4.7

2 files

5.4.6

2 files

5.4.5

2 files

5.4.4

2 files

5.4.3

2 files

5.4.2

2 files

5.4.1

2 files

5.4.0

2 files

5.3.4

2 files

5.3.3

2 files

5.3.2

2 files

5.3.0

2 files

5.2.2

2 files

5.2.1

2 files

5.2.0

2 files

5.1.9

2 files

5.1.8

2 files

5.1.7

2 files

5.1.6

2 files

5.1.5

2 files

5.1.4

2 files

5.1.3

2 files

5.1.2

2 files

5.1.1

2 files

5.1.0

2 files

5.0.9

2 files

5.0.8

2 files

5.0.7

2 files

5.0.6

2 files

5.0.5

2 files

5.0.4

2 files

5.0.3

2 files

5.0.2

2 files

5.0.1

2 files

5.0.0

2 files

4.4.1

2 files

4.4.0

2 files

4.3.2

1 file

4.3.1

1 file

4.3.0

1 file

4.2.1

1 file

4.2.0

1 file

4.1.0

1 file

4.0.1

1 file

4.0.0

1 file

3.1.0

1 file

3.0.12

1 file

3.0.11

1 file

3.0.10

1 file

3.0.9

1 file

3.0.8

1 file

3.0.7

1 file

3.0.6

1 file

3.0.5

1 file

3.0.4

1 file

3.0.3

1 file

3.0.2

1 file

3.0.1

1 file

3.0.0

1 file

2.1.0

1 file

2.0.8

1 file

2.0.7

1 file

2.0.6

1 file

2.0.5

1 file

2.0.4

1 file

2.0.3

1 file

2.0.2

1 file

2.0.1

1 file

1.1.0

1 file

1.0.3

1 file

1.0.2

1 file

1.0.1

1 file

1.0.0

1 file

0.0.1

1 file

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page