Python implementation of CATDAP (CATegorical Data Analysis Program)
Project description
pycatdap
AIC-based EDA and ML error analysis library for categorical data.
pycatdap is a Python implementation of CATDAP (CATegorical Data Analysis Program), developed by Sakamoto & Katsura (1980) at the Institute of Statistical Mathematics. It extends the classic CATDAP toolkit with modern exploratory data analysis (EDA) and machine learning error analysis workflows.
📖 Documentation: https://nbx-liz.github.io/pycatdap/
Why pycatdap?
Unlike general profilers (ydata-profiling, Skrub) or slice discovery tools (DivExplorer, pysubgroup), pycatdap uses AIC as its core relevance measure. This gives it four unique advantages:
| Feature | Most tools | pycatdap |
|---|---|---|
| Variable relevance | Cramér's V, mutual info | AIC — explicit info-vs-complexity trade-off |
| Continuous binning | Equal-width or quantile | AIC-optimal binning |
| Subset discovery | Feature importance ranking | CATDAP-02 combinatorial search |
| Model coupling | Tied to specific frameworks | Model-agnostic (works with y_true, y_pred from anywhere) |
Installation
# Core
pip install pycatdap
# With visualization (matplotlib)
pip install "pycatdap[plot]"
# With interactive Plotly figures + HTML reports
pip install "pycatdap[plotly]"
Supported: Python 3.10 / 3.11 / 3.12 / 3.13
Quickstart
Classic CATDAP
import pycatdap
df = pycatdap.datasets.load_health_data()
# CATDAP-01: pairwise AIC analysis
result = pycatdap.catdap1(df, response_names=["symptoms"])
print(result.aic_order["symptoms"]) # variables ranked by ΔAIC
# CATDAP-02: best explanatory subset
result2 = pycatdap.catdap2(
df,
pool=[2, 2, 2, 0, 0, 0, 0, 2],
response_name="symptoms",
accuracy=[0., 0., 0., 1., 1., 1., 0.1, 0.],
)
for s in result2.subsets[:3]:
print(f"AIC={s.aic:.2f}, vars={s.variables}")
One-call EDA report (v0.5+)
report = pycatdap.profile(df, response="symptoms")
report.show() # Jupyter inline
report.to_html("report.html") # self-contained HTML, inline Plotly
report.to_dict() # JSON-friendly
report.to_plotly_json() # react-plotly.js / LizyStudio
ProfileResult exposes overview, variables (one VariableCard per
column, including ΔAIC vs the response), association (m × m ΔAIC
matrix), top_subsets (CATDAP-02 result when response is given), and
quality_warnings (high_cardinality / constant / id_candidate /
high_missing with overridable thresholds). See
docs/tutorials/08-profile-titanic.ipynb
for an end-to-end walkthrough.
Target analysis and CI-integrable suite (v0.6+)
# Target-driven: rank every column by ΔAIC vs `response`, keep top-K cross-tabs
ta = pycatdap.target_analysis(df, response="symptoms", top_k=5)
ta.ranking # variable / delta_aic / kind / n_obs
ta.top_summaries["cholesterol"] # full TargetSummary for drill-down
# Quality scan only (fast — no catdap2 / association_matrix)
qr = pycatdap.quality_report(df)
assert qr.passed, qr.show()
# CI gate: one-line data contract
suite = pycatdap.suite.AICIndependenceSuite(df, response="symptoms")
result = suite.run()
assert result.passed, result.summary()
# Non-AIC association measures (pure-numpy, no scipy)
m = pycatdap.association_matrix(df, measure="cramers_v")
pycatdap.measures.register("my_measure", my_fn) # pluggable per pysubgroup convention
See docs/tutorials/09-phase-d-target-analysis-and-suite.ipynb.
ML error analysis — Phase G (v0.7.0)
from pycatdap.datasets import load_german_credit
df = load_german_credit()
y_true = df["class"]
y_pred = my_model.predict(df.drop(columns=["class"]))
# Label every prediction "correct" / "incorrect"
labels = pycatdap.error.error_label(y_true, y_pred)
# Or get fine-grained TP / FP / FN / TN (binary classification only)
conf = pycatdap.error.confusion_label(y_true, y_pred, positive="bad")
# Feed the error label back into Phase D to discover which columns
# explain the errors — foundation for v0.8.0 `error_analysis()`
ta = pycatdap.target_analysis(df.assign(was_correct=labels), response="was_correct")
See docs/tutorials/10-phase-g-error-labeling.ipynb.
ML error analysis — Phase H one-call (v0.8.0)
# task auto-detected from y_true / y_pred dtypes
result = pycatdap.error_analysis(
df=test_df,
y_true=y_test, # str column name or array
y_pred=model.predict(X_test), # str column name or array
top_k=5,
)
result.show() # Jupyter / stdout summary
result.to_html("errors.html") # self-contained Plotly report
result.feature_ranking # ΔAIC ranking of explanatories
result.top_slices # single-variable cohort slices
result.confusion # canonical TP/FP/FN/TN (binary)
result.to_divexplorer_format() # DivExplorer-compatible DataFrame
# Larger fairness-relevant benchmarks (needs `pip install 'pycatdap[data]'`)
df = pycatdap.datasets.fetch_california_housing() # regression
df = pycatdap.datasets.fetch_adult_income() # fairness demo
df = pycatdap.datasets.fetch_compas() # fairness demo
See docs/tutorials/11-phase-h-error-analysis.ipynb.
ML error visualisation — Phase I+J (v0.9.0)
r = pycatdap.error_analysis(df, y_true, y_pred)
# Result delegation — visualise straight off the result
r.plot_confusion() # binary OR multi-class
r.plot_confusion(backend="plotly", normalize="true")
r.residual_plot() # regression
# Or call the standalone functions
pycatdap.error.plot_confusion(y_true, y_pred, labels=[0, 1], normalize="true")
pycatdap.error.plot_confusion_by_slice(df, y_true, y_pred, var="age_group")
pycatdap.error.confusion_aic(y_true, y_pred) # ΔAIC, negative = informative
pycatdap.error.residual_plot(y_true, y_pred, kind="histogram")
pycatdap.error.residual_by_category(df, y_true, y_pred, "feature")
pycatdap.error.residual_pool_plot(y_true, y_pred, n_bins=4)
See docs/tutorials/12-phase-i-j-error-visualization.ipynb.
ML calibration — Phase K (v0.10.0)
r = pycatdap.error_analysis(df, y_true, y_pred, y_proba=proba) # binary classification
r.calibration_curve(strategy="aic") # reliability diagram (AIC-binned)
# Or call the standalone functions
pycatdap.error.calibration_curve(y_true, proba, strategy="aic", backend="plotly")
pycatdap.error.brier_score(y_true, proba)
pycatdap.error.expected_calibration_error(y_true, proba) # ECE
pycatdap.error.maximum_calibration_error(y_true, proba) # MCE
strategy="aic" bins the probability axis where the observed positive-rate shifts — sharper than equal-width / quantile bins on skewed predictions. See docs/tutorials/13-phase-k-calibration.ipynb.
ML slice discovery, cohort comparison & drift — Phase L (v0.11.0)
# Auto-discover the multivariable cohorts where the model fails most
result = pycatdap.error.discover_error_slices(
df, y_true, y_pred, max_vars=3, measure="aic", top_k=10, min_support=30
)
for s in result.slices:
print(s.error_metric, s.description) # e.g. "age ∈ [60, 78] × plan = basic"
result.to_divexplorer_format() # DivExplorer-compatible flat table
# Compare two cohorts (distribution + ΔAIC), Sweetviz-style HTML report
pycatdap.error.compare_cohorts(df_a, df_b).to_html("comparison.html")
# Detect train→prod drift, ranked by ΔAIC magnitude
pycatdap.error.detect_drift(df_train, df_prod, y_true=y, y_pred=yhat)
# Calibration beyond binary (deferred from Phase K)
pycatdap.error.regression_calibration_table(y_true, y_pred) # regression
pycatdap.error.multiclass_calibration_table(y_true, y_proba) # one-vs-rest
Slice discovery prunes the search space on support (Apriori, anti-monotone) rather than ΔAIC — sound, and >50% reduction on wide datasets. The interestingness measure is pluggable ("aic", "cramers_v", "mutual_info", or any registered callable). See docs/tutorials/14-phase-l-slice-discovery.ipynb.
Status & Roadmap
| Version | Theme |
|---|---|
| v0.2.0 ✅ | Core CATDAP-01/02 (released) |
| v0.3.0 — v0.6.0 ✅ | EDA workflow (Plotly backend, profile, target analysis) |
| v0.7.0 ✅ | Phase G error labelling building blocks |
| v0.8.0 ✅ | Phase H error_analysis() one-call + D4 benchmarks |
| v0.9.0 ✅ | Phase I+J error visualisation (confusion + residual) |
| v0.10.0 ✅ | Phase K calibration (AIC-binned reliability diagram + Brier/ECE/MCE) |
| v0.11.0 ✅ | Phase L slice discovery + cohort comparison + drift + regression/multi-class calibration |
| v0.12.0 | LizyStudio integration |
| v1.0.0 | API stabilization |
Full roadmap: PLAN.md · Meta Issue #11
Development
git clone https://github.com/nbx-liz/pycatdap.git
cd pycatdap
uv sync --all-groups
uv run pytest # tests (excluding slow R cross-validation)
uv run pytest -m slow # slow tests (requires R + catdap package)
uv run python -m mkdocs serve # local docs preview
make ci # ruff + mypy + pytest + build
Contributing guidelines: CONTRIBUTING.md
Project structure
| Document | Purpose |
|---|---|
| BLUEPRINT.md | Canonical specification (Japanese) |
| HISTORY.md | Proposal-to-decision log (Japanese) |
| PLAN.md | Development roadmap (Japanese) |
| CHANGELOG.md | Release history |
| docs/ | Published documentation site |
Citation
If you use pycatdap in research, please cite the original CATDAP work:
@article{sakamoto1980categorical,
title={Categorical Data Analysis by AIC},
author={Sakamoto, Yosiyuki and Katsura, Koichi},
journal={Mathematical Sciences},
year={1980}
}
License
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file pycatdap-0.12.1.tar.gz.
File metadata
- Download URL: pycatdap-0.12.1.tar.gz
- Upload date:
- Size: 661.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ed6949009adc44048c069b23ef428f20dee40df2cdce5f9f424024c503465e8d
|
|
| MD5 |
eb7e99055154081c93c90d6628ef5e83
|
|
| BLAKE2b-256 |
b2833ca5aeb93643b7a91f703084a407248091984c5acab6b9b6746ca723bf24
|
Provenance
The following attestation bundles were made for pycatdap-0.12.1.tar.gz:
Publisher:
release.yml on nbx-liz/pycatdap
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pycatdap-0.12.1.tar.gz -
Subject digest:
ed6949009adc44048c069b23ef428f20dee40df2cdce5f9f424024c503465e8d - Sigstore transparency entry: 1676068887
- Sigstore integration time:
-
Permalink:
nbx-liz/pycatdap@c5ab2e23a5d9f9638a1d89d47548fac82c16e118 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/nbx-liz
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@c5ab2e23a5d9f9638a1d89d47548fac82c16e118 -
Trigger Event:
workflow_dispatch
-
Statement type:
File details
Details for the file pycatdap-0.12.1-py3-none-any.whl.
File metadata
- Download URL: pycatdap-0.12.1-py3-none-any.whl
- Upload date:
- Size: 214.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5e06fa69fcf3f00d1704e775f663edcf9934168f1816632083836d9a843eb317
|
|
| MD5 |
e44298a9d3e1609d5e8f4afbc6c06ba6
|
|
| BLAKE2b-256 |
8583c8466a6391215b3faebe71b1f41645606de81c594a6a9d5587f8e092ff2b
|
Provenance
The following attestation bundles were made for pycatdap-0.12.1-py3-none-any.whl:
Publisher:
release.yml on nbx-liz/pycatdap
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pycatdap-0.12.1-py3-none-any.whl -
Subject digest:
5e06fa69fcf3f00d1704e775f663edcf9934168f1816632083836d9a843eb317 - Sigstore transparency entry: 1676068920
- Sigstore integration time:
-
Permalink:
nbx-liz/pycatdap@c5ab2e23a5d9f9638a1d89d47548fac82c16e118 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/nbx-liz
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@c5ab2e23a5d9f9638a1d89d47548fac82c16e118 -
Trigger Event:
workflow_dispatch
-
Statement type: