Skip to main content

CompileML

ci PyPI Python License DOI

Keep the predictive power of a tree ensemble. Deploy it as transparent decision logic with no model runtime.

Why this exists

CompileML started in credit risk, where two schools of thought meet and rarely agree.

Risk practitioners build scorecards. A scorecard is transparent, reproducible, and deploys anywhere; it is trusted because a validator can check every point by hand. It also leaves predictive power on the table. Data scientists build tree ensembles. An ensemble predicts better, then arrives with a Python environment, a serving stack, post-hoc explanations, and a model no validator can reproduce independently.

Each side is right about what the other gives up. That leaves teams a bad choice: the model they can defend, or the model that performs.

The choice is not unique to credit. It appears wherever a model makes a decision about an individual case and has to answer for it — to a regulator, an auditor, a customer, a clinician — or has to run somewhere the data-science stack does not. Fraud and anti-money-laundering alerts, insurance underwriting and claims, eligibility screening and clinical decision support all meet it. The details differ; the choice is the same.

CompileML removes that choice.

Train the strongest model you can: XGBoost, LightGBM, a neural network. It sets the ceiling. CompileML then searches for the best whitebox it can compile — trained on the outcomes, on the ceiling's predictions, or on a blend — and chooses on data it will not report on. The whole decision is packaged into one hashed artifact: score, calibrated probability, risk bands, reason codes, attribution. What you get is the artifact, what its guarantees cost against that ceiling, and whether it beats the scorecard you already have.

The ceiling is a measuring instrument. It never ships.

What ships instead is decision logic built from integer addition, comparison, and table lookup. It runs under CompileML's standard-library-only Python implementation, an ordinary application, a small Lambda, a batch job or you export it as standalone SQL or COBOL and run it directly where the decision already happens. No XGBoost in production, no scikit-learn, no scoring service to operate. In exported form there is no model runtime at all.

This is not a lighter way to serve the black box. The black box set the bar and was then compiled out of the system.

That buys three things, and I would not trade any one of them for the other two: the artifact reproduces to the integer on any machine, it explains itself by arithmetic rather than approximation, and it lands where the decision already runs. The sections below are how each one is enforced.

The benchmark puts the cost at about 2% of the ceiling's Gini — 98.04% retained on rows the selection never saw, with a 95% interval of 97.38–98.72% — and the artifact out-scores a WoE logistic regression on the same rows by 3.4%. In exchange, at whitebox depth two or less, every decision reconstructs exactly from a printed scorecard table, and every explanation adds back to the production score with nothing left over.

The examples in this repository come from credit, where the project began and where lenders must explain every adverse decision, reproduce it for a validator, and often run it on systems that predate Python. Where it fits maps the vocabulary to other domains and says where the fit is weaker.

Quick example

Train the strongest model you can. CompileML chooses the whitebox, prices it against that model, and reads the reported figure from rows it never chose on:

from compileml import compile_selected
from compileml.artifact import save_artifact

result = compile_selected(
    X, y,
    ceiling=lambda X_fit, y_fit: XGBClassifier(**tuned).fit(X_fit, y_fit),
    reference="woe",              # the floor: a WoE logistic regression, fitted inside Fit
    reasons=REASON_DICTIONARY,
)

result.selected                   # the target and configuration chosen on Select
report = result.report()          # Report, read once: retention and floor ratio, with intervals
save_artifact(result.artifact, "decision.json")

Every step is also a plain function — train_whitebox, monotone_quantile_bands, build_artifact — for when you want each one in your hands; the quickstart walks that path.

Production does not need the training stack:

from compileml.runtime import load_artifact, decide

artifact = load_artifact("decision.json")
decision = decide(artifact, applicant_row)

# {
#   "band": "G07",
#   "pd": 0.1284,
#   "latent_int": 146,
#   "reasons_negative": [
#       {
#           "code": "HIGH_UTILIZATION",
#           "message": "…",
#           "impact_int": 56,
#       }
#   ],
#   "reasons_positive": [...],
#   "artifact_hash": "84372c36…",
# }

The runtime imports nothing outside the Python standard library.

Or skip the Python runtime entirely:

compileml export decision.json --target sql   --out scorer.sql
compileml export decision.json --target cobol --out scorer.cob

The three objections, answered

I have watched good models become much less impressive on the way to production.

The model starts in Python. Someone rewrites it in SQL. Someone else builds the bands in a spreadsheet. Calibration lives in another script. Reason codes are produced through a separate explanation process. Six months later, everybody is discussing "the model," but they are no longer talking about exactly the same thing.

Each of the three standing objections has a structural answer.

Stability: scores drift

Floating-point arithmetic is not a reassuring foundation for a decision that must be reproduced across languages and systems. Small differences in accumulation, precision, or implementation can move a score near a boundary.

A credit score should be a fact, not a distribution over environments.

CompileML quantizes model leaves once, at compile time. After that, scoring is integer addition, banding is integer comparison, and calibration is integer table lookup. Recalibration refits probabilities while the model and band edges stay byte-identical, so updating a PD table cannot move a single account between bands.

Explainability: explanations do not reconcile

Post-hoc explainers are useful, but an explanation of a regulated decision should not merely resemble the decision.

CompileML computes attribution from the compiled model in integer units. The feature impacts, baseline, and residual satisfy a reconciliation identity that validation can add back independently.

For whiteboxes of depth two or less, the pairwise decomposition is complete and the residual is exactly zero. The runtime refuses to emit an explanation that fails to reconcile.

Deployability: the deployment stack is not the modeling stack

Banks and other large institutions run important decisions on SQL systems, core platforms, and mainframes. Requiring the entire training environment in production is often unrealistic and sometimes unnecessary.

CompileML moves complexity to compile time and leaves production with a small, explicit artifact.

It compiles to a scorecard

At whitebox depth ≤ 2, the compiled artifact collapses into a classic points scorecard — exactly, not as an approximation.

compileml scorecard decision.json --format csv --out scorecard.csv

Depth 1 produces the familiar form: per feature, a bin and its points. Depth 2 adds explicit pairwise interaction grids over the union of the relevant thresholds. Points are the artifact's own integers, and the identity

base_points + Σ main_effect(x) + Σ interaction(x) == score

holds bit-for-bit on every row. score_from_scorecard() re-derives any production decision from the printed tables alone, and the test suite asserts it.

Hand the CSV to a validator and they can reproduce production scores in a spreadsheet. Above depth 2 no clean scorecard exists, and the tool raises instead of approximating — the same boundary as exact attribution, for the same reason.

Expect a depth-2 scorecard to be mostly interaction grids. A tree whose two splits use different features contributes a grid rather than a main effect, and boosting seldom spends both splits on one feature: the fairness notebook's 40-tree model on the 23-feature UCI panel compiles to 55 grids and no main effects at all. It is still exact, but it is not the one-table-per-feature card many validation teams expect. If yours does, compile at max_depth=1 and measure the fidelity that costs with sweep_whitebox.

What the artifact guarantees

Given the same artifact and the same input values, CompileML produces the same governed integer outputs across supported runtimes.

The repository tests this rather than asking you to take it on faith:

  • SQL output is executed in SQLite and compared row by row with the Python runtime.
  • Generated COBOL is compiled and run in CI, then checked against the reference implementation.
  • The same seeded artifact is built on Linux, macOS, and Windows and the hashes are compared.
  • Committed reference decisions are replayed on every OS and Python version in the matrix.
  • Attribution is added back to the decision during validation.
  • Recalibration tests verify that the model and band edges remain unchanged.
  • Scorecard tables are re-summed against the runtime's own integers.
  • The standard-library-only runtime is enforced by inspecting its imports.

The artifact includes a SHA-256 hash. Loaders verify it by default and reject a document whose contents no longer match the stored hash. This detects modification or corruption; it is an integrity check, not a cryptographic signature of who produced the artifact.

Measured performance

The committed benchmark uses deterministic synthetic credit data with 40,000 rows and 23 features, split 60/20/20 into Fit, Select and Report. compile_selected chooses the target, the tree count and the depth on Select from 40 configurations; the winner is refit on Fit ∪ Select; and every retention figure below is read once from Report. It selected α = 0.75 — a blend of labels and the ceiling's cross-fitted predictions — with 80 trees at depth 2, inside a tie band of 4; both pure targets, labels alone and the ceiling alone, scored lower on Select. The ceiling is a 300-tree, depth-4 GBM with a fixed, undeclared-search configuration, recorded as such. Latency runs on a consumer laptop through the pure-Python runtime.

You can reproduce every number from an editable install of the checked-out source:

pip install -e .
python benchmarks/run_benchmarks.py

results.json records the CompileML version it measured, so a run against an older installed copy cannot pass for a measurement of the current code. The table below and the cost figures elsewhere in the docs are written from that file by python benchmarks/sync_docs.py, never by hand.

Metric Value
Ceiling Gini, 300-tree GBM 0.664
Compiled integer artifact Gini 0.651 — 98.0% retained (97.4–98.7)
Floor Gini, WoE logistic regression 0.630 — artifact at 103.4%
Band-ordinal Gini, 10 bands 0.643 — 96.7% retained
Spearman correlation, ceiling vs. artifact 0.980
Selected on Select, of 40 configurations α = 0.75, 80 trees, depth 2
Score + band + calibrated PD 0.02 ms median
Score + band + calibrated PD, p95 0.03 ms
Full explained decision, 80 trees 0.42 ms median
Full explained decision, p95 0.68 ms
Band assignment alone 0.2 µs
Artifact size 70 KB
Same configuration and identical hash on rerun Yes

That 2% of Gini is the price of everything above it. It is stated rather than hidden, it comes with an interval, and compile_selected measures it the same way on your own data. The full selection curve — every configuration scored on Select — is committed beside the results as benchmarks/selection_curve.json; it ranks configurations and does not describe the shipped artifact.

One honest qualification: scoring is very fast; full explanation costs more.

The exact pairwise decomposition is aggregated per tree, which makes it O(trees) and independent of how many features the model has. A depth-2 tree is walked at most eight times however wide the model is. A perturbation-based derivation instead scores the whole ensemble 2 + p + p(p−1)/2 times.

On the benchmark's 120-tree ensemble, attribution alone:

features perturbation per tree tree walks, perturbation tree walks, per tree
8 0.94 ms 0.60 ms 4,560 792
23 6.87 ms 0.69 ms 33,360 912
50 34.01 ms 0.74 ms 153,240 944
100 133.88 ms 0.74 ms 606,240 936

The walk counts are exact and hold on any machine; the milliseconds belong to one laptop. Cost follows walks, and walks follow tree structure rather than width: the sweep's models are fitted to a target that uses every feature, so more of their trees split on three distinct features than the headline model's do, which is why attribution alone at 23 features here costs slightly more than the full decision in the table above. At eight features the difference between the two derivations is modest. At a hundred it is the difference between a quadratic cost and a flat one. It is exact rather than sampled, and both derivations reach identical integers on every timed row.

In practice: explain everything. Under a millisecond per decision is real-time for credit decisioning — the bureau pull costs more — and complete attribution on every decision is what turns portfolio questions (marginal analysis, driver drift, fairness cuts) into census facts instead of sample estimates. It also means every production decision carries its own explanation in the record, computed at decision time under the same artifact hash.

Batch re-explanation over an entire book used to be the one place the cost bit, and wide feature sets made it worse. Per-tree aggregation removed both.

Choosing the configuration

The two capacity knobs are not symmetric, and this is the single most useful thing to know before tuning.

n_estimators buys fidelity at a linear cost in artifact size and explanation time, and costs nothing else. Determinism, portability, and exact attribution are unaffected at any tree count.

max_depth buys fidelity per tree, but above 2 it takes the exactness guarantee with it: attribution residuals appear and no clean scorecard exists.

Spend on trees. Be stingy with depth.

Both are measurable rather than guessable:

from compileml.tune import sweep_whitebox, sweep_bands

sweep_whitebox(X, teacher_latent, y, alpha_grid=(0.0, 0.5, 1.0),
               X_val=X_val, y_val=y_val, teacher_latent_val=t_val)
sweep_bands(latent, y, k_grid=(4, 6, 8, 10, 12, 16))

For banding, band_efficiency() reports what the ladder discards — the Gini gap against the continuous score — and, per band, whether the score still ranks risk internally. A band that can still separate outcomes is a refinement opportunity; a band that cannot is a band you have used up. The tuning guide walks through both.

What CompileML is not

CompileML is not a new training framework. Use XGBoost, LightGBM, scikit-learn, or any model that can set the ceiling for a whitebox, or be compiled into the supported representation directly.

It is not a promise that your data pipelines are identical. Determinism means:

same input values + same artifact = same governed outputs

Producing the same input values across systems remains the caller's responsibility.

It is also not a compliance certification. No library can certify an institution's model, data, policy language, or governance process.

CompileML is infrastructure intended to make those things inspectable instead of asking validators to trust a chain of separate implementations.

Reason codes belong to the institution

CompileML can determine which features moved a decision and by how much. It cannot decide how your institution should explain that result to a customer.

That language is policy, not mathematics.

You provide the reason dictionary:

REASON_DICTIONARY = {
    "BILLS_PAID_LATE": {
        "code": "LATE_PAYMENTS",
        "negative": "Recent payments were made after their due date.",
        "positive": "Consistent on-time payment history.",
    },

    # Add one entry per feature.
    # `suppress: True` hides policy-masked features.
}

Features without an entry still work, but they receive generic fallback messages. CompileML measures reason coverage, records it in the artifact, warns when coverage is incomplete, and can make full coverage a validation requirement.

The tooling should not quietly pretend that generic feature names are suitable adverse-action notices.

Seeing what the model did

The visualization package draws the outputs emitted by decide(). It does not independently recompute the model or explanation.

That matters: the chart cannot disagree with the deployed decision because both come from the same payload.

from compileml.viz import (
    waterfall,
    decision_drivers,
    band_conditioned_decision_drivers,
    band_ladder,
)

waterfall(
    decide(artifact, row, include_contributions=True)
)

decision_drivers(sample_decisions, y=y_sample)
band_conditioned_decision_drivers(sample_decisions, y=y_sample)
band_ladder(score_decisions, y_sample)

The image at the top of this README was rendered by waterfall_svg() from the repository's committed reference artifact. That renderer also uses only the standard library.

Validation

Run the validation framework against a holdout set:

compileml validate artifact.json \
  --csv holdout.csv \
  --y-col DEFAULT

It checks:

  1. artifact integrity;
  2. explanation reconciliation;
  3. fidelity to the source model;
  4. band coverage, score resolution, and band efficiency;
  5. bad-rate monotonicity;
  6. band-ladder churn;
  7. explanation stability;
  8. reason-code coverage;
  9. declared monotone directions, re-verified against the shipped trees;
  10. that the artifact out-scores a reference model on the same data;
  11. that, when the artifact records how it was selected, the protocol was followed: cross-fitted soft targets, disjoint partitions, one Report read.

These checks run against the compiled artifact through the same runtime used for production decisions. There is no separate notebook implementation allowed to become "almost the same" over time.

The command exits with 0 or 1, so it can gate deployment in CI.

Independent evaluation

@deburky evaluated CompileML 0.4.3 on the AWS Fraud Detector sample data with a CatBoost teacher — a different domain, teacher and dataset from anything in this repository — and published the code and notes in deburky/compileml-fraud-scoring.

On that data:

  • Gini retention of 98.5% and 96.9% at depth 2, on two datasets;
  • all ten validation checks passing on both artifacts;
  • the SQL export, executed in SQLite, matching the Python runtime on every holdout row;
  • scorecard re-sums matching the production score on every row;
  • the runtime scoring with no NumPy, pandas, scikit-learn or CatBoost loaded, including from a SageMaker serving image built without them;
  • recalibration after a fourfold drop in base rate leaving the model and band edges byte-identical, with zero band churn.

It found two bugs, both fixed in 0.5.0: invalid SQL for a single-band artifact (#40, fixed by @tote10) and band builders emitting edges the build then rejected on low base rates (#41).

Its notes also recorded five usability observations, all addressed in 0.5.2: a latent-range warning that blamed margin-space models even for train_whitebox output; no pointer onward when certified banding finds a single band; a documented recalibration metadata key that did not match the one written; an undocumented macOS dependency of the XGBoost and LightGBM extras; and nothing warning that a depth-2 scorecard is mostly interaction grids.

It did not exercise the COBOL export (no compiler was installed), determinism across operating systems (CI checks that separately), or anything released after 0.4.3.

Citing

Every release is archived on Zenodo. Cite the concept DOI — it always resolves to the newest version:

Ortiz, C. CompileML. https://doi.org/10.5281/zenodo.22242231

To pin a specific version, use that release's own DOI from its Zenodo record. GitHub's Cite this repository button reads CITATION.cff and will format it for you.

Install

pip install compileml

Optional teacher integrations:

pip install compileml[xgboost]
pip install compileml[lightgbm]

On macOS both libraries load the OpenMP runtime. If importing either fails while loading its shared library, install it with brew install libomp. That is a requirement of XGBoost and LightGBM, not of CompileML, whose runtime needs neither.

Visualization dependencies:

pip install compileml[viz]

The compile side depends on NumPy and scikit-learn. The compileml.runtime package uses only the Python standard library.

If necessary, the runtime directory can be vendored into a constrained environment:

src/compileml/runtime/

Documentation

Roadmap

Recently shipped:

  • exact attribution aggregated per tree, so explaining a decision no longer grows with feature count, and a fairness audit in compileml.fairness (0.5);
  • exact drift decomposition, band calibration and baseline staleness in compileml.monitor (0.6);
  • the calibrated PD and reason codes from the COBOL and SQL exports (0.7);
  • retention by segment and across PD cutoff ranges, and weighting a whitebox toward the segment that pays (0.8);
  • compile_selected: the target and the configuration chosen on data the report never sees, the ceiling as yardstick, a provenance block that carries the reported cost, and a NumPy batch scorer (0.9).

The current priorities are:

  • the gate experiment (#76): where soft targets stop helping, and whether log-loss earns a logit-space latent;
  • a logit-space latent, if that experiment supports it;
  • add Java and C exporters.

License

Apache-2.0.

Copyright 2026 Carlos Ortiz.

Metadata

Release files for compileml 0.9.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for compileml 0.9.0
File Size Uploaded
compileml-0.9.0.tar.gz 180.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for compileml 0.9.0
File Interpreter ABI Platform
compileml-0.9.0-py3-none-any.whl Python 3 none any Details

Total release size: 332.8 kB

Release files / compileml-0.9.0.tar.gz

Download URL compileml-0.9.0.tar.gz
Size 180.5 kB
Tags Source
SHA-256 checksum
How to use checksums
7f3fdf4add3cf7e7ca15654b5616fa04010a8135428496859dd52482c093c73a
BLAKE2b-256 checksum
How to use checksums
2beb98955bd23ea223ed87b8121ea9ae765a118a0038013eb1dfde033da26858
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.

Transparency log

Release files / compileml-0.9.0-py3-none-any.whl

Download URL compileml-0.9.0-py3-none-any.whl
Size 152.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b0d743ac15557a2dc0e949b49d69963450ab0c326c2d16c14a5fd6ca42b7bda4
BLAKE2b-256 checksum
How to use checksums
75f0dd405d5d91cc1e52de640f0af17e2368e8d7fb2f1cd315891f284bed6a0a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.9.0 This release

2 release files

0.8.0

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.3

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page