Skip to main content

ab-analysis-kit

A/B experiment analysis as declarative YAML + SQL — with a chart-first cockpit.

ab-analysis-kit (CLI abk) is an open-source, declarative (dbt / detectkit-style), database-agnostic, numpy-first Python library for analyzing A/B experiments. You define an experiment and its metrics in YAML + SQL; abkit computes per-method effect + confidence interval + p-value + MDE/power cumulatively over the experiment's lifetime (the stabilization chart), writes them to a clean warehouse table any BI can read, and gives you a local cockpit to tune the analysis and a harness to prove your method is actually calibrated.

Status: 0.7.0 (Alpha) — the latest on PyPI (milestones M1–M12 shipped — M12 wired notifications: abk run --notify and abk validate --notify push what a run just decided to nine channel types (Slack, Telegram, email, webhook, Mattermost, Discord, Teams, Google Chat, ntfy), as six routable signals — the readout verdict, a verdict flip, a failed sample-ratio gate, a pipeline error, a slipped schedule, and an A/A cell that broke its false-positive budget. Nothing is recomputed for a message, so it cannot disagree with the report; a repeat run over unchanged data is silent; and no channel failure can change an exit code. M11 added abk dashboard, the project-level cockpit: one row per experiment with its headline verdict, effect + CI, p/α and a sparkline of the cumulative series, plus buttons that spawn real abk subprocesses (Run — for the whole experiment or one metric — Unlock, Clean, Explore, Open report) and stream their logs. It is a launcher: it computes no statistic and never takes the pipeline lock, so every verdict on the page is the readout's own. The 0.6.x interstitial then closed both abk plan sizing gaps: a CUPED comparison is sized on the covariate correlation its own results row already persists — required-N is (1 − ρ²)× the old raw-variance bound (0.6.1) — and --from-history <N d> gives an experiment that has never run a baseline from the days before its start, instead of SKIPPED: no baseline (0.6.2). A second 0.6.x interstitial then gave the dashboard CRUD YAML editing — edit, create, delete an experiment from the cockpit, validated at both levels and archived byte-verbatim before every write — added abk ui as its alias, and made M9's additive read path discoverable: abk run no longer stays quiet about an undecided compute.incremental_reads, --cost-report prints the counterfactual, and abk init scaffolds it on (0.6.4)). The statistical core, the declarative config / DB / pipeline layer, the explore cockpit + self-contained reports, abk validate (numpy-vectorized — minutes → sub-seconds), opt-in sequential analysis + abk plan, and the DX layer (abk init-claude, docs site, Prefect scaffolding) are all shipped. Docs: abkit.pipelab.dev.

Install

pip install ab-analysis-kit          # Python 3.10+; add a DB extra for real data:
pip install "ab-analysis-kit[clickhouse]"   # or [postgres] / [mysql] / [all-db]

(pip install ab-analysis-kit gets 0.7.0 — opt-in notifications across nine channels, abk dashboard with its YAML editor, abk ui, CUPED-aware abk plan sizing, abk plan --from-history and the discoverable additive read path all included.)

abk --version and abk --help work with no database driver; you can even lint a config (abk run --steps validate) with no database at all. See the getting-started guide for the full first run.

What it does

  • Declarative experimentsexperiments/*.yml (assignment + variants + comparisons) referencing a reusable metrics/*.yml library (YAML + SQL).
  • A rigorous statistical engine — t-test, two-proportion z-test, CUPED, ratio (delta-method), and a vectorised bootstrap family (plain/paired/Poisson/ post-normed), with relative & absolute effects, MDE/power, and multiple-testing correction. Ported from a battle-tested legacy engine and improved deliberately.
  • The cumulative stabilization chart — effect + CI per day from experiment start, so you see the estimate converge and call a winner only once it stabilizes.
  • abk dashboard — the project-level cockpit: every experiment as one row (verdict, effect + CI, p/α, sparkline), with Run / Unlock / Clean / Explore / Open-report buttons that spawn real abk subprocesses and stream their logs. It never computes a statistic itself, so what you read is the readout's own verdict.
  • abk explore — a local, chart-first cockpit to turn method knobs (CUPED, stratification, alpha…) and watch the result recompute live, with A/A calibration always in view. The priority interface.
  • abk validate — an A/A false-positive + power matrix that measures your method's real α (including the honest cumulative-peeking FPR), not the nominal.
  • BI-agnostic — results land in one clean table; connect Grafana, Lightdash, Metabase, or Superset. Orchestrate with Prefect.
  • AI-nativeabk init-claude sets up assistant context + skills so an assistant can scaffold and tune experiments with (or for) you.

Design at a glance

experiment (YAML)  ──▶ load exposures ──▶ SRM gate ──▶ compute (t/z/CUPED/bootstrap) ──▶ readout
  └ references reusable metrics (YAML + SQL)                                          └ _ab_results → your BI

abkit is the sibling of detectkit: same DNA (CLI-first, db-agnostic, numpy-first, self-contained reports, a chart-first cockpit, init-claude), with the anomaly detect stage replaced by a statistical compute stage and the primary entity flipped from metric to experiment.

Documentation

License

MIT.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

ab_analysis_kit-0.7.0.tar.gz (608.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

ab_analysis_kit-0.7.0-py3-none-any.whl (716.8 kB view details)

Uploaded Python 3

File details

Details for the file ab_analysis_kit-0.7.0.tar.gz.

File metadata

  • Download URL: ab_analysis_kit-0.7.0.tar.gz
  • Upload date:
  • Size: 608.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for ab_analysis_kit-0.7.0.tar.gz
Algorithm Hash digest
SHA256 11a767a384e2ae4b19c79be84289f80c0936d754dbcb0f32267be0d9e53cd41c
MD5 4772acde6d3373b0da5dc1431d5843e9
BLAKE2b-256 cffa12361b6f9050da5e6f17dc63a4c5aa2de216b637eded7dd9fabe6576f0bc

See more details on using hashes here.

Provenance

The following attestation bundles were made for ab_analysis_kit-0.7.0.tar.gz:

Publisher: publish.yml on alexeiveselov92/ab-analysis-kit

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file ab_analysis_kit-0.7.0-py3-none-any.whl.

File metadata

  • Download URL: ab_analysis_kit-0.7.0-py3-none-any.whl
  • Upload date:
  • Size: 716.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for ab_analysis_kit-0.7.0-py3-none-any.whl
Algorithm Hash digest
SHA256 75d130a38054292edfc7c4d51cb626a126867a2431a781e675cbe04af92146e4
MD5 e9d47acf88272debc8140f673067482b
BLAKE2b-256 bfc571f6a81fcd1d32a66030a66bae74a356e2e6ac483ddcbbbc3d99e6c14b53

See more details on using hashes here.

Provenance

The following attestation bundles were made for ab_analysis_kit-0.7.0-py3-none-any.whl:

Publisher: publish.yml on alexeiveselov92/ab-analysis-kit

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page