A lightweight A/B testing framework: fixed-horizon tests, CUPED variance reduction, ratio metrics, power analysis, and multiple-comparison corrections.
Project description
abpy — Frequentist A/B Testing Library
A Python library for rigorous frequentist A/B and multivariate testing, with built-in support for CUPED variance reduction, Delta method for ratio metrics, and multiple comparison corrections.
Installation
pip install abpy-toolkit
Quick Start
from abpy.router import Experiment
# binary metric — direct input
exp = Experiment(type='ab', metric='binary', input='direct', alpha=0.05)
exp.fit({
'control': {'n': 1000, 'successes': 100},
'treatment': {'n': 1000, 'successes': 130}
})
exp.results()
exp.plot()
Features
- Three metric types — binary, continuous, count
- Two input modes — CSV file or direct summary statistics
- Automatic test selection — z-test, t-test, Mann-Whitney U, Poisson rate test
- CUPED — variance reduction using pre-experiment data
- Delta method — correct variance estimation for ratio metrics
- Multiple comparison corrections — Bonferroni, Holm, Benjamini-Hochberg
- Multivariate testing — all pairwise comparisons with corrections
- Power analysis — achieved power for every test
- Dashboard visualisation — 6-panel results plot
API Reference
Experiment(type, metric, input, alpha, cuped, correction, is_ratio,power,alternative)
| Parameter | Type | Default | Description |
|---|---|---|---|
type |
str | 'ab' |
'ab' or 'multivariate' |
metric |
str | 'continous' |
'binary', 'continous', 'discrete' |
input |
str | required | 'csv' or 'direct' |
alpha |
float | 0.05 |
significance level |
cuped |
bool | False |
variance reduction (CSV only) |
correction |
str | None |
'bonf', 'holm', 'bh' |
is_ratio |
bool | False |
ratio metric — applies Delta method (CSV only) |
power |
float | 0.8 |
power |
alternative |
str | two-tail |
type of alternative hypothesis |
Usage Examples
Binary Metric — CSV Input
import pandas as pd
from abpy.router import Experiment
exp = Experiment(type='ab', metric='binary', input='csv', alpha=0.05)
exp.fit('experiment.csv') # requires columns: variant, metric
exp.results()
exp.plot()
CSV structure:
variant,metric
control,0
control,1
treatment,1
treatment,0
...
Continuous Metric — Direct Input
exp = Experiment(type='ab', metric='continous', input='direct', alpha=0.05)
exp.fit({
'control': {'n': 1000, 'mean': 45.2, 'std': 12.3},
'treatment': {'n': 1000, 'mean': 48.7, 'std': 11.9}
})
exp.results()
Discrete / Count Metric
exp = Experiment(type='ab', metric='discrete', input='direct', alpha=0.05)
exp.fit({
'control': {'n': 1000, 'events': 300},
'treatment': {'n': 1000, 'events': 375}
})
exp.results()
Multivariate with Holm Correction
exp = Experiment(
type='multivariate',
metric='binary',
input='direct',
alpha=0.05,
correction='holm'
)
exp.fit({
'control': {'n': 2000, 'successes': 200},
'treatment1': {'n': 2000, 'successes': 240},
'treatment2': {'n': 2000, 'successes': 260}
})
exp.results()
CUPED — Variance Reduction
Reduces variance using pre-experiment data, requiring fewer users to reach significance.
exp = Experiment(type='ab', metric='continous', input='csv', alpha=0.05, cuped=True)
exp.fit('experiment_with_pre.csv')
exp.results()
CSV structure for CUPED:
variant,metric,pre_metric
control,45.2,44.1
treatment,48.7,43.9
...
Ratio Metric — Delta Method
For metrics like CTR = clicks / impressions where both vary per user.
exp = Experiment(type='ab', metric='continous', input='csv', alpha=0.05, is_ratio=True)
exp.fit('ratio_experiment.csv')
exp.results()
CSV structure for ratio metrics:
variant,metric1,metric2
control,3,10
treatment,5,10
...
CSV structure for ratio metrics (if Cuped applied)
variant,metric1,metric2,pre_metric1,pre_metric2
control,3,10,2,6
treatment,5,10,3,5
...
Test Selection Logic
Binary metric → z-test (pooled proportion)
Continuous
n >= 30 → t-test (Welch's approximation)
n < 30 → Shapiro-Wilk normality check
normal → t-test
not normal → Mann-Whitney U
is_ratio=True → Delta method → t-test
Discrete/Count → Poisson rate test
Output
Every test returns:
Test used which statistical test was selected
Alternative type
Test statistic z, t, or U value
P-value raw p-value
Adjusted p-value if correction applied
Confidence interval 95% CI on the lift
Effect size Cohen's h (binary) or d (continuous)
Power achieved power given sample size
Required sample size The sample size for the given power to be achieved
Recommendation "2nd variant wins" / "no significant difference" / "underpowered"
Corrections Reference
| Method | Controls | Best For |
|---|---|---|
bonf |
FWER | Few comparisons, false positives very costly |
holm |
FWER | More powerful than Bonferroni, same control |
bh |
FDR | Many comparisons, some false positives acceptable |
Visualisation
exp.plot() # renders 6-panel dashboard
Dashboard panels:
- P-value vs significance threshold
- Confidence interval on lift
- Effect size (Cohen's benchmark)
- Power curve vs sample size
- Metric distribution by variant
- Cuped applied metric distribution by variant
Dependencies
numpy >= 1.21.0
scipy >= 1.7.0
pandas >= 1.3.0
matplotlib >= 3.4.0
License
MIT License — see LICENSE file for details.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file abpy_toolkit-0.1.1.tar.gz.
File metadata
- Download URL: abpy_toolkit-0.1.1.tar.gz
- Upload date:
- Size: 16.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d054453c198b95dbe030ead4c00f75c2eff82dfe026efd6866d99ce5856b1ca5
|
|
| MD5 |
1ad4eafcf8fced0b5c24e2cd0e517991
|
|
| BLAKE2b-256 |
d1e2c9f6c5e21eb3e41775d91d9067d51eb50ed1395c2d8595a1de6f54a1cec0
|
File details
Details for the file abpy_toolkit-0.1.1-py3-none-any.whl.
File metadata
- Download URL: abpy_toolkit-0.1.1-py3-none-any.whl
- Upload date:
- Size: 17.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
cc6e2fe029dd4eb6aaf931d27789d002d0443708a86b1122690bb7cdba502ca9
|
|
| MD5 |
0ab5a33037409eba4638ee1ecaf484f9
|
|
| BLAKE2b-256 |
40112bdfd247a4af0c5f7d5984345cc3002d1d5e59f868935f5aa1780fccbbfe
|