This release is a pre-release and may not be stable for production use.
HypEx: Advanced Causal Inference and AB Testing Toolkit
HypEx (Hypotheses and Experiments) is a library for causal inference and AB testing. It runs the same experiments on pandas for everyday analysis and on Apache Spark for data that does not fit on one machine.
What's new in 2.0
HypEx 2.0 adds an Apache Spark backend, is substantially faster than 1.0.x (most of all on large data, in A/A loops and in matching), and extends the statistical toolkit. Alpha: the API may still change. See the release notes for the full list.
- Spark backend.
Datasetruns on pandas or Spark with the same API for AA tests, AB tests, homogeneity tests and matching. - Speed. Vectorised splitting, Spark checkpointing, batched statistical tests, FAISS
shufflemode and co-partitioned search for distributed matching. - New statistics and blocks.
StatsUTest, matching bias correction and metrics,NaDropper,Float32Caster,AATest(dry_test=...). - CUPED and CUPAC variance reduction in
ABTest, with a variance reduction report. - Multiple-testing corrections in
ABTest(Holm by default, plus Bonferroni, Sidak, FDR and others). - Breaking changes against 1.0.7: Python 3.13 is not supported,
Output.resumeis nowOutput.summary, A/A splits use a different hash,Matchingextracts full data and indexes only on request. Version 0.1.x is no longer supported (pip install hypex==0.1.10).
Introduction
HypEx employs Rubin's Causal Model (RCM) for matching closely related pairs, ensuring equitable group comparisons when estimating treatment effects. Its automated pipeline calculates the Average Treatment Effect (ATE), Average Treatment Effect on the Treated (ATT) and Average Treatment Effect on the Control (ATC), with a standardized interface for running the estimations.
Beyond causal inference, HypEx provides AA tests, homogeneity tests and AB tests (including A/B/n tests with multiple testing corrections and CUPED/CUPAC variance reduction) to rigorously test hypotheses and validate experimental results.
Features
- Matching: Faiss-based nearest neighbor search (Mahalanobis or L2 distance), several neighbors per object, feature weights, matching inside groups, optional categorical encoding and bias estimation.
- Matching quality tests: SMD, KS, PSI, Repeats, t-test and chi2-test to check the robustness of the matching.
- AA test: repeated random splits with a selection of the most homogeneous split, optional stratification, early stopping and custom group sizes.
- Homogeneity test: compares target and feature distributions between groups.
- AB test: group difference with t-test, u-test and chi2-test, A/B/n tests with multiple testing corrections, CUPED and CUPAC variance reduction with a detailed report.
- Roles: describe your data once (
TreatmentRole,TargetRole,FeatureRole,StratificationRole,InfoRole) and every experiment picks what it needs. - Pandas and Spark backends behind one
DatasetAPI.
Warnings
Some functions in HypEx help to solve auxiliary tasks but cannot automate decisions on experiment design.
Note: For Matching, it is recommended to use no more than 7 features: more may lead to the curse of dimensionality and make the results unrepresentative.
Installation
pip install --pre -U hypex # 2.0 is in alpha: --pre is required
Without --pre pip installs the latest stable release (1.0.x). To pin the alpha: pip install hypex==2.0.0a1.
Optional extras for CUPAC models:
pip install -U "hypex[cat]" # CatBoost
pip install -U "hypex[lgbm]" # LightGBM
pip install -U "hypex[all]" # everything above
Requirements: Python >=3.8, <3.13. PySpark 3.5.1 is installed with the library; running on Spark also needs Java
(8, 11 or 17).
Quick start
Explore usage examples and tutorials here.
Matching example
from hypex.dataset import Dataset, InfoRole, TreatmentRole, TargetRole, FeatureRole
from hypex import Matching
data = Dataset(
roles={
"user_id": InfoRole(int), # InfoRole for ID
"treat": TreatmentRole(int), # TreatmentRole identifies the group (control or target)
"post_spends": TargetRole(float), # TargetRole for the target
},
data="data.csv",
default_role=FeatureRole(), # All remaining columns are features (used to search for similar objects)
)
test = Matching() # Classic matching (Mahalanobis distance + quality tests)
test = Matching(distance="l2") # Choose the distance
test = Matching(n_neighbors=3) # Several neighbors per object
test = Matching(group_match=True) # Match inside groups of categorical features
test = Matching(weights={"age": 2.0}) # Custom feature weights
test = Matching(extract_full_data=True, compute_indexes=True) # Also build full_data and indexes
result = test.execute(data)
result.summary # Summary of results (ATE, ATT, ATC)
result.quality_results # Matching quality tests
result.full_data # Wide dataset with pairs (requires extract_full_data=True)
result.indexes # Matched pairs of indexes, good for join (requires compute_indexes=True)
full_data and indexes are skipped by default because they are expensive on large data.
More about Matching here
AA-test example
from hypex.dataset import Dataset, InfoRole, TargetRole, StratificationRole
from hypex import AATest
data = Dataset(
roles={
"user_id": InfoRole(int), # InfoRole for ID
"pre_spends": TargetRole(), # TargetRole for the homogeneity check
"post_spends": TargetRole(), # TargetRole for the homogeneity check
"gender": StratificationRole(str), # StratificationRole for strata
},
data="data.csv",
)
aa = AATest(n_iterations=10)
result = aa.execute(data)
result.summary # Summary of the whole test
result.aa_score # AA score
result.best_split # The most homogeneous split
result.best_split_statistic # Statistics of the best split
More about AA test here
Homogeneity test example
from hypex import HomogeneityTest
# `data` has a TreatmentRole column and TargetRole columns, as above
result = HomogeneityTest().execute(data)
result.summary
More about the homogeneity test here
AB-test example
from hypex.dataset import Dataset, InfoRole, TreatmentRole, TargetRole
from hypex import ABTest
data = Dataset(
roles={
"user_id": InfoRole(int), # InfoRole for ID
"treat": TreatmentRole(), # TreatmentRole identifies the group (control or target)
"pre_spends": TargetRole(), # TargetRole for A/B(n) tests
"post_spends": TargetRole(), # TargetRole for A/B(n) tests
},
data="data.csv",
)
test = ABTest() # Classic A/B test
test = ABTest(multitest_method="bonferroni") # A/Bn test with Bonferroni correction
test = ABTest(additional_tests=["t-test", "u-test", "chi2-test"]) # Choose the tests
test = ABTest(cuped_features={"post_spends": "pre_spends"}) # CUPED variance reduction
test = ABTest(enable_cupac=True, cupac_models=["linear", "ridge"]) # CUPAC variance reduction
result = test.execute(data)
result.summary # Summary of results
result.multitest # Multiple testing corrections
result.sizes # Group sizes
result.variance_reduction_report # Variance reduction report for CUPED/CUPAC
CUPED and CUPAC
CUPED uses one pre-experiment value of the target. CUPAC predicts the target from pre-experiment covariates, one model
per time period transition. Its configuration comes from the roles: PreTargetRole(parent=..., lag=...) marks a
historical value of a target, FeatureRole(parent=..., lag=...) a covariate of the same period, and
TargetRole(cofounders=[...]) lists the covariates of the target.
from hypex import ABTest
from hypex.dataset import Dataset, FeatureRole, InfoRole, PreTargetRole, TargetRole, TreatmentRole
data = Dataset(
roles={
"d": TreatmentRole(),
"y": TargetRole(cofounders=["X1", "X2"]),
"y_lag1": PreTargetRole(parent="y", lag=1), # target one period ago
"X1_lag1": FeatureRole(parent="X1", lag=1),
"X2_lag1": FeatureRole(parent="X2", lag=1),
"y_lag2": PreTargetRole(parent="y", lag=2), # target two periods ago
"X1_lag2": FeatureRole(parent="X1", lag=2),
"X2_lag2": FeatureRole(parent="X2", lag=2),
},
data=df,
default_role=InfoRole(), # all other columns are ignored
)
result = ABTest(cuped_features={"y": "y_lag1"}).execute(data) # CUPED
result.variance_reduction_report
# CUPAC: the best model is selected for every transition
result = ABTest(enable_cupac=True, cupac_models=["linear", "ridge", "lasso", "catboost"]).execute(data)
result.cupac.variance_reductions # variance reduction per model
result.cupac.feature_importances # which covariates helped the most
"catboost" requires pip install "hypex[cat]". If cupac_models is omitted, all available models are tried.
Full guide: CUPED & CUPAC tutorial.
More about AB test here
Spark backend
Every experiment above runs on Spark with the same API. Create a SparkSession, pass it to Dataset and select the
backend explicitly. data can be a pandas.DataFrame, it is converted to a Spark DataFrame.
from pyspark.sql import SparkSession
from hypex import AATest
from hypex.dataset import Dataset, InfoRole, TargetRole
from hypex.utils import BackendsEnum
spark = SparkSession.builder.master("local[*]").appName("HypEx").getOrCreate()
data = Dataset(
roles={
"user_id": InfoRole(int),
"pre_spends": TargetRole(),
"post_spends": TargetRole(),
},
data=pandas_df,
session=spark,
backend=BackendsEnum.spark,
)
result = AATest(n_iterations=10, float32=True).execute(data)
result.summary
Notes:
- Computation is lazy and distributed. Spark pays off on large data; on small data pandas is faster.
Dataset.datais apyspark.sql.DataFrameon this backend, so inspect it with the Spark API.- Convert between backends with
dataset.to_backend(BackendsEnum.pandas)ordataset.to_backend(BackendsEnum.spark, session=spark). Converting a large Spark dataset to pandas is refused when it would not fit in memory. SmallDatasetis pandas-based by design: it holds compactanalysis_tablesand reporting results.- Supported runtime: Python
>=3.8, PySpark3.5.1, Java 8/11/17.
Each tutorial has a "Working with the Spark backend" section.
Documentation
For more detailed information about the library, visit our documentation on ReadTheDocs. It has guides and tutorials to get started and detailed API documentation for advanced use cases.
The architecture of the package (executors, ExperimentData, backends) is described in
hypex/README.md.
Contributions
Join our community! For guidelines on contributing, reporting issues or seeking support, please refer to our Contributing Guidelines.
More Information & Resources
Habr (ru) - discover how HypEx is revolutionizing causal
inference in various fields.
A/B testing seminar - Seminar in NoML about
matching and A/B testing
Matching with HypEx: Simple Guide -
Simple matching guide with explanation
Matching with HypEx: Grouping - Matching
with grouping guide
HypEx vs Causal Inference and DoWhy -
discover why HypEx is the best solution for causal inference
HypEx vs Causal Inference and DoWhy: part 2 -
discover why HypEx is the best solution for causal inference
Testing different libraries for the speed of matching
Visit this notebook on Kaggle and estimate results by yourself. The benchmark was run on an earlier HypEx version.
| Group size | 32 768 | 65 536 | 131 072 | 262 144 | 524 288 | 1 048 576 | 2 097 152 | 4 194 304 |
|---|---|---|---|---|---|---|---|---|
| Causal Inference | 46s | 169s | None | None | None | None | None | None |
| DoWhy | 9s | 19s | 40s | 77s | 159s | 312s | 615s | 1 235s |
| HypEx with grouping | 2s | 6s | 16s | 42s | 167s | 509s | 1 932s | 7 248s |
| HypEx without grouping | 2s | 7s | 21s | 101s | 273s | 982s | 3 750s | 14 720s |
Join Our Community
Have questions or want to discuss HypEx? Join our Telegram chat and connect with the community and the developers.
Metadata
Release files for HypEx 2.0.0a1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| hypex-2.0.0a1.tar.gz | 299.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| hypex-2.0.0a1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 655.9 kB
Release files / hypex-2.0.0a1.tar.gz
| Download URL | hypex-2.0.0a1.tar.gz |
|---|---|
| Size | 299.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
af0276e31e749d93d585937450dba6540a361179e9a2990502ede2f9f16bd397
|
|
BLAKE2b-256 checksum How to use checksums |
dc069713679b8cffce7961be994a0156be43b02e1db109ca6ef853b9ae39b100
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.
Transparency logRelease files / hypex-2.0.0a1-py3-none-any.whl
| Download URL | hypex-2.0.0a1-py3-none-any.whl |
|---|---|
| Size | 356.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
af9a572e9aff1b5ef62f82a58bf6d721e824186c6b031b21a2ff0dee79f9e02f
|
|
BLAKE2b-256 checksum How to use checksums |
46e1b4eeaf18aa07dfb6f54d1d1cbbeaeffa353f09f272e1d9331116ba4d5f3b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.
Transparency log