Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

HypEx: Advanced Causal Inference and AB Testing Toolkit

PyPI version Python versions License Telegram

HypEx (Hypotheses and Experiments) is a library for causal inference and AB testing. It runs the same experiments on pandas for everyday analysis and on Apache Spark for data that does not fit on one machine.

What's new in 2.0

HypEx 2.0 adds an Apache Spark backend, is substantially faster than 1.0.x (most of all on large data, in A/A loops and in matching), and extends the statistical toolkit. Alpha: the API may still change. See the release notes for the full list.

  • Spark backend. Dataset runs on pandas or Spark with the same API for AA tests, AB tests, homogeneity tests and matching.
  • Speed. Vectorised splitting, Spark checkpointing, batched statistical tests, FAISS shuffle mode and co-partitioned search for distributed matching.
  • New statistics and blocks. StatsUTest, matching bias correction and metrics, NaDropper, Float32Caster, AATest(dry_test=...).
  • CUPED and CUPAC variance reduction in ABTest, with a variance reduction report.
  • Multiple-testing corrections in ABTest (Holm by default, plus Bonferroni, Sidak, FDR and others).
  • Breaking changes against 1.0.7: Python 3.13 is not supported, Output.resume is now Output.summary, A/A splits use a different hash, Matching extracts full data and indexes only on request. Version 0.1.x is no longer supported (pip install hypex==0.1.10).

Introduction

HypEx employs Rubin's Causal Model (RCM) for matching closely related pairs, ensuring equitable group comparisons when estimating treatment effects. Its automated pipeline calculates the Average Treatment Effect (ATE), Average Treatment Effect on the Treated (ATT) and Average Treatment Effect on the Control (ATC), with a standardized interface for running the estimations.

Beyond causal inference, HypEx provides AA tests, homogeneity tests and AB tests (including A/B/n tests with multiple testing corrections and CUPED/CUPAC variance reduction) to rigorously test hypotheses and validate experimental results.

Features

  • Matching: Faiss-based nearest neighbor search (Mahalanobis or L2 distance), several neighbors per object, feature weights, matching inside groups, optional categorical encoding and bias estimation.
  • Matching quality tests: SMD, KS, PSI, Repeats, t-test and chi2-test to check the robustness of the matching.
  • AA test: repeated random splits with a selection of the most homogeneous split, optional stratification, early stopping and custom group sizes.
  • Homogeneity test: compares target and feature distributions between groups.
  • AB test: group difference with t-test, u-test and chi2-test, A/B/n tests with multiple testing corrections, CUPED and CUPAC variance reduction with a detailed report.
  • Roles: describe your data once (TreatmentRole, TargetRole, FeatureRole, StratificationRole, InfoRole) and every experiment picks what it needs.
  • Pandas and Spark backends behind one Dataset API.

Warnings

Some functions in HypEx help to solve auxiliary tasks but cannot automate decisions on experiment design.

Note: For Matching, it is recommended to use no more than 7 features: more may lead to the curse of dimensionality and make the results unrepresentative.

Installation

pip install --pre -U hypex   # 2.0 is in alpha: --pre is required

Without --pre pip installs the latest stable release (1.0.x). To pin the alpha: pip install hypex==2.0.0a1.

Optional extras for CUPAC models:

pip install -U "hypex[cat]"   # CatBoost
pip install -U "hypex[lgbm]"  # LightGBM
pip install -U "hypex[all]"   # everything above

Requirements: Python >=3.8, <3.13. PySpark 3.5.1 is installed with the library; running on Spark also needs Java (8, 11 or 17).

Quick start

Explore usage examples and tutorials here.

Matching example

from hypex.dataset import Dataset, InfoRole, TreatmentRole, TargetRole, FeatureRole
from hypex import Matching

data = Dataset(
    roles={
        "user_id": InfoRole(int),  # InfoRole for ID
        "treat": TreatmentRole(int),  # TreatmentRole identifies the group (control or target)
        "post_spends": TargetRole(float),  # TargetRole for the target
    },
    data="data.csv",
    default_role=FeatureRole(),  # All remaining columns are features (used to search for similar objects)
)

test = Matching()  # Classic matching (Mahalanobis distance + quality tests)
test = Matching(distance="l2")  # Choose the distance
test = Matching(n_neighbors=3)  # Several neighbors per object
test = Matching(group_match=True)  # Match inside groups of categorical features
test = Matching(weights={"age": 2.0})  # Custom feature weights
test = Matching(extract_full_data=True, compute_indexes=True)  # Also build full_data and indexes

result = test.execute(data)
result.summary  # Summary of results (ATE, ATT, ATC)
result.quality_results  # Matching quality tests
result.full_data  # Wide dataset with pairs (requires extract_full_data=True)
result.indexes  # Matched pairs of indexes, good for join (requires compute_indexes=True)

full_data and indexes are skipped by default because they are expensive on large data.

More about Matching here

AA-test example

from hypex.dataset import Dataset, InfoRole, TargetRole, StratificationRole
from hypex import AATest

data = Dataset(
    roles={
        "user_id": InfoRole(int),  # InfoRole for ID
        "pre_spends": TargetRole(),  # TargetRole for the homogeneity check
        "post_spends": TargetRole(),  # TargetRole for the homogeneity check
        "gender": StratificationRole(str),  # StratificationRole for strata
    },
    data="data.csv",
)

aa = AATest(n_iterations=10)
result = aa.execute(data)

result.summary  # Summary of the whole test
result.aa_score  # AA score
result.best_split  # The most homogeneous split
result.best_split_statistic  # Statistics of the best split

More about AA test here

Homogeneity test example

from hypex import HomogeneityTest

# `data` has a TreatmentRole column and TargetRole columns, as above
result = HomogeneityTest().execute(data)
result.summary

More about the homogeneity test here

AB-test example

from hypex.dataset import Dataset, InfoRole, TreatmentRole, TargetRole
from hypex import ABTest

data = Dataset(
    roles={
        "user_id": InfoRole(int),  # InfoRole for ID
        "treat": TreatmentRole(),  # TreatmentRole identifies the group (control or target)
        "pre_spends": TargetRole(),  # TargetRole for A/B(n) tests
        "post_spends": TargetRole(),  # TargetRole for A/B(n) tests
    },
    data="data.csv",
)

test = ABTest()  # Classic A/B test
test = ABTest(multitest_method="bonferroni")  # A/Bn test with Bonferroni correction
test = ABTest(additional_tests=["t-test", "u-test", "chi2-test"])  # Choose the tests
test = ABTest(cuped_features={"post_spends": "pre_spends"})  # CUPED variance reduction
test = ABTest(enable_cupac=True, cupac_models=["linear", "ridge"])  # CUPAC variance reduction

result = test.execute(data)
result.summary  # Summary of results
result.multitest  # Multiple testing corrections
result.sizes  # Group sizes
result.variance_reduction_report  # Variance reduction report for CUPED/CUPAC

CUPED and CUPAC

CUPED uses one pre-experiment value of the target. CUPAC predicts the target from pre-experiment covariates, one model per time period transition. Its configuration comes from the roles: PreTargetRole(parent=..., lag=...) marks a historical value of a target, FeatureRole(parent=..., lag=...) a covariate of the same period, and TargetRole(cofounders=[...]) lists the covariates of the target.

from hypex import ABTest
from hypex.dataset import Dataset, FeatureRole, InfoRole, PreTargetRole, TargetRole, TreatmentRole

data = Dataset(
    roles={
        "d": TreatmentRole(),
        "y": TargetRole(cofounders=["X1", "X2"]),
        "y_lag1": PreTargetRole(parent="y", lag=1),  # target one period ago
        "X1_lag1": FeatureRole(parent="X1", lag=1),
        "X2_lag1": FeatureRole(parent="X2", lag=1),
        "y_lag2": PreTargetRole(parent="y", lag=2),  # target two periods ago
        "X1_lag2": FeatureRole(parent="X1", lag=2),
        "X2_lag2": FeatureRole(parent="X2", lag=2),
    },
    data=df,
    default_role=InfoRole(),  # all other columns are ignored
)

result = ABTest(cuped_features={"y": "y_lag1"}).execute(data)  # CUPED
result.variance_reduction_report

# CUPAC: the best model is selected for every transition
result = ABTest(enable_cupac=True, cupac_models=["linear", "ridge", "lasso", "catboost"]).execute(data)
result.cupac.variance_reductions  # variance reduction per model
result.cupac.feature_importances  # which covariates helped the most

"catboost" requires pip install "hypex[cat]". If cupac_models is omitted, all available models are tried.

Full guide: CUPED & CUPAC tutorial.

More about AB test here

Spark backend

Every experiment above runs on Spark with the same API. Create a SparkSession, pass it to Dataset and select the backend explicitly. data can be a pandas.DataFrame, it is converted to a Spark DataFrame.

from pyspark.sql import SparkSession

from hypex import AATest
from hypex.dataset import Dataset, InfoRole, TargetRole
from hypex.utils import BackendsEnum

spark = SparkSession.builder.master("local[*]").appName("HypEx").getOrCreate()

data = Dataset(
    roles={
        "user_id": InfoRole(int),
        "pre_spends": TargetRole(),
        "post_spends": TargetRole(),
    },
    data=pandas_df,
    session=spark,
    backend=BackendsEnum.spark,
)

result = AATest(n_iterations=10, float32=True).execute(data)
result.summary

Notes:

  • Computation is lazy and distributed. Spark pays off on large data; on small data pandas is faster.
  • Dataset.data is a pyspark.sql.DataFrame on this backend, so inspect it with the Spark API.
  • Convert between backends with dataset.to_backend(BackendsEnum.pandas) or dataset.to_backend(BackendsEnum.spark, session=spark). Converting a large Spark dataset to pandas is refused when it would not fit in memory.
  • SmallDataset is pandas-based by design: it holds compact analysis_tables and reporting results.
  • Supported runtime: Python >=3.8, PySpark 3.5.1, Java 8/11/17.

Each tutorial has a "Working with the Spark backend" section.

Documentation

For more detailed information about the library, visit our documentation on ReadTheDocs. It has guides and tutorials to get started and detailed API documentation for advanced use cases.

The architecture of the package (executors, ExperimentData, backends) is described in hypex/README.md.

Contributions

Join our community! For guidelines on contributing, reporting issues or seeking support, please refer to our Contributing Guidelines.

More Information & Resources

Habr (ru) - discover how HypEx is revolutionizing causal inference in various fields.
A/B testing seminar - Seminar in NoML about matching and A/B testing
Matching with HypEx: Simple Guide - Simple matching guide with explanation
Matching with HypEx: Grouping - Matching with grouping guide
HypEx vs Causal Inference and DoWhy - discover why HypEx is the best solution for causal inference
HypEx vs Causal Inference and DoWhy: part 2 - discover why HypEx is the best solution for causal inference

Testing different libraries for the speed of matching

Visit this notebook on Kaggle and estimate results by yourself. The benchmark was run on an earlier HypEx version.

Group size 32 768 65 536 131 072 262 144 524 288 1 048 576 2 097 152 4 194 304
Causal Inference 46s 169s None None None None None None
DoWhy 9s 19s 40s 77s 159s 312s 615s 1 235s
HypEx with grouping 2s 6s 16s 42s 167s 509s 1 932s 7 248s
HypEx without grouping 2s 7s 21s 101s 273s 982s 3 750s 14 720s

Join Our Community

Have questions or want to discuss HypEx? Join our Telegram chat and connect with the community and the developers.

Metadata

Release files for HypEx 2.0.0a1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for HypEx 2.0.0a1
File Size Uploaded
hypex-2.0.0a1.tar.gz 299.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for HypEx 2.0.0a1
File Interpreter ABI Platform
hypex-2.0.0a1-py3-none-any.whl Python 3 none any Details

Total release size: 655.9 kB

Release files / hypex-2.0.0a1.tar.gz

Download URL hypex-2.0.0a1.tar.gz
Size 299.5 kB
Tags Source
SHA-256 checksum
How to use checksums
af0276e31e749d93d585937450dba6540a361179e9a2990502ede2f9f16bd397
BLAKE2b-256 checksum
How to use checksums
dc069713679b8cffce7961be994a0156be43b02e1db109ca6ef853b9ae39b100
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release files / hypex-2.0.0a1-py3-none-any.whl

Download URL hypex-2.0.0a1-py3-none-any.whl
Size 356.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
af9a572e9aff1b5ef62f82a58bf6d721e824186c6b031b21a2ff0dee79f9e02f
BLAKE2b-256 checksum
How to use checksums
46e1b4eeaf18aa07dfb6f54d1d1cbbeaeffa353f09f272e1d9331116ba4d5f3b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page