Skip to main content

Probabilistic Win Rate Comparisons

prob_wins is a Python package for performing pairwise comparisons of machine learning models with win rates. It computes the proportion of wins, losses and ties, generates several win rate effect size statistics, and can perform statistical inference tests.

Installation

Installation can be done using from pypi can be done using pip:

pip install prob_wins

Or, if you're using uv, simply run:

uv add prob_wins

The project currently depends on the following packages:

Dependency tree
prob-wins
├── jaxtyping
├── numpy
└── scipy

Development Environment

This project was developed using uv. To install the development environment, simply clone this github repo:

git clone https://github.com/ioverho/prob_wins.git

And then run the uv sync --dev command:

uv sync --dev

The development dependencies should automatically install into the .venv folder.

Documentation

Let's say you have some evaluation results from your model, eval_results_model, and from a baseline model, eval_results_baseline, and you want to assess how much better your model is than baseline. prob_wins provides two frameworks for doing this assessment.

A more detailed explanation and derivation of the various win rate statistics and testing is provided in an accompanying blog post.

Check out the ./analysis/ folder for some examples.

Frequentist

Running the prob_wins.compare_paired_win_rates_frequentist method will return a FrequentistComparisonResult object that contains all necessary statistics for assessing your win rates in a Frequentist framework.

import prob_wins

pairwise_comparison_results = prob_wins.compare_paired_win_rates_frequentist(
    results=eval_results_model,
    baseline_results=eval_results_baseline,
    confidence_level=0.95,
)

Specifically, it contains:

  1. The proportion of times your model beat the baseline, and vice versa
    1. prob_win (float): Estimated probability that the system outperforms the baseline
    2. prob_win_ci (ConfidenceInterval): Confidence interval for prob_win
    3. prob_loss (float): Estimated probability that the system underperforms the baseline
    4. prob_loss_ci (ConfidenceInterval): Confidence interval for prob_loss
    5. prob_tie (float): Estimated probability of a tie
    6. prob_tie_ci (ConfidenceInterval): Confidence interval for prob_tie
  2. Win rate statistics
    1. wr (float): Win ratio — ratio of wins to losses (prob_win / prob_loss)
    2. wr_ci (ConfidenceInterval): Confidence interval for win ratio
    3. wo (float): Win odds — (wins + 0.5ties) / (losses + 0.5ties)
    4. wo_ci (ConfidenceInterval): Confidence interval for win odds
    5. nb (float): Net benefit — difference of win and loss probabilities (prob_win - prob_loss)
    6. nb_ci (ConfidenceInterval): Confidence interval for net benefit
  3. Statistical test results
    1. test_statistic (float): Sign test statistic: (wins - losses) / sqrt(wins + losses)
    2. test_p_val_one_sided (float): One-sided p-value for the sign test (H1: system > baseline)
    3. test_p_val_two_sided (float): Two-sided p-value for the sign test (H1: system =/= baseline)

Bayesian

Running the prob_wins.compare_paired_win_rates_bayesian method will return a BayesianComparisonResult object that contains all necessary statistics for assessing your win rates in a Bayesian framework.

import prob_wins

pairwise_comparison_results = prob_wins.compare_paired_win_rates_bayesian(
    results=eval_results_model,
    baseline_results=eval_results_baseline,
    seed=0,
    confidence_level=0.95,
    prior_strategy=[1, 1, 1],
    num_samples=10000,
    min_sig_diff=1.00,
)

The BayesianComparisonResult object contains largely the same elements (except that point statistics are now posterior medians, and confidence intervals become credibility intervals), except for the p values. These are replaced with:

  1. p_direction (float): Probability of direction — posterior mass where test_statistic > 0
  2. p_sig_neg (float): Posterior mass below the negative ROPE boundary (system is worse)
  3. p_sig_pos (float): Posterior mass above the positive ROPE boundary (system is better)
  4. p_sig_bidirectional (float): Total posterior mass outside the ROPE (p_sig_neg + p_sig_pos)

The ROPE is the region of practical equivalence, wherein the difference between the models might not be exactly 0, but it might as well be. The ROPE boundaries are determined by the min_sig_diff parameter.

Citation

@software{ioverho_prob_wins,
    author = {Verhoeven, Ivo},
    license = {MIT},
    title = {{prob\_wins}},
    url = {https://github.com/ioverho/prob_wins}
}

Metadata

Release files for prob_wins 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for prob_wins 0.1.0
File Size Uploaded
prob_wins-0.1.0.tar.gz 12.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for prob_wins 0.1.0
File Interpreter ABI Platform
prob_wins-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 28.1 kB

Release files / prob_wins-0.1.0.tar.gz

Download URL prob_wins-0.1.0.tar.gz
Size 12.4 kB
Tags Source
SHA-256 checksum
How to use checksums
27e6bdac1a917396c4a2028792a0893a432cb1550dedbf374796ada0299c3bea
BLAKE2b-256 checksum
How to use checksums
045544253a6c917ab5a3f9210d9884f12fb1613ac2ce096411716745e90b09d4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.11.24 {"installer":{"name":"uv","version":"0.11.24","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"SlimbookOS","version":"24.0","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / prob_wins-0.1.0-py3-none-any.whl

Download URL prob_wins-0.1.0-py3-none-any.whl
Size 15.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
025cfa539763f9084ead7bf2d19afe85d94b5f79f3c2fd115d267eeb8f508480
BLAKE2b-256 checksum
How to use checksums
71b79b27dcd7da18b210e90cb4fbeb7a06cb954c5c2d62de20dd2c102ef6b200
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.11.24 {"installer":{"name":"uv","version":"0.11.24","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"SlimbookOS","version":"24.0","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page