Skip to main content

ssepy: A Library for Efficient Model Evaluation through Stratification, Sampling, and Estimation in Python

Paper Apache-2.0

Given an unlabeled dataset and model predictions, how can we select which instances to annotate in one go to maximize the precision of our estimates of model performance on the entire dataset?

The ssepy package helps you do that! The implementation of the ssepy package revolves around the following sequential framework:

  1. Predict: Predict the expected model performance for each example.
  2. Stratify: Divide the dataset into strata using the base predictions.
  3. Sample: Sample a data subset using the chosen sampling method.
  4. Annotate: Acquire annotations for the sampled subset.
  5. Estimate: Estimate model performance.

See our paper here for a technical overview of the framework.

Getting started

In order to intall the package, run

pip install ssepy

Alternatively, clone the repo, cd into it, and run

pip install .

You may want to initialize a conda environment before running this operation.

Test your setup using this example, which demonstrates data stratification, n allocation for annotation via proportional allocation, sampling via stratified simple random sampling, and estimation using the Horvitz-Thompson estimator:

import numpy as np
from sklearn.cluster import KMeans
from ssepy import ModelPerformanceEvaluator

np.random.seed(0)
# Generate data
N = 100000
Y = np.random.normal(0, 1, N) # Ground truth

# Unobserved target
print(np.mean(Y))

n = 100 # Annotation n
# 1. Proxy for ground truth
Yh = Y + np.random.normal(0, 0.1, N)
evaluator = ModelPerformanceEvaluator(Yh = Yh, budget = n) # Initialize evaluator
# 2. Stratify on Yh
evaluator.stratify_data(clustering_algo=KMeans(n_clusters=5, random_state=0, n_init="auto"), X=Yh) # 5 strata
# 3. Allocate n with proportional allocation and sample
evaluator.allocate_budget(allocation_type="proportional")
sampled_idx = evaluator.sample()
# 4. Annotate
Yl = Y[sampled_idx]
# 5. Estimate target and variance of estimate
estimate, variance_estimate = evaluator.compute_estimate(Yl, estimator="ht")
print(estimate, variance_estimate)

For the difference estimator under simple random sampling, run

evaluator = ModelPerformanceEvaluator(Yh=Yh, budget=n) # initialize sampler
sampled_idx = evaluator.sample(sampling_method="srs") # 3. sample
Yl = Y[sampled_idx] # 4. annotate
estimate, variance_estimate = evaluator.compute_estimate(Yl, estimator="df") # 5. estimate
print(estimate, variance_estimate)

See also some examples in the associated folder.

Features

The supported sample designs are: (SRS) simple random sampling without replacement, (SSRS) stratified simple random sampling without replacement with proportional and optimal/Neyman allocation, (Poisson) sampling. All sampling methods have associated (HT) Horvitz-Thompson and (DF) difference estimators.

Bugs and contribute

Feel free to reach out if you find any bugs or you would like other features to be implemented in the package.

Release files for ssepy 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ssepy 0.1.1
File Size Uploaded
ssepy-0.1.1.tar.gz 11.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ssepy 0.1.1
File Interpreter ABI Platform
ssepy-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 24.3 kB

Release files / ssepy-0.1.1.tar.gz

Download URL ssepy-0.1.1.tar.gz
Size 11.2 kB
Tags Source
SHA-256 checksum
How to use checksums
78b6f2bcc366d3a96ce42cfa00932b4964583205a6b38714fc5c82d3f5a4dfa0
BLAKE2b-256 checksum
How to use checksums
4495e3a9c89996aa640c5bace762297fdde43c288f878582becbdfb933988295
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/2.1.1 CPython/3.10.16 Darwin/24.3.0

Release files / ssepy-0.1.1-py3-none-any.whl

Download URL ssepy-0.1.1-py3-none-any.whl
Size 13.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
70a84c1bc84b04e10a7e289df6d4748cd3fb419809d85f678b287fb7e32f3fcc
BLAKE2b-256 checksum
How to use checksums
adab5ac9271abf276280016f7416951a2514edc6c9f390580fd131fe59123424
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/2.1.1 CPython/3.10.16 Darwin/24.3.0

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page