PowerBench
Statistical power analysis that shows its work.
PowerBench computes required sample sizes for 27 study designs. What distinguishes it is not the calculations — those are standard — but that every number is checked three ways and the checks are published:
- an analytic reference implementing a named published formula,
- an independent Monte Carlo simulation that shares no code with the reference,
- external tools (R
pwr,pwrss,metafor,TOSTER,MASS) where they can represent the same design.
When those disagree, PowerBench says so and explains why, rather than presenting one confident number.
This package is the calculation engine behind PowerMate. It is published separately so the mathematics can be audited, cited, and reported against without running the web application.
Why this exists
On 2026-08-03 an audit of this engine found six analytic formulas producing wrong sample
sizes. The worst recommended 191 participants for a design needing about 614 — a study
designed at roughly a third of the power its authors would believe it had. MASS::polr
measured 0.34 power where the tool promised 0.80.
Every one of those defects passed the entire test suite.
The reason is worth stating plainly, because it generalises to any power tool: each sample-size function is a search loop of the form
while power(n) < target:
n += 1
so assert power >= target_power is a tautology. It holds no matter how wrong
power() is. A test suite built on that assertion cannot fail.
What can falsify a formula is an independent implementation of the same design. That is what the simulation layer and the external adapters are for, and why their agreement is tested rather than merely displayed.
All six defects are fixed, each confirmed against an independent R implementation. See
CHANGELOG.md.
Install
pip install powerbench
Requires Python 3.11+. Runtime dependencies are SciPy and NumPy. R is optional and needed only to regenerate external-tool comparisons.
Quick start
from powerbench.schema import Scenario, DecisionRule
from powerbench.references import analytic_reference
from powerbench.simulation import simulate
scenario = Scenario(
id="my-study",
title="Two independent groups, d = 0.4",
framework="frequentist",
goal="required_sample_size",
model="two_sample_t",
target_power=0.80,
effect={"cohens_d": 0.4},
design={},
decision_rule=DecisionRule(alpha=0.05, alternative="two.sided"),
simulation={"replications": 5000, "seed": 20260803},
assumptions=["Independent observations", "Equal variances"],
)
scenario.validate()
reference = analytic_reference(scenario)
print(reference["n_total"], reference["method"])
# 200 noncentral_t
check = simulate(scenario, n1=reference["n1"], n2=reference["n2"])
print(check["power"], check["mc_ci95"])
# 0.8092 (0.7983, 0.8201) <- independent confirmation, not the same code path
The second call is the point. If the analytic answer falls outside the simulation's Monte Carlo interval, one of them is wrong, and you have learned something the single number would never have told you.
What it covers
27 designs, each with a declared estimand:
| Family | Designs |
|---|---|
| t tests | pooled two-sample, Welch unequal-variance, paired |
| Proportions | one-sample (exact binomial), two-sample |
| Chi-square | goodness-of-fit, independence |
| ANOVA | one-way, factorial, ANCOVA, planned contrast |
| Regression | overall, incremental block, moderation, logistic, Poisson, ordinal |
| Association | Pearson correlation |
| Mediation | indirect effect via joint significance |
| Equivalence | TOST, non-inferiority, Bayesian ROPE |
| Meta-analysis | fixed-effect, random-effects (Hartung–Knapp) |
| Sequential | group-sequential t with Lan–DeMets O'Brien–Fleming boundaries |
| Survival | two-arm log-rank (Schoenfeld) |
| Bayesian | JZS Bayes-factor design analysis |
Coverage is not the same as verification. See Verification status below for which designs have been checked against what.
Verification status
The alignment gate
tests/test_reference_simulation_alignment.py runs every committed scenario through both
the analytic reference and the independent simulator, across three fixed seeds at 12,000
replications, and requires agreement on a majority.
pytest -m slow tests/test_reference_simulation_alignment.py
Latest full run: 29 passed, 0 failed.
Multiple seeds are not decoration. At the 2,000 replications the scenarios ship with, agreement is a seed lottery: one method read as broken on one seed and agreed comfortably at 12,000, while another looked fine at 2,000 and was genuinely wrong.
Two scenarios carry documented divergences with written case studies in
docs/discrepancies/. They are still tested — against a declared
envelope with an expected direction and a maximum gap — so a sign flip or a widening gap
still fails. An entry that stops being needed is forced out by
test_allowlist_has_no_stale_entries, so nothing gets permanent amnesty.
External tools
12 of 28 scenarios have been checked against an external implementation:
| Design | PowerBench | External | Tool |
|---|---|---|---|
| Two-sample t | 200 | 200 | R pwr |
| Paired t | 52 | 52 | R pwr |
| Chi-square GOF | 181 | 181 | R pwr |
| One-way ANOVA | 159 | 159 | R pwr |
| Linear regression | 77 | 77 | R pwr |
| Incremental regression | 58 | 58 | R pwr |
| Correlation | 194 | 194 | R pwr |
| Logistic regression | 856 | 856 | R pwrss |
| Poisson regression | 796 | 787 | R pwrss |
| TOST equivalence | 140 | 140 | TOSTER |
| Ordinal regression | 614 | 624 | MASS::polr |
| Random-effects meta | 6,504 | 5,088 | metafor — investigated |
The remaining 16 have no external coverage, and the tooling reports that honestly rather than implying agreement.
Disagreements are attributed, not scored
A comparison that only compares numbers cannot distinguish "these tools disagree" from "these tools are answering different questions".
Concretely: pwrss.z.logreg returns N = 249 where PowerBench returns 856 for what
looks like the same logistic design — a 3.4× gap that reads as catastrophic. It is not a
defect. pwrss defaults to a standard-normal predictor, so it powers a one-standard-deviation
change, while PowerBench's binary branch contrasts two equally sized groups. Direct glm
Monte Carlo measures 0.31 power at N = 249 for a balanced binary design. Called with a
matched Bernoulli predictor, pwrss returns 856 exactly.
So every adapter declares its own effect definition, predictor distribution, test variant,
degrees-of-freedom convention, allocation and tails
(powerbench/parameterization.py), and declarations are
compared before the numbers. Two tools answering different questions cannot launder a
coincidental agreement into "compatible".
probable_defect is never assigned automatically. An unexplained gap becomes
needs_review for a human, because asserting that a named third-party tool is wrong is a
claim a person should make after looking. An undeclared field is treated as unknown, so a
silent adapter can never explain away a gap it never accounted for.
Versioning
Two levels, because they answer different questions:
powerbench.__version__— the package release.powerbench.METHOD_VERSIONS[method_id]— that method's answers, bumped only when a computed number changes for some input.
import powerbench
powerbench.method_fingerprint("ordinal_regression")
# 'ordinal_regression@2.0.0'
for change in powerbench.changes_since("ordinal_regression", "1.0.0"):
print(change.numeric_impact)
# Required N for OR=1.5, k=4, 80% power moves from 191 to 614...
powerbench.changes_since("two_sample_t", "1.0.0")
# [] <- the ordinal correction did not touch the t test
Cache keys, published study plans and validation records should bind to
(method_id, method_version). A correction to one method must not invalidate work that
depended on another, and changes_since reports the numeric impact so a stale result is
actionable rather than merely void.
Reproducing the checks
pytest -q # fast suite
pytest -m slow tests/test_reference_simulation_alignment.py # the correctness gate (~11 min)
External-tool fixtures in data/benchmarks/ each ship with the script that generated them
in scripts/benchmarks/, so a reviewer can re-derive every committed
number rather than trusting a prose source field:
Rscript scripts/benchmarks/ordinal_regression_polr.R
Requires R with pwr, pwrss, TOSTER, metafor and MASS. Tests that need R skip
cleanly where it is absent rather than failing for an environment reason.
Reporting a suspected error
This is the contribution we most want. If a number here disagrees with a tool you trust, please open an issue using the calculation error template, which asks for the method version, exact inputs, both numbers, and what produced your expectation.
You do not need to be certain. Many apparent disagreements turn out to be estimand
differences — and those still become useful case studies. See
CONTRIBUTING.md for the bar a formula change has to clear.
Status and limits
Beta. The engine is used in production by PowerMate, the alignment gate passes on every committed scenario, and the corrections above are confirmed against independent R implementations. It has not yet been reviewed by anyone outside the project — which is precisely why it is public.
Known boundaries, stated rather than buried:
- Mixed, clustered and longitudinal models are not covered by closed-form paths.
- Effect-size choice is yours. No power tool can tell you the smallest effect worth detecting, and planning from a small pilot's observed effect systematically under-powers studies.
- Random-effects meta-analysis has a nearly flat power curve once between-study heterogeneity dominates; the required N is only weakly identified there, which the case study explains.
Licence
MIT. See LICENSE.
Citation
If PowerBench contributes to published work, please cite the repository and record the
method version you used (powerbench.method_fingerprint(...)), so the calculation can be
reproduced exactly.
Metadata
Release files for powerbench 0.8.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| powerbench-0.8.2.tar.gz | 96.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| powerbench-0.8.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 176.1 kB
Release files / powerbench-0.8.2.tar.gz
| Download URL | powerbench-0.8.2.tar.gz |
|---|---|
| Size | 96.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
3bd12698e5c9551329f5afd38c7cec6e51bd29a0e27e9977602fd0ed8d13138e
|
|
BLAKE2b-256 checksum How to use checksums |
c7ec332c8f45ad53ac998cfa960270abdb492e6756622110153327839ba268c3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 5, 2026.
Transparency logRelease files / powerbench-0.8.2-py3-none-any.whl
| Download URL | powerbench-0.8.2-py3-none-any.whl |
|---|---|
| Size | 79.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
130a7d7fefbe6a0e6a420b5db9b63948a1bb3552839a24d44b220b852de7a73f
|
|
BLAKE2b-256 checksum How to use checksums |
1c2310fcddc880e332421003c8a3911aef0b719cc8315f7110f44e3ef4f20667
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 5, 2026.
Transparency log