benchbro
benchbro is a Python benchmarking library and CLI with parameterized discovery,
noise-aware statistics, reproducible environments, and rich terminal output.
Read the documentation or jump straight to the quick start.
Quick start
Add the CLI to a project:
uv add --dev benchbro
For a cloned development checkout, use uv sync --all-groups instead. A global
installation with uv tool install benchbro is also supported; in that case run
benchbro directly instead of uv run benchbro in the examples below.
BenchBro bundles a version-matched agent skill. After installing the package, add it to the current project with:
uvx library-skills --skill use-benchbro --yes
Create benchmark cases in any importable module:
from benchbro import Case, system
@system(scope="module")
def salt() -> bytes:
return b"-fixture"
case = Case(name="hashing", case_type="cpu", metric_type="time", tags=["fast", "core"])
@case.input()
def payload() -> bytes:
return b"benchbro"
@case.benchmark()
def sha1(payload: bytes, salt: bytes) -> str:
import hashlib
return hashlib.sha1(payload + salt).hexdigest()
@case.benchmark()
def sha256(payload: bytes, salt: bytes) -> str:
import hashlib
return hashlib.sha256(payload + salt).hexdigest()
Standalone systems behave like lightweight fixtures:
- declare them with
@system()outside anyCase - inject them into benchmarks and other systems by parameter name
- choose
scope="function"(default) orscope="module" - systems may be sync, async, yield-based, or async-yield-based
- async inputs, systems, benchmarks, and teardown share one managed event loop
- keep them separate from inputs: inputs cannot depend on systems, and systems cannot depend on inputs
Subprocess benchmarks use a dedicated subprocess API:
import sys
from benchbro import Case, CommandSpec
case = Case(name="cli", metric_type="time")
@case.subprocess()
def python_cli() -> CommandSpec:
return CommandSpec(argv=[sys.executable, "-c", "pass"])
@case.subprocess()benchmarks are time-only in v1.- Commands are explicit argv, not shell strings.
- Unexpected exit codes and timeouts fail the run.
- Long-lived subprocesses should be started in
@system(...)and consumed by benchmarks or subprocess benchmarks.
Regression thresholds default to 50.0 percent warning and 100.0 percent error at the case level, and can be overridden per benchmark:
case = Case(name="hashing", warning_threshold_pct=5.0, regression_threshold_pct=10.0)
@case.benchmark(warning_threshold_pct=2.0, regression_threshold_pct=3.0)
def critical_path(payload: bytes) -> str: ...
Comparison metric can be configured at case-level and benchmark-level:
case = Case(name="hashing", comparison_metric="p95_s")
@case.benchmark(comparison_metric="median_s")
def critical_path(payload: bytes) -> str: ...
Valid comparison metrics:
- time:
median_s(default),mean_s,iqr_s,p95_s,stddev_s,ops_per_sec - memory:
peak_alloc_bytes(default),net_alloc_bytes,peak_alloc_bytes_max
GC is disabled during measured iterations by default. To keep the interpreter GC behavior unchanged, set:
Case(name="hashing", gc_control="inherit")
Run benchmarks:
uv run benchbro run --repeats 10 --warmup 2
The defaults are 20 repeats, 50 measured iterations per repeat, and 5 warmup iterations. Percentiles are calculated across repeat-level per-iteration means.
For adaptive sampling, give BenchBro a precision target and/or time budget:
case = Case(
name="hashing",
adaptive=True,
min_repeats=5,
repeats=100, # maximum
target_relative_margin_pct=2.0,
max_time_s=10.0,
noise_threshold_pct=10.0,
)
Adaptive runs stop after the minimum sample count once the 95% confidence interval reaches the requested relative margin, or at the time/repeat limit. Results include confidence bounds, variance, standard error, coefficient of variation, outlier count, and a noisy/stable quality flag.
Parameters, scopes, and measurement boundaries
Parameter sets create independently named benchmarks. Put parametrize below
benchmark, as shown here:
case = Case(name="encoding")
@case.benchmark()
@case.parametrize("size", [100, 10_000], ids=["small", "large"])
@case.parametrize("sort_keys", [False, True], ids=["unsorted", "sorted"])
def encode(size: int, sort_keys: bool) -> str:
import json
return json.dumps(list(range(size)), sort_keys=sort_keys)
Multiple parameter declarations form a Cartesian product. Parameters are injected by name and recorded in JSON, CSV, Markdown, listing, and history data.
Systems support iteration, benchmark, and session scopes. The legacy names
function and module remain aliases for benchmark and session:
@system(scope="iteration")
def temporary_resource():
resource = create_resource()
yield resource
resource.close()
Use setup_timing="exclude" or teardown_timing="exclude" on a Case or
benchmark override to keep fixture work outside the measured interval. Warmup,
iterations, repeats, timeout, isolation, timing boundaries, and profiling can be
overridden on individual @case.benchmark(...) declarations.
Isolation and profiling
Run each selected benchmark in its own worker process:
uv run benchbro run benchmarks --isolation process
Or configure isolation="process" on a case. Python benchmark functions must be
importable when the operating system does not support fork. timeout_s is a
hard limit for subprocess benchmarks and isolated workers; in-process synchronous
benchmarks are checked immediately after each invocation. Linux runners can be
pinned with --cpu-affinity 2,3 or cpu_affinity=(2, 3).
Built-in profiler hooks support cprofile and tracemalloc allocation snapshots:
uv run benchbro run benchmarks --profile cprofile \
--profile-output '.benchbro/profiles/{case}-{benchmark}.prof'
Custom integrations—including wrappers around external profilers—can implement
ProfilerHook and register a factory with register_profiler(...).
When no target is provided, benchbro discovers benchmarks from:
benchmarks/**/*.py(relative to repo root)
You can configure discovery in pyproject.toml or a standalone benchbro.toml:
[tool.benchbro.ini_options]
benchmark_paths = ["benchmarks"]
file_pattern = ["bench_*.py", "*_bench.py", "*benchmark.py", "*benchmarks.py"]
benchmark_paths: directories to scan when no CLI target is providedfile_pattern: glob pattern(s) for benchmark file names in directory discovery
A standalone configuration can also hold common run and output options:
benchmark_paths = ["benchmarks"]
baseline = "feature-branch"
save_history = true
[run]
repeats = 100
warmup = 5
min_iterations = 50
adaptive = true
min_repeats = 5
min_time_s = 0.25
max_time_s = 10.0
target_relative_margin_pct = 2.0
noise_threshold_pct = 10.0
stabilization_delay_s = 0.5
isolation = "in_process"
# cpu_affinity = [2, 3] # Linux only
[output]
json = "artifacts/current.json"
markdown = "artifacts/current.md"
Command-line values take precedence over configuration. Bootstrap a project with
benchbro init; it creates benchbro.toml and a parameterized starter benchmark.
benchbro compares against the baseline by default (.benchbro/baseline.local.json).
If the baseline is missing, benchbro creates it automatically.
If new cases/benchmarks are introduced later, missing entries are merged into baseline.
Pass --new-baseline to replace the entire baseline with the current run.
Pass --ci to use .benchbro/baseline.ci.json for baseline read/write/compare.
Pass --no-compare to skip comparison while still backfilling missing benchmark entries in baseline.
CI mode is deliberately strict: it fails when baseline.ci.json is missing or
does not contain every selected benchmark. Create or replace that file explicitly
with --ci --new-baseline, then commit it. Runs also reject comparisons across a
different Python major/minor version, implementation, operating system, or machine
architecture. Use --allow-environment-mismatch only when that difference is
intentional.
Named baselines keep environments or branches separate:
uv run benchbro baseline update benchmarks --baseline macos-arm64
uv run benchbro run benchmarks --baseline macos-arm64
uv run benchbro baseline list
Names other than local and ci live under .benchbro/baselines/. Result files
carry schema_version = 2; the reader migrates version-1 artifacts and rejects
unknown future schemas rather than silently misreading them.
By default, regular runs do not write artifacts.
Use explicit output flags (--output-json, --output-csv, --output-md) when needed.
The baseline is always written to:
.benchbro/baseline.local.json(default local mode).benchbro/baseline.ci.jsonwhen using--ci
Recommended:
- ignore
.benchbro/for machine-local benchmarking artifacts. - commit
.benchbro/baseline.ci.jsonfor CI comparisons.
Recommended .gitignore
# Benchbro local artifacts
.benchbro/*
!.benchbro/baseline.ci.json
If requested, markdown output can also be written with --output-md.
JSON artifacts include environment metadata for reproducibility (Python/runtime/platform/CPU fields) both at run level and on each benchmark entry.
CLI basics
The command-oriented interface is:
benchbro run [target]
benchbro list [target] [--verbose]
benchbro compare BASELINE.json CURRENT.json
benchbro baseline update [target] --baseline NAME
benchbro baseline list
benchbro history list|show|compare
benchbro init [path]
benchbro completion bash|zsh|fish
The original benchbro TARGET [options] syntax remains supported.
Run selected cases/tags and write outputs:
uv run benchbro my_benchmarks.py \
--case hashing \
--tag fast \
--output-json artifacts/current.json \
--output-csv artifacts/current.csv \
--output-md artifacts/current.md
Compare against baseline:
uv run benchbro my_benchmarks.py
Render time benchmark histograms in terminal output:
uv run benchbro my_benchmarks.py --histogram
Skip comparison for a run while still maintaining baseline structure:
uv run benchbro my_benchmarks.py --no-compare
Regression status uses each benchmark's effective thresholds (benchmark override -> case threshold -> defaults):
- warning default:
50% - error threshold default:
100%
The comparison table shows warning and threshold values for each row.
Histograms are terminal-only in v1 and are shown for time benchmarks.
Exit codes distinguish outcomes:
0: successful or non-regressing run1: invalid input, discovery, configuration, or comparison setup2: statistically supported regression3: benchmark execution failure
Threshold crossings with insufficient evidence are reported as LIKELY or
INCONCLUSIVE without returning the regression exit code. A single-sample run
cannot claim statistical confidence.
Local history
Enable save_history = true or pass --save-history to retain schema-versioned
runs beneath .benchbro/history/:
uv run benchbro history list
uv run benchbro history show COMMIT_OR_FILENAME_FRAGMENT
uv run benchbro history compare OLDER NEWER
History selectors accept an exact path or an unambiguous filename fragment.
Pytest integration
Install the optional integration with uv add --dev 'benchbro[pytest]', then
load the plugin from a pytest configuration:
[tool.pytest.ini_options]
addopts = "-p benchbro.pytest_plugin"
Then reuse pytest fixtures in a measured callable:
def test_parser_speed(benchbro_runner, parsed_fixture):
result = benchbro_runner(
parse,
parsed_fixture,
repeats=20,
warmup=5,
min_iterations=50,
)
assert result.metrics["median_s"] < 0.01
The integration is opt-in, so BenchBro does not add pytest as a runtime dependency or interfere with normal test collection.
End-to-end example
For a complete runnable workflow (baseline + candidate comparison), use:
examples/README.mdmake examples
Development
Run the same checks used by CI:
uv sync --all-groups
make ci
make docs
uv run tox
See CONTRIBUTING.md for the contributor workflow and CHANGELOG.md for release
history. Maintainers should also complete the account-level controls in
docs/maintenance.md.
Metadata
Release files for benchbro 1.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| benchbro-1.0.0.tar.gz | 4.7 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| benchbro-1.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 4.8 MB
Release files / benchbro-1.0.0.tar.gz
| Download URL | benchbro-1.0.0.tar.gz |
|---|---|
| Size | 4.7 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
11f478eba28615ca91e1fc6ca88f403b956b72e2883301f15bbe59ebe0230645
|
|
BLAKE2b-256 checksum How to use checksums |
9c09f5ab1ab67cc06914798e6b9ee0c47fc53dccf5ead649e998fed089cd6764
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / benchbro-1.0.0-py3-none-any.whl
| Download URL | benchbro-1.0.0-py3-none-any.whl |
|---|---|
| Size | 53.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
87771b3999f16c4816b0d1baa160266419579afbe3fa3aaf15efd954b49055e1
|
|
BLAKE2b-256 checksum How to use checksums |
d123ce5ff768972ff2b54114f8c1c5e9d8239058a47847fa071487eebbaf6340
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|