Skip to main content

benchbro

benchbro is a Python benchmarking library and CLI with parameterized discovery, noise-aware statistics, reproducible environments, and rich terminal output.

Read the documentation or jump straight to the quick start.

Bench Bro Mascot

Quick start

Add the CLI to a project:

uv add --dev benchbro

For a cloned development checkout, use uv sync --all-groups instead. A global installation with uv tool install benchbro is also supported; in that case run benchbro directly instead of uv run benchbro in the examples below.

BenchBro bundles a version-matched agent skill. After installing the package, add it to the current project with:

uvx library-skills --skill use-benchbro --yes

Create benchmark cases in any importable module:

from benchbro import Case, system


@system(scope="module")
def salt() -> bytes:
    return b"-fixture"


case = Case(name="hashing", case_type="cpu", metric_type="time", tags=["fast", "core"])


@case.input()
def payload() -> bytes:
    return b"benchbro"


@case.benchmark()
def sha1(payload: bytes, salt: bytes) -> str:
    import hashlib

    return hashlib.sha1(payload + salt).hexdigest()


@case.benchmark()
def sha256(payload: bytes, salt: bytes) -> str:
    import hashlib

    return hashlib.sha256(payload + salt).hexdigest()

Standalone systems behave like lightweight fixtures:

  • declare them with @system() outside any Case
  • inject them into benchmarks and other systems by parameter name
  • choose scope="function" (default) or scope="module"
  • systems may be sync, async, yield-based, or async-yield-based
  • async inputs, systems, benchmarks, and teardown share one managed event loop
  • keep them separate from inputs: inputs cannot depend on systems, and systems cannot depend on inputs

Subprocess benchmarks use a dedicated subprocess API:

import sys

from benchbro import Case, CommandSpec

case = Case(name="cli", metric_type="time")


@case.subprocess()
def python_cli() -> CommandSpec:
    return CommandSpec(argv=[sys.executable, "-c", "pass"])
  • @case.subprocess() benchmarks are time-only in v1.
  • Commands are explicit argv, not shell strings.
  • Unexpected exit codes and timeouts fail the run.
  • Long-lived subprocesses should be started in @system(...) and consumed by benchmarks or subprocess benchmarks.

Regression thresholds default to 50.0 percent warning and 100.0 percent error at the case level, and can be overridden per benchmark:

case = Case(name="hashing", warning_threshold_pct=5.0, regression_threshold_pct=10.0)


@case.benchmark(warning_threshold_pct=2.0, regression_threshold_pct=3.0)
def critical_path(payload: bytes) -> str: ...

Comparison metric can be configured at case-level and benchmark-level:

case = Case(name="hashing", comparison_metric="p95_s")


@case.benchmark(comparison_metric="median_s")
def critical_path(payload: bytes) -> str: ...

Valid comparison metrics:

  • time: median_s (default), mean_s, iqr_s, p95_s, stddev_s, ops_per_sec
  • memory: peak_alloc_bytes (default), net_alloc_bytes, peak_alloc_bytes_max

GC is disabled during measured iterations by default. To keep the interpreter GC behavior unchanged, set:

Case(name="hashing", gc_control="inherit")

Run benchmarks:

uv run benchbro run --repeats 10 --warmup 2

The defaults are 20 repeats, 50 measured iterations per repeat, and 5 warmup iterations. Percentiles are calculated across repeat-level per-iteration means.

For adaptive sampling, give BenchBro a precision target and/or time budget:

case = Case(
    name="hashing",
    adaptive=True,
    min_repeats=5,
    repeats=100,  # maximum
    target_relative_margin_pct=2.0,
    max_time_s=10.0,
    noise_threshold_pct=10.0,
)

Adaptive runs stop after the minimum sample count once the 95% confidence interval reaches the requested relative margin, or at the time/repeat limit. Results include confidence bounds, variance, standard error, coefficient of variation, outlier count, and a noisy/stable quality flag.

Parameters, scopes, and measurement boundaries

Parameter sets create independently named benchmarks. Put parametrize below benchmark, as shown here:

case = Case(name="encoding")


@case.benchmark()
@case.parametrize("size", [100, 10_000], ids=["small", "large"])
@case.parametrize("sort_keys", [False, True], ids=["unsorted", "sorted"])
def encode(size: int, sort_keys: bool) -> str:
    import json

    return json.dumps(list(range(size)), sort_keys=sort_keys)

Multiple parameter declarations form a Cartesian product. Parameters are injected by name and recorded in JSON, CSV, Markdown, listing, and history data.

Systems support iteration, benchmark, and session scopes. The legacy names function and module remain aliases for benchmark and session:

@system(scope="iteration")
def temporary_resource():
    resource = create_resource()
    yield resource
    resource.close()

Use setup_timing="exclude" or teardown_timing="exclude" on a Case or benchmark override to keep fixture work outside the measured interval. Warmup, iterations, repeats, timeout, isolation, timing boundaries, and profiling can be overridden on individual @case.benchmark(...) declarations.

Isolation and profiling

Run each selected benchmark in its own worker process:

uv run benchbro run benchmarks --isolation process

Or configure isolation="process" on a case. Python benchmark functions must be importable when the operating system does not support fork. timeout_s is a hard limit for subprocess benchmarks and isolated workers; in-process synchronous benchmarks are checked immediately after each invocation. Linux runners can be pinned with --cpu-affinity 2,3 or cpu_affinity=(2, 3).

Built-in profiler hooks support cprofile and tracemalloc allocation snapshots:

uv run benchbro run benchmarks --profile cprofile \
  --profile-output '.benchbro/profiles/{case}-{benchmark}.prof'

Custom integrations—including wrappers around external profilers—can implement ProfilerHook and register a factory with register_profiler(...).

When no target is provided, benchbro discovers benchmarks from:

  • benchmarks/**/*.py (relative to repo root)

You can configure discovery in pyproject.toml or a standalone benchbro.toml:

[tool.benchbro.ini_options]
benchmark_paths = ["benchmarks"]
file_pattern = ["bench_*.py", "*_bench.py", "*benchmark.py", "*benchmarks.py"]
  • benchmark_paths: directories to scan when no CLI target is provided
  • file_pattern: glob pattern(s) for benchmark file names in directory discovery

A standalone configuration can also hold common run and output options:

benchmark_paths = ["benchmarks"]
baseline = "feature-branch"
save_history = true

[run]
repeats = 100
warmup = 5
min_iterations = 50
adaptive = true
min_repeats = 5
min_time_s = 0.25
max_time_s = 10.0
target_relative_margin_pct = 2.0
noise_threshold_pct = 10.0
stabilization_delay_s = 0.5
isolation = "in_process"
# cpu_affinity = [2, 3] # Linux only

[output]
json = "artifacts/current.json"
markdown = "artifacts/current.md"

Command-line values take precedence over configuration. Bootstrap a project with benchbro init; it creates benchbro.toml and a parameterized starter benchmark.

benchbro compares against the baseline by default (.benchbro/baseline.local.json). If the baseline is missing, benchbro creates it automatically. If new cases/benchmarks are introduced later, missing entries are merged into baseline. Pass --new-baseline to replace the entire baseline with the current run. Pass --ci to use .benchbro/baseline.ci.json for baseline read/write/compare. Pass --no-compare to skip comparison while still backfilling missing benchmark entries in baseline.

CI mode is deliberately strict: it fails when baseline.ci.json is missing or does not contain every selected benchmark. Create or replace that file explicitly with --ci --new-baseline, then commit it. Runs also reject comparisons across a different Python major/minor version, implementation, operating system, or machine architecture. Use --allow-environment-mismatch only when that difference is intentional.

Named baselines keep environments or branches separate:

uv run benchbro baseline update benchmarks --baseline macos-arm64
uv run benchbro run benchmarks --baseline macos-arm64
uv run benchbro baseline list

Names other than local and ci live under .benchbro/baselines/. Result files carry schema_version = 2; the reader migrates version-1 artifacts and rejects unknown future schemas rather than silently misreading them.

By default, regular runs do not write artifacts. Use explicit output flags (--output-json, --output-csv, --output-md) when needed.

The baseline is always written to:

  • .benchbro/baseline.local.json (default local mode)
  • .benchbro/baseline.ci.json when using --ci

Recommended:

  • ignore .benchbro/ for machine-local benchmarking artifacts.
  • commit .benchbro/baseline.ci.json for CI comparisons.

Recommended .gitignore

# Benchbro local artifacts
.benchbro/*
!.benchbro/baseline.ci.json

If requested, markdown output can also be written with --output-md.

JSON artifacts include environment metadata for reproducibility (Python/runtime/platform/CPU fields) both at run level and on each benchmark entry.

CLI basics

The command-oriented interface is:

benchbro run [target]
benchbro list [target] [--verbose]
benchbro compare BASELINE.json CURRENT.json
benchbro baseline update [target] --baseline NAME
benchbro baseline list
benchbro history list|show|compare
benchbro init [path]
benchbro completion bash|zsh|fish

The original benchbro TARGET [options] syntax remains supported.

Run selected cases/tags and write outputs:

uv run benchbro my_benchmarks.py \
  --case hashing \
  --tag fast \
  --output-json artifacts/current.json \
  --output-csv artifacts/current.csv \
  --output-md artifacts/current.md

Compare against baseline:

uv run benchbro my_benchmarks.py

Render time benchmark histograms in terminal output:

uv run benchbro my_benchmarks.py --histogram

Skip comparison for a run while still maintaining baseline structure:

uv run benchbro my_benchmarks.py --no-compare

Regression status uses each benchmark's effective thresholds (benchmark override -> case threshold -> defaults):

  • warning default: 50%
  • error threshold default: 100%

The comparison table shows warning and threshold values for each row.

Histograms are terminal-only in v1 and are shown for time benchmarks.

Exit codes distinguish outcomes:

  • 0: successful or non-regressing run
  • 1: invalid input, discovery, configuration, or comparison setup
  • 2: statistically supported regression
  • 3: benchmark execution failure

Threshold crossings with insufficient evidence are reported as LIKELY or INCONCLUSIVE without returning the regression exit code. A single-sample run cannot claim statistical confidence.

Local history

Enable save_history = true or pass --save-history to retain schema-versioned runs beneath .benchbro/history/:

uv run benchbro history list
uv run benchbro history show COMMIT_OR_FILENAME_FRAGMENT
uv run benchbro history compare OLDER NEWER

History selectors accept an exact path or an unambiguous filename fragment.

Pytest integration

Install the optional integration with uv add --dev 'benchbro[pytest]', then load the plugin from a pytest configuration:

[tool.pytest.ini_options]
addopts = "-p benchbro.pytest_plugin"

Then reuse pytest fixtures in a measured callable:

def test_parser_speed(benchbro_runner, parsed_fixture):
    result = benchbro_runner(
        parse,
        parsed_fixture,
        repeats=20,
        warmup=5,
        min_iterations=50,
    )
    assert result.metrics["median_s"] < 0.01

The integration is opt-in, so BenchBro does not add pytest as a runtime dependency or interfere with normal test collection.

End-to-end example

For a complete runnable workflow (baseline + candidate comparison), use:

  • examples/README.md
  • make examples

Development

Run the same checks used by CI:

uv sync --all-groups
make ci
make docs
uv run tox

See CONTRIBUTING.md for the contributor workflow and CHANGELOG.md for release history. Maintainers should also complete the account-level controls in docs/maintenance.md.

Metadata

Release files for benchbro 1.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for benchbro 1.0.0
File Size Uploaded
benchbro-1.0.0.tar.gz 4.7 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for benchbro 1.0.0
File Interpreter ABI Platform
benchbro-1.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 4.8 MB

Release files / benchbro-1.0.0.tar.gz

Download URL benchbro-1.0.0.tar.gz
Size 4.7 MB
Tags Source
SHA-256 checksum
How to use checksums
11f478eba28615ca91e1fc6ca88f403b956b72e2883301f15bbe59ebe0230645
BLAKE2b-256 checksum
How to use checksums
9c09f5ab1ab67cc06914798e6b9ee0c47fc53dccf5ead649e998fed089cd6764
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / benchbro-1.0.0-py3-none-any.whl

Download URL benchbro-1.0.0-py3-none-any.whl
Size 53.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
87771b3999f16c4816b0d1baa160266419579afbe3fa3aaf15efd954b49055e1
BLAKE2b-256 checksum
How to use checksums
d123ce5ff768972ff2b54114f8c1c5e9d8239058a47847fa071487eebbaf6340
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

1.0.0 This release

2 release files

0.5.0

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page