Skip to main content
█                   █    █
█▀▀▄ █▀▀█ █▀▀▄ █▀▀▀ █▀▀▄ █▀▀▄ ▄▀▀▄ ▀▄▄▀
█▄▄▀ █▄▄▄ █  █ █▄▄▄ █  █ █▄▄▀ ▀▄▄▀ ▄▀▀▄

BenchBox

License: MIT Beta Software codecov PyPI Release PyPI Downloads

BenchBox is an open-source Python toolkit for benchmarking analytical data platforms. It generates and loads data, runs repeatable workloads, validates execution, and records comparable results through one workflow.

Use BenchBox to evaluate local databases, cloud data warehouses, and DataFrame runtimes with the same benchmark definitions. It is built for data engineers, platform evaluators, and performance practitioners who need evidence they can inspect and reproduce.

BenchBox focuses on online analytical processing (OLAP). If you need an online transaction processing (OLTP) benchmark, see the comparison of database benchmarking tools.

Why BenchBox?

  • Run recognized workloads. Use TPC-H, TPC-DS, TPC-DI, ClickBench, SSB, Join Order Benchmark, and BenchBox's focused primitive workloads.
  • Compare different kinds of engines. Run the same benchmark against SQL databases, cloud warehouses, and native DataFrame APIs.
  • Control the full run. Generate data, create schemas, load tables, execute queries, validate results, and capture metrics from one command.
  • Reproduce results. Record the benchmark, scale, platform, configuration, tuning, environment, timings, and validation evidence in structured result bundles.
  • Inspect before you spend. Preview planned queries, files, phases, and configuration with dry-run support.
  • Analyze and share evidence. Compare runs, render terminal charts, export reports, and contribute results to the public Results Explorer.

Quick Start

The quickest local path uses DuckDB and a small TPC-H dataset.

1. Install BenchBox with DuckDB

Using uv:

uv add benchbox --extra duckdb

Using pip:

python -m pip install "benchbox[duckdb]"

DuckDB is an optional dependency. A plain benchbox installation includes SQLite but does not include DuckDB.

The commands below use the benchbox executable installed by either method. If you used uv add and have not activated the project environment, prefix each command with uv run --.

2. Run a benchmark

benchbox run \
  --platform duckdb \
  --benchmark tpch \
  --scale 0.01

This command generates about 10 MB of TPC-H data, loads it into DuckDB, runs the benchmark, validates the execution, and stores the result under benchmark_runs/.

3. Inspect the result

benchbox results

The summary lists each recent run's benchmark, platform, timestamp, duration, query count, and BenchBox version.

To preview a run without executing it:

benchbox run \
  --dry-run ./preview \
  --platform duckdb \
  --benchmark tpch \
  --scale 0.01

Continue with the five-minute guide, or see the installation guide for other platforms and package extras.

What can you benchmark?

Benchmarks

BenchBox includes several kinds of analytical workloads:

  • TPC standards: TPC-H, TPC-DS, and TPC-DI
  • Academic benchmarks: SSB, AMPLab, and Join Order Benchmark
  • Industry and real-world workloads: ClickBench, H2O DB Benchmark, NYC Taxi, Flight Data, TSBS DevOps, and CoffeeShop
  • Focused primitives: read, write, transaction, metadata, and AI operations
  • AI and machine learning: Vector Search
  • Experimental variants: TPC-DS One Big Table, TPC-Havoc, TPC-H Skew, and TPC-H Data Vault

See the benchmark catalog for workload details, resource guidance, and selection help.

Platforms

BenchBox uses adapters to run workloads across three broad groups:

  • local and embedded SQL engines, such as DuckDB, SQLite, and DataFusion;
  • cloud warehouses and distributed SQL engines, such as Snowflake, BigQuery, Databricks, Redshift, ClickHouse, Spark, Trino, and PrestoDB; and
  • native DataFrame runtimes, such as Polars, Pandas, PySpark, DataFusion, Dask, and cuDF.

Support status and published evidence answer different questions. A supported adapter can be available before the public results corpus contains a run for that platform. Check the platform guides, platform comparison matrix, and public support contract before planning a comparison.

  • Platform registry: 50 metadata entries; 46 SQL-capable; 18 DataFrame-capable; 14 dual-mode; support status counts: stable=5, beta=28, experimental=16, deprecated=1.
  • Benchmark registry: 23 metadata entries; 22 public discovery entries.

These counts come from BenchBox's registries and are checked in the test suite. Use benchbox platforms list and benchbox benchmarks list for the current names and support details.

Results you can inspect and compare

Each run records more than a headline time. Result bundles can include query timings, validation status, resource measurements, platform configuration, tuning evidence, environment metadata, cost data, and captured query plans. The available fields depend on the platform and run configuration.

Use the CLI to inspect, visualize, and export local results:

uv run -- benchbox results
uv run -- benchbox visualize benchmark_runs/results/*.json
uv run -- benchbox export --last --format html

The public BenchBox Results Explorer lets you inspect published runs and their provenance. Comparisons are meaningful only when the workload, scale, phase, configuration, tuning, hardware, and evidence are compatible; a shared benchmark name alone is not enough.

To contribute a complete run, follow the result contribution guide. It explains local validation, privacy safeguards, trust labels, and the published-results submission process.

Learn more

Goal Start here
Install BenchBox Installation and environment setup
Run your first benchmark Getting started in five minutes
Learn the CLI CLI quick reference
Choose a benchmark Benchmark catalog
Choose a platform Platform selection guide
Use a DataFrame runtime DataFrame platforms
Use the Python API Python API reference
Find examples Examples guide
Troubleshoot a run Troubleshooting guide
Understand the design Architecture overview
Add a platform Adding new platforms
Create a custom benchmark Custom benchmark guide

The documentation index links to the complete user, reference, design, and contributor documentation.

Installation and platform setup

The quick start intentionally installs only the DuckDB extra. BenchBox offers separate extras for cloud services, database drivers, DataFrame libraries, and development tools so that you install only what you need.

See the installation guide for the supported package managers and extras. Then use the dependency checker for your target platform:

uv run -- benchbox check-deps --platform databricks

Platform guides cover credentials, connection settings, and platform-specific options. Do not put credentials in configuration files that you commit.

Project status

BenchBox is BETA software. The CLI and core workflows are usable, but public APIs may change before 1.0.

Current release: v0.4.1.

Support labels describe the stability of each public surface. The benchbox.experimental namespace has no compatibility guarantee and can change or be removed without notice. Read the public contracts and support taxonomy before depending on an experimental or beta interface.

See PyPI for published releases and DISCLAIMER.md for project limitations. Release and versioning details live in the backward-compatibility policy and release guide.

Contributing

Bug reports, documentation improvements, platform adapters, benchmark work, and result contributions are welcome.

Disclaimer

BenchBox is an independent open-source project. It is not affiliated with the Transaction Processing Performance Council or with Joe Harris's past or present employers. See DISCLAIMER.md for details.

License

BenchBox is available under the MIT License.

Release files for benchbox 0.4.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for benchbox 0.4.1
File Size Uploaded
benchbox-0.4.1.tar.gz 9.7 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for benchbox 0.4.1
File Interpreter ABI Platform
benchbox-0.4.1-py3-none-any.whl Python 3 none any Details

Total release size: 20.2 MB

Release files / benchbox-0.4.1.tar.gz

Download URL benchbox-0.4.1.tar.gz
Size 9.7 MB
Tags Source
SHA-256 checksum
How to use checksums
feea87cf379486231e7679b1507e0a450541b8ff201add86d9148c3d1f2ee923
BLAKE2b-256 checksum
How to use checksums
ad25107cb326ffb55907529d2f347acccf00d0e30e11e0b40a2fe2d16c3f92aa
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / benchbox-0.4.1-py3-none-any.whl

Download URL benchbox-0.4.1-py3-none-any.whl
Size 10.4 MB
Tags Python 3
SHA-256 checksum
How to use checksums
b4e75c6c397033af0719eb1584d3d94a27bc2206878e8dcd9dda78db92603c2f
BLAKE2b-256 checksum
How to use checksums
99ce10ecc2e823dc824a883967752901bc53d03c4b15f850b51eba05b1401e4d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

0.4.1 This release

2 release files

0.4.0

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page