Skip to main content

CrowDB TPC Loader

Generate TPC-H or TPC-DS Parquet, upload it to CROWDB Iceberg, and register complete tables through the REST Catalog. The loader creates 8 TPC-H or 24 TPC-DS tables. It does not run benchmark SQL queries or claim TPC certification.

Install

Python 3.10–3.12 is required. Install the published package from PyPI:

python3 -m venv .venv
. .venv/bin/activate
python -m pip install crowdb-tpc-loader

For local development, install from a checkout with python -m pip install --only-binary=:all: -e .. The first TPC-H run may download tpchgen-cli 3.0.0; TPC-DS may download DuckDB's tpcds extension. Use --no-download and provide these components ahead of time for an offline run.

Load

Start crowdb/crowdb-iceberg:latest and export the ICEBERG_URI and ICEBERG_TOKEN values printed by docker exec <container> crowdb-monitor credentials show --format env. Keep the token private. Use a new namespace for each run:

crowdb-tpc-loader load --benchmark tpch --sf 1 \
  --namespace tpch_demo --report-file ./tpch-demo.json
crowdb-tpc-loader load --benchmark tpcds --sf 1 \
  --namespace tpcds_demo --upload-workers 4 --report-file ./tpcds-demo.json

The loader validates the entire dataset before creating a table. TPC-H and TPC-DS use the same load flow. Different tables calculate MD5 and upload concurrently, with 8 workers by default. --upload-workers N selects 1–24 scheduled table writes; the S3 path caps active workers and connections at 8. One S3 client shares its connection pool, with each request signed using that table's catalog credentials. Files stream from disk without a whole-file buffer. S3 PUT and multipart parts send Content-MD5; file payload SHA256 is disabled while SigV4 authentication remains enabled. PUT uses If-None-Match: * without an existence HEAD.

Each table's files are registered in one snapshot. Required recovery records are durable before table creation, each upload and commit; concurrently ready records share one report save. Catalog configuration and namespace setup happen once per run. A successful commit reuses its returned table metadata, then verifies the snapshot's file inventory. Ambiguous commits are checked without repeating add_files. One table's failure does not roll back successful tables. See recovery before retrying. An existing table stops the default load; --on-exists skip leaves it unchanged without verifying it.

After all tables are committed and verified, local staging is removed unless --keep-files is set. Failed and uncertain runs retain local files. Reports contain md5, md5_duration_seconds, transfer_duration_seconds, and the combined upload_duration_seconds. For non-S3 FileIO the checksum is calculated inline while copying, and separate checksum time is unavailable. Performance checks use SF=1; record generation, transfer and commit time separately.

For local Parquet only, use crowdb-tpc-loader generate --benchmark tpch --sf 1 --output-dir ./tpch-sf1.

Normal load does not download uploaded files to recheck MD5, decode every row, or execute SQL. The server validates Content-MD5 during upload. PyIceberg reads Parquet footers for registration and small manifest metadata for commit verification.

Check the result

Run the read-only verifier from a checkout after loading:

python scripts/verify_crowdb.py ./tpch-demo.json --require-complete --iceberg-scan

With DuckDB's iceberg and httpfs extensions, attach the REST Catalog using its token and query tpch_demo.region or run TPC-H Q1 against tpch_demo.lineitem. The website guide has the SQL. At SF 0.01, DuckDB 1.5.6 ran all 22 TPC-H and 99 TPC-DS queries against the published latest container; results matched DuckDB reading the same local Parquet. This is a development check, not a benchmark result. See compatibility and the test record.

Develop and publish

python -m pip install --only-binary=:all: -e '.[dev,sql-test]'
ruff check src tests scripts
ruff format --check src tests scripts
python -m pytest
python -m build
python -m twine check dist/*

CI runs lint, format, tests, and package checks. Publishing is manual through the PyPI workflow from a release/v<version> branch after configuring PyPI Trusted Publishing. See testing for optional real generator and CROWDB runs. The package is Apache-2.0; third-party generators and libraries keep their own licenses.

To publish a later version:

  1. Keep the existing PyPI Trusted Publisher for project crowdb-tpc-loader: GitHub owner buzzcrow, repository crowdb-tpc-loader, workflow publish.yml, environment pypi. Keep the matching pypi environment in GitHub.
  2. Bump the version in pyproject.toml and src/crowdb_tpc_loader/__init__.py, then run CI. Create and push release/v<version> from the commit to publish.
  3. In GitHub Actions, open Publish to PyPI, click Run workflow, select that release branch, then run it. The workflow verifies the branch name against the package version, reruns checks, builds distributions, and publishes through OIDC. No PyPI API token is stored in GitHub.
  4. Confirm the files on PyPI, then test python -m pip install --no-cache-dir crowdb-tpc-loader==<version> in a clean environment and run crowdb-tpc-loader --version.

Metadata

Release files for crowdb-tpc-loader 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for crowdb-tpc-loader 0.1.1
File Size Uploaded
crowdb_tpc_loader-0.1.1.tar.gz 84.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for crowdb-tpc-loader 0.1.1
File Interpreter ABI Platform
crowdb_tpc_loader-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 143.5 kB

Release files / crowdb_tpc_loader-0.1.1.tar.gz

Download URL crowdb_tpc_loader-0.1.1.tar.gz
Size 84.5 kB
Tags Source
SHA-256 checksum
How to use checksums
8e0b72f15c7184a163d3c3ee363be0ba58c94352e69e1db85b8241c8d22309f0
BLAKE2b-256 checksum
How to use checksums
231a0c76e57d27b6060d89dce2ebe24cca88fd086bbea6074dfa59cf86308571
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release files / crowdb_tpc_loader-0.1.1-py3-none-any.whl

Download URL crowdb_tpc_loader-0.1.1-py3-none-any.whl
Size 59.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7ad434d5ddb8565e2de26fecbeb2b3ac5acfdd049eaf448abcc6fdf74ed99347
BLAKE2b-256 checksum
How to use checksums
729620f90500610452c9ca6a352c8ddfdf0cc2e852b9aa22c8874d5f6633c7a7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page