CrowDB TPC Loader
Generate TPC-H or TPC-DS Parquet, upload it to CROWDB Iceberg, and register complete tables through the REST Catalog. The loader creates 8 TPC-H or 24 TPC-DS tables. It does not run benchmark SQL queries or claim TPC certification.
- Repository: buzzcrow/crowdb-tpc-loader
- End-to-end guide: CROWDB TPC loader documentation
Install
Python 3.10–3.12 is required. Install the published package from PyPI:
python3 -m venv .venv
. .venv/bin/activate
python -m pip install crowdb-tpc-loader
For local development, install from a checkout with python -m pip install --only-binary=:all: -e .. The first TPC-H run may download tpchgen-cli 3.0.0; TPC-DS may download DuckDB's tpcds extension. Use --no-download and provide these components ahead of time for an offline run.
Load
Start crowdb/crowdb-iceberg:latest and export the ICEBERG_URI and ICEBERG_TOKEN values printed by docker exec <container> crowdb-monitor credentials show --format env. Keep the token private. Use a new namespace for each run:
crowdb-tpc-loader load --benchmark tpch --sf 1 \
--namespace tpch_demo --report-file ./tpch-demo.json
crowdb-tpc-loader load --benchmark tpcds --sf 1 \
--namespace tpcds_demo --upload-workers 4 --report-file ./tpcds-demo.json
The loader validates the entire dataset before creating a table. TPC-H and TPC-DS use the same load flow. Different tables calculate MD5 and upload concurrently, with 8 workers by default. --upload-workers N selects 1–24 scheduled table writes; the S3 path caps active workers and connections at 8. One S3 client shares its connection pool, with each request signed using that table's catalog credentials. Files stream from disk without a whole-file buffer. S3 PUT and multipart parts send Content-MD5; file payload SHA256 is disabled while SigV4 authentication remains enabled. PUT uses If-None-Match: * without an existence HEAD.
Each table's files are registered in one snapshot. Required recovery records are durable before table creation, each upload and commit; concurrently ready records share one report save. Catalog configuration and namespace setup happen once per run. A successful commit reuses its returned table metadata, then verifies the snapshot's file inventory. Ambiguous commits are checked without repeating add_files. One table's failure does not roll back successful tables. See recovery before retrying. An existing table stops the default load; --on-exists skip leaves it unchanged without verifying it.
After all tables are committed and verified, local staging is removed unless --keep-files is set. Failed and uncertain runs retain local files. Reports contain md5, md5_duration_seconds, transfer_duration_seconds, and the combined upload_duration_seconds. For non-S3 FileIO the checksum is calculated inline while copying, and separate checksum time is unavailable. Performance checks use SF=1; record generation, transfer and commit time separately.
For local Parquet only, use crowdb-tpc-loader generate --benchmark tpch --sf 1 --output-dir ./tpch-sf1.
Normal load does not download uploaded files to recheck MD5, decode every row, or execute SQL. The server validates Content-MD5 during upload. PyIceberg reads Parquet footers for registration and small manifest metadata for commit verification.
Check the result
Run the read-only verifier from a checkout after loading:
python scripts/verify_crowdb.py ./tpch-demo.json --require-complete --iceberg-scan
With DuckDB's iceberg and httpfs extensions, attach the REST Catalog using its token and query tpch_demo.region or run TPC-H Q1 against tpch_demo.lineitem. The website guide has the SQL. At SF 0.01, DuckDB 1.5.6 ran all 22 TPC-H and 99 TPC-DS queries against the published latest container; results matched DuckDB reading the same local Parquet. This is a development check, not a benchmark result. See compatibility and the test record.
Develop and publish
python -m pip install --only-binary=:all: -e '.[dev,sql-test]'
ruff check src tests scripts
ruff format --check src tests scripts
python -m pytest
python -m build
python -m twine check dist/*
CI runs lint, format, tests, and package checks. Publishing is manual through the PyPI workflow from a release/v<version> branch after configuring PyPI Trusted Publishing. See testing for optional real generator and CROWDB runs. The package is Apache-2.0; third-party generators and libraries keep their own licenses.
To publish a later version:
- Keep the existing PyPI Trusted Publisher for project
crowdb-tpc-loader: GitHub ownerbuzzcrow, repositorycrowdb-tpc-loader, workflowpublish.yml, environmentpypi. Keep the matchingpypienvironment in GitHub. - Bump the version in
pyproject.tomlandsrc/crowdb_tpc_loader/__init__.py, then run CI. Create and pushrelease/v<version>from the commit to publish. - In GitHub Actions, open Publish to PyPI, click Run workflow, select that release branch, then run it. The workflow verifies the branch name against the package version, reruns checks, builds distributions, and publishes through OIDC. No PyPI API token is stored in GitHub.
- Confirm the files on PyPI, then test
python -m pip install --no-cache-dir crowdb-tpc-loader==<version>in a clean environment and runcrowdb-tpc-loader --version.
Metadata
Release files for crowdb-tpc-loader 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| crowdb_tpc_loader-0.1.1.tar.gz | 84.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| crowdb_tpc_loader-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 143.5 kB
Release files / crowdb_tpc_loader-0.1.1.tar.gz
| Download URL | crowdb_tpc_loader-0.1.1.tar.gz |
|---|---|
| Size | 84.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8e0b72f15c7184a163d3c3ee363be0ba58c94352e69e1db85b8241c8d22309f0
|
|
BLAKE2b-256 checksum How to use checksums |
231a0c76e57d27b6060d89dce2ebe24cca88fd086bbea6074dfa59cf86308571
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency logRelease files / crowdb_tpc_loader-0.1.1-py3-none-any.whl
| Download URL | crowdb_tpc_loader-0.1.1-py3-none-any.whl |
|---|---|
| Size | 59.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7ad434d5ddb8565e2de26fecbeb2b3ac5acfdd049eaf448abcc6fdf74ed99347
|
|
BLAKE2b-256 checksum How to use checksums |
729620f90500610452c9ca6a352c8ddfdf0cc2e852b9aa22c8874d5f6633c7a7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency log