Ephemeral, hash-pinned intelligence-benchmark runner for CrowdBench.
Project description
crowdbench-run
The ephemeral, hash-pinned, open-source intelligence-benchmark runner. Pure Python
(httpx + psutil + platformdirs + stdlib; never torch/transformers — HTTP only),
uv-managed, invoked via uvx identically on macOS/Linux/Windows.
It probes hardware/engine, normalizes config against the vendored packages/shared JSON
artifacts, administers the benchmark item set against a locally-served OpenAI-compatible
endpoint, captures per-item timing forensics, and uploads model outputs — scoring is
entirely server-side and answer keys never touch this package. Resumable state lives in
~/.crowdbench/ (via platformdirs).
Dark-period invocation
No real version is published to PyPI until the repo flip; run from source.
Bare crowdbench-run on a terminal launches the guided wizard — a zero-file, zero-flag
contribution flow (splash → one-time consent → detect engines → pick a model → pick suites by
category → time budget + a live throughput estimate → reasoning → confirm → run + poll to a final
state). It runs only what this build can administer (core-v1 and code-v1 directly); industry lm-eval suites are
shown with their cost and the exact crowdbench-run bridge … command, never run under the wizard's
pretense. On a non-TTY (piped/CI) the bare command prints usage and exits — it never hangs for input.
uvx --from ./packages/runner crowdbench-run # the guided wizard (TTY)
uvx --from ./packages/runner crowdbench-run detect
uvx --from ./packages/runner crowdbench-run inventory
# The flag path (what the wizard drives; agents use it directly). --dataset is now OPTIONAL —
# omit it and the public prompts dataset is auto-fetched from the API into the content-addressed cache.
uvx --from ./packages/runner crowdbench-run core-v1-rc \
--endpoint http://localhost:8080/v1 --model <model> --agree-tos --yes
The RUN flow is probe → confirm → run → upload: probe hardware + engine, confirm the plan
(interactive, or --yes/--dry-run for agents), administer every item capturing output +
per-item timing (capture requirement 12), then POST to /v1/runs/outputs, which returns a
pre-issued run URL (status pending-score until the server scores). An interactive run shows
the same ASCII splash as the wizard (TTY only — never under --json, --yes, or a pipe), and the
benchmark id is validated before the consent screen, so an unsupported id bounces immediately
rather than after you have agreed.
An interactive run adds three things (all TTY-only; --yes/--json surfaces are unchanged):
- Model picker. With several loaded models and no
--model, the wizard's numbered picker runs — never a silent first-pick. Non-interactive stays documented behaviour:--model, else the endpoint's first loaded model. - Pre-flight, before the proceed prompt. How many items against which model; an estimated duration on this machine from a short (<30 s) throughput probe — reasoning-aware: the estimate multiplies by the suite's reasoning multiplier when reasoning is configured on OR the probe observed native reasoning output (thinking models emit thousands of reasoning tokens per item); what uploads (outputs + timings), that you get a run URL, and that scoring is server-side (for code-v1: in a sandboxed executor, a few extra minutes after upload).
- Honest progress + safe abort. Progress lines read
problem 12/100 · ~1h 40m left(rolling ETA from measured per-item times); raw item ids appear only with--verbose(they always ride the--jsonpayload).Ctrl-C anytime — partial runs are never uploaded: SIGINT aborts cleanly (exit 130) and uploads nothing — a partial suite never scores, so it is never submitted, and a started suite is never time-truncated.
code-v1 runs the identical path: the runner administers the pooled code prompts and uploads
OUTPUTS + timings, and the held-out tests never leave the server (they are the answer key — see
executor/README.md). Because code-v1 is scored asynchronously (queue → ephemeral container →
judge) rather than inline, the runner polls it to a final state on a longer per-benchmark
window (15 min vs. core-v1's 5); if the window closes first it says so and tells you to report the
run — it never assumes success. code-v1 uses the same sampler defaults as core-v1 (temperature 0,
fixed seed): no in-repo code-v1 spec prescribes a different profile, and the runner does not invent
one. Consent is shown once
(state in ~/.crowdbench); --agree-tos attests a human has seen and agreed to the terms (what
uploads; the public CC-BY-NC-SA 4.0 dataset; immutability) and is required to upload non-interactively.
Commands
| command | what it does |
|---|---|
| (no arguments, on a TTY) | the guided wizard — zero-file, zero-flag contribution flow |
<benchmark-id> |
run one benchmark end-to-end (the flag path); core-v1 / core-v1-rc / code-v1 / throughput-std-v1 in this phase |
bridge lmeval:<task>:<variant> |
run/parse an allowlisted lm-eval suite → validate → upload |
detect |
scan common local serving ports; report apps + loaded models (by route signature) |
inventory |
list local models (Ollama / LM Studio / HF cache / user dirs) + benchmark tooling |
doctor |
probe + connectivity + state report for support |
cache list|clear |
manage the content-addressed dataset cache |
Key flags (dual-surface parity — every wizard prompt has a flag): --endpoint, --api-base,
--dataset (omit to auto-fetch), --model, --model-path, --reps N, --seed,
--reasoning-effort, --reasoning-budget, --agree-tos, --minutes,
--hf-repo/--hf-revision/--unattributed, --endpoint-api-key
(env CROWDBENCH_ENDPOINT_API_KEY; never logged/uploaded/stored), --dry-run,
--no-upload/--save/--upload-only, --json, --yes, --verbose (raw item ids on the
progress lines), --sudo-probes, --ports. Auth uses CROWDBENCH_API_KEY (else an anonymous
submitter is minted and cached).
Trust posture
- Outputs only. The runner uploads model outputs + telemetry; it never sees, computes, or transmits answer keys. Benchmark item content is inert data — never a tool, never agentic.
- Dataset integrity. A dataset's sha256 must match the benchmark's
runner_specpin before every run (capture requirement 11); a mismatch refuses to run (a tampered set never runs). - Vendored artifacts, byte-matched.
crowdbench_run/_artifacts/*.jsonare copied verbatim frompackages/shared/artifactsso both languages read one source; CI fails on drift (scripts/check_artifact_drift.py— the third leg of the cross-language drift check). - Endpoint key privacy.
--endpoint-api-keyis forwarded as the target endpoint's Authorization header only, never logged/uploaded/stored.
Supply-chain audit trail: packages/runner/BUILD-TIME-CHECKS.md.
Publishing (PyPI Trusted Publishing)
Releases go out through .github/workflows/publish.yml using PyPI Trusted Publishing
(OIDC) — there is no long-lived PyPI API token in this repo, ever. GitHub mints a
short-lived OIDC token per run and pypa/gh-action-pypi-publish (pinned by full commit SHA)
exchanges it with PyPI. Three human gates stand between a dispatch and a live release:
workflow_dispatch only (no push/tag/schedule), a typed confirm input that must equal
publish-to-pypi, and environment: pypi (owner approval before any publish step runs). A
preceding job builds the sdist+wheel, runs twine check --strict, install-smokes both
artifacts, and refuses the placeholder 0.0.x version line.
Owner one-time setup — verify on pypi.org. Create the Trusted Publisher for this project under Manage project → Publishing (or Your projects → Publishing for a pending publisher). These three fields must match the workflow exactly, or PyPI rejects the OIDC exchange:
| PyPI Trusted-Publisher field | Must be |
|---|---|
| PyPI project (package) name | crowdbench-run |
| GitHub repository + workflow filename | crowdbench-ai/crowdbench-dev, workflow publish.yml |
| Environment name | pypi |
(The owner also configures the pypi GitHub Environment's protection rule — required
reviewers — so gate #3 actually pauses for approval.) This PR only prepares the workflow and
docs; wiring the Trusted Publisher and the environment protection rule is an owner action on
pypi.org / GitHub, and nothing publishes during the dark period.
Development
cd packages/runner
python -m pip install -e ".[dev]"
python -m pytest -q
The test suite runs on ubuntu/macos/windows in CI (a release gate — cross-platform posture).
CrowdBench is a service of Hooch Labs LLC. © 2026 Hooch Labs LLC · Apache-2.0.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file crowdbench_run-0.1.3.tar.gz.
File metadata
- Download URL: crowdbench_run-0.1.3.tar.gz
- Upload date:
- Size: 148.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
fb04f7cc2d2518e296972a66c5622618a68fd563e0746d39a8d5146d710b3881
|
|
| MD5 |
136b6ffc554c7d167abea177c2eabc8f
|
|
| BLAKE2b-256 |
6ddc1413ae7c29d31e54ae89786b3d904f6f96d049991bf7c0e175f51b6e5e05
|
Provenance
The following attestation bundles were made for crowdbench_run-0.1.3.tar.gz:
Publisher:
publish.yml on crowdbench-ai/crowdbench-dev
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
crowdbench_run-0.1.3.tar.gz -
Subject digest:
fb04f7cc2d2518e296972a66c5622618a68fd563e0746d39a8d5146d710b3881 - Sigstore transparency entry: 2195526384
- Sigstore integration time:
-
Permalink:
crowdbench-ai/crowdbench-dev@f61e7a82d7b25469ab47cde26d4f9d7d15938c69 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/crowdbench-ai
-
Access:
private
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@f61e7a82d7b25469ab47cde26d4f9d7d15938c69 -
Trigger Event:
workflow_dispatch
-
Statement type:
File details
Details for the file crowdbench_run-0.1.3-py3-none-any.whl.
File metadata
- Download URL: crowdbench_run-0.1.3-py3-none-any.whl
- Upload date:
- Size: 92.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7126839f427c1d9f60bc81da706d3ef9e0cb131f8006d455bdf6985fd37f509c
|
|
| MD5 |
dab0f73d95244e37948f968dfd8fa871
|
|
| BLAKE2b-256 |
9bfaedb7d3aa5df026814a6754ca148f5d53ca88402caedbb4e5a2f851b50e6a
|
Provenance
The following attestation bundles were made for crowdbench_run-0.1.3-py3-none-any.whl:
Publisher:
publish.yml on crowdbench-ai/crowdbench-dev
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
crowdbench_run-0.1.3-py3-none-any.whl -
Subject digest:
7126839f427c1d9f60bc81da706d3ef9e0cb131f8006d455bdf6985fd37f509c - Sigstore transparency entry: 2195526396
- Sigstore integration time:
-
Permalink:
crowdbench-ai/crowdbench-dev@f61e7a82d7b25469ab47cde26d4f9d7d15938c69 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/crowdbench-ai
-
Access:
private
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@f61e7a82d7b25469ab47cde26d4f9d7d15938c69 -
Trigger Event:
workflow_dispatch
-
Statement type: