Skip to main content

Benchmark runner for local-bench.ai - the community quality leaderboard for local AI setups.

Reason this release was yanked:

Fails its own execution-contract self-check at the agentic phase after the full static suite

Project description

local-bench-ai

CLI benchmark runner for local-bench.ai — a community quality leaderboard for local AI setups. It benchmarks GGUF models served by llama.cpp (or any OpenAI-compatible endpoint), scores them on the five-axis Local Intelligence Index (Agentic, Knowledge, Instruction-Following, execution-verified Coding, Math — with call-formatting and long-context tracked as unweighted diagnostics), and packs signed, reproducible result bundles that publish to the board immediately on submission.

Quickstart

pip install "local-bench-ai[hf]"   # Python 3.11+

# 1. Fetch the complete benchmark suite (hash-verified)
localbench fetch-suite --site https://local-bench.ai \
  --suite suite-v1-full-exec-6axis-v1 --accept-suite-terms

# 2. Optional pre-cache (required with --offline; online advanced bench auto-caches a miss)
localbench cache-tokenizer <hf-model-id>

# 3. Run the full suite; explicitly consent to restricted model-generated code execution
localbench bench <catalog-model-or-hf-repo> \
  --llama-server-path <path-to-llama-server> \
  --allow-untrusted-code

# 4. Advanced managed-harness path
localbench bench \
  --runtime llama.cpp --server-bin <path-to-llama-server> \
  --model-file <model.gguf> --model-id <model-slug> \
  --hf-model-id <hf-model-id> \
  --suite suite-v1-full-exec-6axis-v1 --bench all \
  --wsl-venv-python <managed-wsl-python> \
  --appworld-root <managed-appworld-root> \
  --lane bounded-final-v2 --profile auto --tier standard \
  --allow-untrusted-code \
  --ctx 32768 --seed 1234 --out runs/my-bench

# 5. Submit — complete runs publish to the board immediately, attributed to you
localbench submit run --run runs/my-bench

Full-suite execution requires the AppWorld harness (localbench setup-agentic) and Docker. Agentic runs use the same signed, pinned appliance on both supported host paths: Windows hosts run it through managed WSL2, while Linux hosts materialize it natively and launch it under mandatory bubblewrap isolation. The other axes run wherever llama.cpp and Docker do. Runtime attestations preserve the shared canonical identity fields on both paths. Worker topology evidence is conditional: Windows/WSL emits wsl_distro, wsl_kernel, and appworld_root_under_mnt; native Linux emits runtime_topology, linux_kernel, and linux_os_release and omits those WSL-only fields. --allow-untrusted-code acknowledges the warning that model-generated code executes in a restricted container. Before model download, the CLI actively verifies its non-root, network-disabled, read-only, capability-free, seccomp-filtered, resource-bounded sandbox; missing consent or an unenforceable control fails the coding axis closed. Existing result bundles with pending coding artifacts can be completed with localbench grade-coding --allow-untrusted-code. Safetensors/vLLM execution is a separate maintainer-operated lane documented in docs/benchmark-build/vllm-maintainer-runbook.md; it does not change the public llama.cpp/GGUF path.

Troubleshooting

Windows CLI with a Docker engine inside WSL2

Do not use tcp://localhost:2375: the WSL2 localhost relay can drop Docker attach output even when ordinary daemon requests succeed. Connect through the current WSL adapter IP, pull the pinned image into the same rootful daemon store, and keep the distribution alive for the run. The complete setup, including rootless-vs-rootful stores, safe TCP exposure, transient systemd units, and a standalone version-matched Windows client, is in the Windows + WSL-engine coding sandbox guide.

The site's submit page generates these commands for your exact model and runtime, including the full identity flag set for bring-your-own-server runs. Publishable bounded-final-v2 runs require a 32k server context.

What makes rows trustworthy

  • Suites are hash-pinned releases; sampler settings are pinned (greedy, seeded).
  • Coding is BigCodeBench-Hard, executed locally in a network-disabled, digest-pinned Docker sandbox with no host mounts; coding and agentic verdicts are carried as client-reported evidence and labeled as such on the board.
  • Every number on the board links to a receipt with the full run manifest.
  • Complete runs publish and rank immediately, attributed to the submitter; maintainers moderate post-hoc and can suppress rows that fail scrutiny.

Methodology: https://local-bench.ai/methodology

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

local_bench_ai-0.4.6.tar.gz (1.2 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

local_bench_ai-0.4.6-py3-none-any.whl (963.6 kB view details)

Uploaded Python 3

File details

Details for the file local_bench_ai-0.4.6.tar.gz.

File metadata

  • Download URL: local_bench_ai-0.4.6.tar.gz
  • Upload date:
  • Size: 1.2 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.9.22 {"installer":{"name":"uv","version":"0.9.22","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for local_bench_ai-0.4.6.tar.gz
Algorithm Hash digest
SHA256 0e8f5e34c67254fa0993ed8ad0a8a30799a13226aec9fe6b404b86fbb6764cad
MD5 5fdd1c9de93e9686d4b69d13c34d6266
BLAKE2b-256 cb6101c5e204044cfffadb9ea08e4cba848455ab4ad4734e0ce55c69387aebf2

See more details on using hashes here.

File details

Details for the file local_bench_ai-0.4.6-py3-none-any.whl.

File metadata

  • Download URL: local_bench_ai-0.4.6-py3-none-any.whl
  • Upload date:
  • Size: 963.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.9.22 {"installer":{"name":"uv","version":"0.9.22","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for local_bench_ai-0.4.6-py3-none-any.whl
Algorithm Hash digest
SHA256 629d089133a5977f8413058cab10b58d973f7fa000dce55f35dee150e68e60fc
MD5 399aff3f41c2dc8bd200a9ecd593619a
BLAKE2b-256 2afbf5ad6550fa3204e0dc48199168397d84c9da3645def3bb20ad89df7bdd3d

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page