Skip to main content

szl-engine-bench

An honest, standard-library-only benchmark client for OpenAI-compatible streaming endpoints exposed by vLLM, SGLang, llama.cpp, MLX, TGI, and Transformers Serve.

Version 0.3.0 reports separate distributions for time to first token/chunk (TTFT), inter-chunk latency (ITL), end-to-end latency, and the fraction of runs meeting a declared TTFT service-level objective. It does not collapse an engine comparison into one winner.

The distinction is deliberate. SGLang's serving guide reports TTFT, TPOT, ITL, and throughput as separate online-serving metrics, while vLLM Bench separately exposes TTFT/ITL percentiles and SLO goodput. Cache shape, request rate, concurrency, model, and hardware can all change the result; a benchmark from a different workload is not evidence for this one.

States and measurement semantics

  • MEASURED: every requested run produced at least two non-empty streaming response chunks and valid monotonic arrival times.
  • BLOCKED: the engine endpoint was not configured, or fewer than two engines produced measurements for comparison.
  • INVALID: configuration, stream shape, or fairness inputs were unsuitable for the requested metric.
  • FAILED: the configured endpoint or streaming protocol actually failed.

ITL is measured between non-empty SSE response chunks. An SSE chunk is not assumed to equal one model token. chunks_per_s is always labeled as such; the request asks the server for streamed usage with stream_options.include_usage=true, and tok_per_s is UNAVAILABLE unless the endpoint actually supplies an integer usage.completion_tokens value. This avoids turning packet counts into fake token throughput.

The comparison keeps the v0.1.0 fairness gates: measured engines must use the same model string and run count for low-level numerical diagnostics. Version 0.3.0 additionally requires the declared identity contract below for CLI comparisons; the old model-name check alone is explicitly unqualified. goodput_at_slo is the auditable fraction of runs whose TTFT is less than or equal to --slo-ttft-ms; the boundary is inclusive. It is not claimed to be a maximum sustainable request rate.

Install and verify

python -m pip install -e . pytest
python -m pytest tests -q
python -m compileall -q engine_bench.py benchmark_manifest.py tests
python engine_bench.py --engines vllm --runs 1

The last command is a fail-closed smoke test: without a declaration it emits a BLOCKED result and a valid receipt, even when an endpoint is configured. It issues no network request. No model or GPU is needed for the test suite.

The deterministic burst fixture from issue #2 can be reproduced without an endpoint:

python -c "import engine_bench as e; print(e.run_stats([.05,.06,.07,.4,.45,.5,.55,.6], 8))"

Its TTFT is 50 ms. The inter-chunk gaps are 10, 10, 330, 50, 50, 50, and 50 ms, so linear interpolation yields ITL p50/p95/p99 of 50, 246, and 313.2 ms. This is a unit fixture, not a hardware benchmark.

Exact experiment declarations (v0.3.0)

The default CLI requires --manifest comparison.json before contacting any configured engine. The complete selected cohort is checked first: a missing or changed endpoint, extra engine, mismatched request configuration, invalid digest, or different hardware declaration cannot become a successful partial comparison.

The manifest uses the closed szl.engine-comparison/v1 contract. Its template is examples/comparison.synthetic.json. Every identity in that file is a synthetic test fixture, not a qualified model, hardware profile, or live endpoint. Populate real operator-reviewed declarations before measuring a deployment. The loader never downloads or executes a manifest reference and rejects duplicate JSON keys, non-finite values, oversized inputs, unknown fields, control characters, mutable revision names, and newline-suffixed hashes. SHA checks use full-string matching, not permissive end anchors.

Declaration Meaning
Model/tokenizer revisions Exact 40-character lowercase Git subjects; not branch names
Weights/tokenizer SHA-256 Exact single artifact bytes, or SHA-256 of a separately retained canonical multi-file digest manifest
Template/adapter SHA-256 Deployed template/adapter artifact identity; null adapter means explicitly none
Quantization/precision Operator-declared configuration; no implicit conversion between engines
Suite/prompt SHA-256 Retained suite artifact plus SHA-256 of the exact UTF-8 prompt sent by the CLI
Runs/token limit/timeout Must match the requested invocation; timeout is the client's socket timeout, not a promised total wall-clock deadline
Sampling Explicit temperature, top-p, and integer seed, included in every actual HTTP request
Engine build/configuration Separate SHA-256 identities for each engine; these may differ in an engine comparison
Hardware SHA-256 Same retained hardware-profile digest across the declared cohort
Endpoint SHA-256 SHA-256 of the exact configured base URL after removing trailing slashes; no credentials allowed

benchmark_manifest.endpoint_digest() computes the endpoint digest. File digests refer to exact bytes; they are not a cryptographic demonstration that a remote server loaded those bytes. Engine names must match the supported registry. The client runs serially, performs no warmup/reset, and requires concurrency: 1 and cache_state: UNCONTROLLED; comparing cold and warm caches is not silently claimed. It freezes the preflight endpoints and refuses redirects for bound requests.

Every sample commits the exact serialized request bytes. A run binding includes the experiment, model/artifact subject, workload, per-engine build/configuration, hardware, endpoint, and request hashes. compare_bound() rejects altered bindings, missing samples, different wire requests, and omitted/failed cohort members. No failed engine is discarded to obtain a better-looking comparison.

Identity evidence remains OPERATOR_DECLARED_NOT_RUNTIME_ATTESTED. A matching manifest is not server identity attestation. Comparison reports keep runtime_identity: NOT_ATTESTED, quality_evaluation: NOT_PERFORMED, production_admission: NOT_EVALUATED, and uncontrolled cache state. They measure client-observed latency only and never decide model promotion or an engine winner. Different hardware should be evaluated as a separately designed system experiment, not misrepresented as an engine-only improvement.

An incomplete or invalid explicit --manifest invocation exits 2 after emitting its available receipt. The default offline smoke command retains its historical zero exit. Library run_engine(), compare(), and compare_engines() remain available for numerical diagnostics, labeled UNBOUND_METRICS_ONLY in comparisons. The CLI requires --unbound-diagnostics to explicitly opt into that legacy path; it cannot be combined with --manifest and is not release admission evidence.

Measure configured engines

PowerShell:

$env:VLLM_ENDPOINT = "http://127.0.0.1:8000"
$env:SGLANG_ENDPOINT = "http://127.0.0.1:30000"
python engine_bench.py --engines vllm sglang --model YOUR_MODEL --runs 10 --max-tokens 64 --slo-ttft-ms 200 --manifest comparison.json

POSIX shells:

export VLLM_ENDPOINT=http://127.0.0.1:8000
export SGLANG_ENDPOINT=http://127.0.0.1:30000
python engine_bench.py --engines vllm sglang --model YOUR_MODEL --runs 10 --max-tokens 64 --slo-ttft-ms 200 --manifest comparison.json

Supported variables are VLLM_ENDPOINT, SGLANG_ENDPOINT, LLAMACPP_ENDPOINT, MLX_ENDPOINT, TGI_ENDPOINT, and TRANSFORMERS_ENDPOINT. Each must be an absolute HTTP(S) base URL without embedded credentials, query parameters, fragments, or an invalid TCP port. Duplicate names in --engines are rejected as INVALID before any request, so receipt maps cannot silently overwrite a result. The client posts to /v1/completions.

Receipts

Every CLI invocation emits one canonical-JSON SHA-256 chain receipt using the repository's prev_hash / self_hash schema and the all-zero genesis hash. The receipt anchors:

  • a canonical SHA-256 of every complete engine result;
  • a canonical SHA-256 of each measured engine's nested ITL gap arrays;
  • the ordered engine selection, benchmark configuration, and a prompt hash; and
  • the comparison verdict and every engine state; and
  • the canonical manifest digest and each exact wire-request/run binding (schema version 3).

Each run sample also includes its raw itl_gaps_ms array and matching itl_gaps_sha256, so the hash can be independently recomputed. UNSIGNED_HONEST proves integrity and order only; it does not claim signer identity or external attestation.

Verification scope

The test suite covers strict declarations, complete-cohort comparisons, actual loopback HTTP request bytes/sampling fields, redirect refusal, preflight zero- request behavior, endpoint freezing, bad JSON/digests, failure retention, and receipt commitments, in addition to the original streaming and numerical tests. Loopback test responses are fixtures: no trained model, remote production system, GPU performance, energy measurement, or quality result is represented by CI.

Apache-2.0 · Doctrine v11 · SZL Holdings

Metadata

Release files for szl-engine-bench 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for szl-engine-bench 0.3.0
File Size Uploaded
szl_engine_bench-0.3.0.tar.gz 30.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for szl-engine-bench 0.3.0
File Interpreter ABI Platform
szl_engine_bench-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 51.6 kB

Release files / szl_engine_bench-0.3.0.tar.gz

Download URL szl_engine_bench-0.3.0.tar.gz
Size 30.8 kB
Tags Source
SHA-256 checksum
How to use checksums
31fe42946b813ae51ef66b5f331ac4310d32923fd1935f71ca693bbda4ca6296
BLAKE2b-256 checksum
How to use checksums
28270697af910699a223658cc7e3ac50c9323a20c1703075298ad70fc4103e8a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / szl_engine_bench-0.3.0-py3-none-any.whl

Download URL szl_engine_bench-0.3.0-py3-none-any.whl
Size 20.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1f27dc8f0ec13aedfd5f14aa42d4245e52f3fb87d415ce86a90ef8a37232cd9c
BLAKE2b-256 checksum
How to use checksums
deb43a88b00a8deb4bca9708df853586a8be3f574b29fd4371746a756f58caba
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page