AI Infra Bench
AI Infra Bench measures OpenAI-compatible inference endpoints, evaluates model outputs, and finds the highest load that satisfies a service-level objective. It is backend-independent and does not require a serving framework SDK.
Install
pip install ai-infra-bench
Python 3.10 or newer is required.
CLI
Send one request:
aib req --base-url http://127.0.0.1:30000 --prompt "Who are you?"
Run a random-token benchmark:
aib bench \
--base-url http://127.0.0.1:30000 \
--dataset random \
--input-len 1024 \
--output-len 256 \
--random-range-ratio 0.5 \
--num-requests 100 \
--max-concurrency 16
When --random-range-ratio is set, each random request samples its input and
output length uniformly from the configured ratio of the target length up to
the target length. The default value of 1.0 preserves fixed-length requests.
Run the packaged GPQA Diamond multiple-choice workload (198 questions):
aib bench \
--base-url http://127.0.0.1:30000 \
--dataset gpqa \
--num-requests 198 \
--max-concurrency 16
Evaluate the text-only portion of Humanity's Last Exam directly through a chat-completions endpoint. Local JSONL files are supported; image questions are skipped because this adapter does not require a multimodal runtime:
aib eval-dataset \
--evals hle \
--dataset-path ./hle.jsonl \
--num-shots 0 \
--num-questions 20 \
--base-url http://127.0.0.1:30000 \
--model your-model
Other commands cover dataset evaluation, logits and hidden-state comparison, metric plotting, and local Prometheus monitoring:
aib --help
Replay session-shaped JSONL payloads with one concurrency slot per session:
aib session-bench 'sessions/**/*.jsonl' \
--max-concurrency 16 \
--num-warmup-sessions 3
Requests in each file are sent in order. Different files run concurrently up
to --max-concurrency.
SLO Search
SLO searches are configured in YAML. The command probes the configured range
and returns the highest max_concurrency or request_rate that satisfies every
condition:
endpoint:
base_url: http://127.0.0.1:30000
api_key: EMPTY
model: null
request:
payload:
messages:
- role: user
content: hello
benchmark:
num_requests: 100
warmup_requests: 5
max_concurrency: 32
request_rate: inf
search:
parameter: max_concurrency
min: 1
max: 64
conditions:
- metric: success_rate
operator: ">="
value: 0.99
- metric: p99_latency_ms
operator: "<"
value: 3000
Run it with:
aib slo examples/slo.yaml -o slo-result.yaml
Output
SLO runs write a YAML report with every probe, its metrics, and condition results. Request benchmarks can also write machine-readable JSON or JSONL metrics for later visualization.
Development
python -m pip install -e .
pytest
Licensed under the Apache License 2.0.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file ai_infra_bench-0.1.4.post7.tar.gz.
File metadata
- Download URL: ai_infra_bench-0.1.4.post7.tar.gz
- Upload date:
- Size: 82.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.12.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
38a9fb2e4360afc2ee09ffab17d2252bc804e22d21406a7fc0774cd5afc221d2
|
|
| MD5 |
d8519d2b3e622e8d7cd99c61bb47ed86
|
|
| BLAKE2b-256 |
725ddd0e116fefc2561bab655c2ce273666b79ed3ea5b8262aa154c275204d4f
|
File details
Details for the file ai_infra_bench-0.1.4.post7-py3-none-any.whl.
File metadata
- Download URL: ai_infra_bench-0.1.4.post7-py3-none-any.whl
- Upload date:
- Size: 102.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.12.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a4f9dd1708a160576f40bea6eb3aeff24b2c00be3e3f9cb8d1a37761e4f5a1bf
|
|
| MD5 |
50d89410a6a1e0fe2d1d5f4d24fa61b4
|
|
| BLAKE2b-256 |
acc93ada64a3aaa09dbf888797fe8e09ff3a96078a611b66db0b52f006156184
|