Skip to main content

llm-bench

CI PyPI Python License Tests

LLM API latency benchmarker with the statistical rigor of HFT latency analysis. Measures TTFT, total latency, inter-token latency, throughput, and cost per call, and reports p50/p95/p99 instead of a single misleading average.

Most benchmarking tools report mean latency. In trading systems, reporting the mean is the mistake that gets you laughed out of the room: latency lives in the tail, and p99 is what your worst user actually feels. This tool treats LLM APIs the same way.

Works with Anthropic, OpenAI, and any OpenAI-compatible endpoint (Groq, Together, vLLM, OpenRouter). Ships with a mock provider so you can try the whole pipeline with no API key.

Built by Saksham Arora.


Quickstart

# Install from PyPI
pip install llm-latency-bench

# Or run instantly with no install, no key (mock provider)
uvx --from llm-latency-bench llm-bench \
  --provider mock --model demo --prompt "hi" -n 20
# Clone and develop locally
git clone https://github.com/saksham10arora-dotcom/llm-bench
cd llm-bench && uv sync && uv run llm-bench --help

The demo that sold me on building it

A real run, 100 requests against Groq (llama-3.3-70b-versatile, free tier):

llm-bench  openai/llama-3.3-70b-versatile
                 (n=100)
+--------+-------+-------+-------+-------+
| Metric |   p50 |   p95 |   p99 |  mean |
+--------+-------+-------+-------+-------+
| TTFT   | 2.20s | 2.32s | 2.33s | 1.71s |
| Total  | 2.28s | 2.39s | 2.41s | 1.79s |
| ITL    |   0ms |  10ms |  13ms |   2ms |
+--------+-------+-------+-------+-------+
Throughput: 536.1 tok/s  |  Cost/call: $0.000058  |  Errors: 0/100

Look at TTFT. The mean (1.71s) is lower than the median (2.20s). That should be impossible for a normal latency distribution, and it is: this one is bimodal. The first ~30 requests returned in ~200ms, then the free tier's rate limiter kicked in and everything after waited ~2.2s. The mean blended two completely different latency regimes into one number that describes neither. A mean-only tool would have printed "1.7s, fine" and hidden the whole story. Catching exactly this is why the tool exists.


What it measures

Metric What it is Why it is computed this way
TTFT Dispatch to first content chunk What the user perceives as "it's responding." Queueing + prefill included on purpose.
Total Dispatch to last content chunk Trailing usage frames and connection teardown excluded.
ITL Gap between consecutive chunks Pooled across all requests, percentiles on raw gaps, so a single stall survives to p99.
Throughput (tokens - 1) / (total - TTFT) Pure decode rate. Including prefill would mislabel two phases as one.
Cost/call prompt + completion tokens x price From exact API usage counts when available, not word-count estimates.
Errors failed / total Counted and reported, never averaged into the latency percentiles.

Every metric is reported at p50, p95, p99, and mean.

Cache-aware TTFT (cold vs warm)

Prompt caching splits latency into two regimes: a cold miss processes the whole prompt, a warm hit reuses cached work and returns first token much sooner. Averaging across them gives a p99 no request actually saw, the same tail-hiding trap the mean falls into. So llm-bench reads the cache-read token counts the APIs already return (prompt_tokens_details.cached_tokens on OpenAI-compatible endpoints, cache_read_input_tokens on Anthropic), tags each request, and reports cold and warm TTFT separately whenever any warm hits show up. The JSON output always carries a cache block with the cold/warm counts and p99s. Added after a reader pointed out that an untagged average silently blends the two.


Usage

# Benchmark Anthropic
llm-bench --provider anthropic --model claude-sonnet-4-6 \
  --prompt "Explain recursion in one sentence" -n 100 --warmup 5

# Any OpenAI-compatible endpoint (Groq, Together, vLLM, OpenRouter)
llm-bench --provider openai --model llama-3.3-70b-versatile \
  --base-url https://api.groq.com/openai/v1 \
  --prompt "Explain recursion" -n 100

# Compare two providers side-by-side (prints a delta row)
llm-bench --provider anthropic --model claude-sonnet-4-6 \
  --prompt "Explain recursion" -n 100 \
  --compare openai:gpt-4o

# Latency under load
llm-bench --provider anthropic --model claude-sonnet-4-6 \
  --prompt "Explain recursion" -n 100 --concurrency 8

# Export for further analysis
llm-bench --provider anthropic --model claude-sonnet-4-6 \
  --prompt "Explain recursion" -n 100 \
  --markdown results.md --json results.json

Prompts can be inlined or read from a file with --prompt @prompt.txt.

Flags

Flag Default Description
--provider required anthropic, openai, or mock
--model required Model ID
--prompt required Prompt string, or @file.txt to read from a file
-n 50 Number of measured requests
--concurrency 1 Max in-flight requests
--max-tokens 256 Max completion tokens
--warmup 3 Warmup requests, excluded from stats
--timeout 30 Per-request timeout in seconds
--base-url None OpenAI-compatible endpoint override
--api-key None API key override (defaults to env var)
--markdown None Write a markdown report
--json None Write raw per-request JSON + run metadata
--compare None provider:model for a side-by-side comparison

Auth

export ANTHROPIC_API_KEY=sk-ant-...
export OPENAI_API_KEY=sk-...   # also used for Groq / Together / vLLM keys with --base-url

Supported providers

Provider How Token counts Cost
Anthropic native streaming SDK exact (usage) priced
OpenAI streaming SDK exact (usage) priced
Groq / Together / vLLM / OpenRouter OpenAI-compatible, via --base-url exact, or chunk-count fallback priced when model is known
Mock scripted delays, no network synthetic n/a

Architecture

The core is fully decoupled from the CLI, so the runner and stats engine can be reused (a planned MCP wrapper will import them directly without touching cli.py).

src/llm_bench/
  cli.py              # Click CLI: arg parsing + wiring only, no business logic
  core/
    adapters/
      base.py         # StreamEvent / RequestResult / Adapter protocol
      anthropic.py    # native Anthropic streaming
      openai.py       # OpenAI + base_url override (Groq / vLLM / Together)
      mock.py         # scripted-delay fake adapter, zero network
    runner.py         # async runner: semaphore concurrency, warmup, timeout
    stats.py          # nearest-rank percentiles, pooled ITL, decode-rate throughput
    pricing.py        # per-model price table, cost estimate
    report.py         # rich table + markdown / JSON export

Data flows one direction: cli -> runner -> adapter.stream() -> stats.compute() -> report. Adapters only ever yield timestamped stream events; all the math lives in stats.py, which has zero knowledge of HTTP.


Methodology

The numbers are only worth sharing if the measurement is honest. Precisely what this tool does:

  • Clock: all timestamps come from time.monotonic_ns(). Wall clocks can step backwards; monotonic clocks can't.
  • TTFT: time from request dispatch to the first content chunk. Includes connection setup, queueing, and prefill, because that is what the user waits through.
  • Total: dispatch to the last content chunk. Trailing usage frames and connection teardown are excluded.
  • ITL: every gap between consecutive content chunks is measured individually and pooled across all requests. Percentiles are over raw gaps, never per-request averages, which would hide exactly the tail spikes you are hunting. Note: providers stream chunks, not tokens, so a chunk may carry several tokens.
  • Throughput: decode rate, (completion_tokens - 1) / (total - TTFT). Including TTFT would blend prefill into a number that claims to be generation speed.
  • Warmup: warmup requests fully complete before measurement begins, even under --concurrency. They never overlap the measured window.
  • Errors: failed and timed-out requests are counted and reported, but excluded from latency percentiles. A failed call has no valid latency.
  • Percentiles: nearest-rank on the raw samples. No interpolation, no smoothing.
  • Token counts: exact counts from API usage when the endpoint provides them, chunk-count fallback otherwise. Cost is computed from exact counts, not estimates.
  • Sample size: with n=50, "p99" is just your second-slowest request. The tool warns below n=100. For stable tails, use n=200+.
  • Compare-mode caveat: the two targets run sequentially, not interleaved, so a network blip during one run biases that target. Treat small deltas (<20%) as noise.

Development

uv sync --group dev
uv run pytest -v          # 28 tests

The test suite makes no network calls. A scripted mock adapter drives the runner with controllable delays, so percentile math, concurrency, warmup ordering, and timeout handling are all verified deterministically. The credibility-critical code (percentiles, pooled ITL) is tested against known distributions, including a case where two tail spikes hide among 98 fast samples and must surface at p99. CI runs the full suite on Python 3.11 and 3.12 plus a keyless end-to-end run on every push.


Roadmap

  • Publish to PyPI (pip install llm-latency-bench)
  • Task-type context flag (--task text/code/pdf/image/chat)
  • MCP server wrapper (expose benchmark as a tool to Claude / MCP clients)
  • Interleaved compare mode (alternate requests to cancel out network drift)
  • Histogram / sparkline output for distribution shape at a glance

License

MIT. See LICENSE.

Release files for llm-latency-bench 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llm-latency-bench 0.3.0
File Size Uploaded
llm_latency_bench-0.3.0.tar.gz 42.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llm-latency-bench 0.3.0
File Interpreter ABI Platform
llm_latency_bench-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 60.2 kB

Release files / llm_latency_bench-0.3.0.tar.gz

Download URL llm_latency_bench-0.3.0.tar.gz
Size 42.0 kB
Tags Source
SHA-256 checksum
How to use checksums
7b9ef2da2ed9db41f37b9f7f17393d8f1b16175e9c9bcbee84c75f4d65ffe050
BLAKE2b-256 checksum
How to use checksums
83196784496327b6f8d8405ac7136df0c3552b75dce7d546a0a960b7dcd8e43a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.11.28 {"installer":{"name":"uv","version":"0.11.28","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / llm_latency_bench-0.3.0-py3-none-any.whl

Download URL llm_latency_bench-0.3.0-py3-none-any.whl
Size 18.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f3db9ed10430e7977da932979f76487e73c99d64bb923ddc872b5d7c61f80fcb
BLAKE2b-256 checksum
How to use checksums
fb8414b71c03265f51ec6f9538a74ce64c9a40942e3d1c1ebf9f45b24e0d567c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.11.28 {"installer":{"name":"uv","version":"0.11.28","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

0.2.1

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page