Inference Profiler
A profiling tool for LLM inference servers — the perf / PyTorch Profiler equivalent for serving.
Point it at any OpenAI-compatible endpoint (vLLM, SGLang, TensorRT-LLM, TGI, llama.cpp server, …)
and it measures where time goes, diagnoses bottlenecks, and renders a self-contained HTML report.
Everything runs locally. No cloud backend, no auth, no telemetry.
Why
Engineers usually know latency is bad, but not why. Inference Profiler answers:
- Why is TTFT high — prefill compute, or requests stuck in the scheduler queue?
- Is batching actually working, or is the GPU idle between requests?
- Where is the throughput/latency knee as concurrency grows?
- Are a few long prompts ruining tail latency for everyone?
- Is the server compute-saturated, or starved for work?
Install
pip install inference-profiler # core
pip install "inference-profiler[gpu]" # + NVML GPU metrics (NVIDIA)
Quick start
inference-profiler profile \
--endpoint http://localhost:8000/v1 \
--model llama3 \
--requests 500 \
--concurrency 32 \
--stream
Console output shows per-stage latency/throughput, findings, and an ASCII waterfall.
Full artifacts land in results/:
results/
├── config.json # exact run configuration (reproducible)
├── metrics.jsonl # one line per request: timestamps, tokens, errors
├── system.jsonl # GPU/CPU/RAM/network samples over time
├── summary.json # aggregated metrics + derived estimates
├── timeline.json # per-request phase spans (queue/waiting/decode)
├── findings.json # automatic analysis results with evidence
├── report.html # self-contained interactive report
└── charts/ # each chart as a standalone HTML file
No GPU? Try the mock server
python examples/mock_server.py --port 8399 --batch-size 8 &
inference-profiler profile --endpoint http://127.0.0.1:8399/v1 --model demo \
--requests 100 --concurrency 4,8,16,32
The mock server queues requests beyond its batch size — the profiler will find the knee.
What it measures
| Category | Metrics |
|---|---|
| Latency | TTFT, inter-token latency, end-to-end · P50/P90/P95/P99 |
| Throughput | requests/s, output tokens/s, total tokens/s |
| Tokens | prompt/output totals, means, min/max, per-request decode speed |
| System | GPU utilization & memory (NVML), CPU, RAM, network |
| Derived estimates | batch utilization, prefill vs scheduler-wait split, GPU idle %, saturation knee |
Measured numbers and heuristic estimates are kept strictly separate — estimates are
suffixed *_estimate and their method is documented in
docs/methodology.md. Findings are only ever generated from
measured data.
Common recipes
# Concurrency sweep: find the throughput/latency knee
inference-profiler profile --endpoint http://localhost:8000/v1 --model llama3 \
--requests 200 --concurrency 8,16,32,64,128
# Open-loop (Poisson arrivals at 25 req/s) — reveals queueing under overload
inference-profiler profile --endpoint http://localhost:8000/v1 --model llama3 \
--requests 500 --arrival-mode open --rate 25
# Burst traffic
inference-profiler profile --endpoint http://localhost:8000/v1 --model llama3 \
--requests 300 --arrival-mode burst --rate 50 --burst-size 40 --burst-idle 5
# Your own prompts (JSONL with a "prompt" field), controlled output lengths
inference-profiler profile --endpoint http://localhost:8000/v1 --model llama3 \
--dataset examples/prompts.jsonl --output-tokens 64-256 --requests 200
# Everything in a config file (CLI flags override it)
inference-profiler profile --config examples/sweep_config.json
# Re-generate the report from raw data (tweak nothing, re-analyze)
inference-profiler report --results results/
Reading the report
- Findings — ranked observations with the evidence that triggered them, e.g. “GPU utilization averaged only 42%” or “Throughput saturates near concurrency 16”.
- Timeline — a waterfall of every request split into client-queue, waiting (TTFT), and decode phases. Growing gray/yellow bands = queueing; long blue = decode-bound.
- Charts — latency/TTFT histograms, percentile curves, tokens/s over time, in-flight concurrency, GPU utilization & memory, prompt/output length vs latency, and concurrency scaling (on sweeps).
Architecture
cli ─▶ config (pydantic) ─▶ runner ──▶ workload (synthetic / JSONL prompts)
│ ──▶ scheduler (closed / open-poisson / burst)
│ ──▶ collector (httpx + SSE timing per request)
│ ──▶ gpu_monitor (NVML/psutil background sampling)
▼
metrics ─▶ timeline ─▶ analysis ─▶ charts ─▶ report.html
See docs/architecture.md. Adding a non-OpenAI backend means
implementing one execute_request variant; everything downstream operates on
RequestRecords and is backend-agnostic.
Development
git clone https://github.com/SarnadAbhilash/inference_profiler
cd inference-profiler
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest # no GPU or server required — uses a mock transport
ruff check .
pre-commit install
License
MIT
Metadata
Release files for inference-profiler 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| inference_profiler-0.1.0.tar.gz | 36.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| inference_profiler-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 70.5 kB
Release files / inference_profiler-0.1.0.tar.gz
| Download URL | inference_profiler-0.1.0.tar.gz |
|---|---|
| Size | 36.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
a957f7c16449f75328a72f9fa8ac268a1b5974345237f18f8f363f71950c243e
|
|
BLAKE2b-256 checksum How to use checksums |
2c604f9800844ec519e31c4ca7b65cc8c04d2e310bc10f03154e3db26e82411a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.11
|
Release files / inference_profiler-0.1.0-py3-none-any.whl
| Download URL | inference_profiler-0.1.0-py3-none-any.whl |
|---|---|
| Size | 34.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
b3e53b0c15e34133792693aee7ae9c94487248f7b83654ab5a662476dda0658b
|
|
BLAKE2b-256 checksum How to use checksums |
90d88bb5f50246c5f2d98a125a70293df250f87141126f9c9ddcda489dfa4789
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.11
|