Skip to main content

Inference Profiler

A profiling tool for LLM inference servers — the perf / PyTorch Profiler equivalent for serving. Point it at any OpenAI-compatible endpoint (vLLM, SGLang, TensorRT-LLM, TGI, llama.cpp server, …) and it measures where time goes, diagnoses bottlenecks, and renders a self-contained HTML report.

Everything runs locally. No cloud backend, no auth, no telemetry.

Why

Engineers usually know latency is bad, but not why. Inference Profiler answers:

  • Why is TTFT high — prefill compute, or requests stuck in the scheduler queue?
  • Is batching actually working, or is the GPU idle between requests?
  • Where is the throughput/latency knee as concurrency grows?
  • Are a few long prompts ruining tail latency for everyone?
  • Is the server compute-saturated, or starved for work?

Install

pip install inference-profiler          # core
pip install "inference-profiler[gpu]"   # + NVML GPU metrics (NVIDIA)

Quick start

inference-profiler profile \
    --endpoint http://localhost:8000/v1 \
    --model llama3 \
    --requests 500 \
    --concurrency 32 \
    --stream

Console output shows per-stage latency/throughput, findings, and an ASCII waterfall. Full artifacts land in results/:

results/
├── config.json      # exact run configuration (reproducible)
├── metrics.jsonl    # one line per request: timestamps, tokens, errors
├── system.jsonl     # GPU/CPU/RAM/network samples over time
├── summary.json     # aggregated metrics + derived estimates
├── timeline.json    # per-request phase spans (queue/waiting/decode)
├── findings.json    # automatic analysis results with evidence
├── report.html      # self-contained interactive report
└── charts/          # each chart as a standalone HTML file

No GPU? Try the mock server

python examples/mock_server.py --port 8399 --batch-size 8 &
inference-profiler profile --endpoint http://127.0.0.1:8399/v1 --model demo \
    --requests 100 --concurrency 4,8,16,32

The mock server queues requests beyond its batch size — the profiler will find the knee.

What it measures

Category Metrics
Latency TTFT, inter-token latency, end-to-end · P50/P90/P95/P99
Throughput requests/s, output tokens/s, total tokens/s
Tokens prompt/output totals, means, min/max, per-request decode speed
System GPU utilization & memory (NVML), CPU, RAM, network
Derived estimates batch utilization, prefill vs scheduler-wait split, GPU idle %, saturation knee

Measured numbers and heuristic estimates are kept strictly separate — estimates are suffixed *_estimate and their method is documented in docs/methodology.md. Findings are only ever generated from measured data.

Common recipes

# Concurrency sweep: find the throughput/latency knee
inference-profiler profile --endpoint http://localhost:8000/v1 --model llama3 \
    --requests 200 --concurrency 8,16,32,64,128

# Open-loop (Poisson arrivals at 25 req/s) — reveals queueing under overload
inference-profiler profile --endpoint http://localhost:8000/v1 --model llama3 \
    --requests 500 --arrival-mode open --rate 25

# Burst traffic
inference-profiler profile --endpoint http://localhost:8000/v1 --model llama3 \
    --requests 300 --arrival-mode burst --rate 50 --burst-size 40 --burst-idle 5

# Your own prompts (JSONL with a "prompt" field), controlled output lengths
inference-profiler profile --endpoint http://localhost:8000/v1 --model llama3 \
    --dataset examples/prompts.jsonl --output-tokens 64-256 --requests 200

# Everything in a config file (CLI flags override it)
inference-profiler profile --config examples/sweep_config.json

# Re-generate the report from raw data (tweak nothing, re-analyze)
inference-profiler report --results results/

Reading the report

  • Findings — ranked observations with the evidence that triggered them, e.g. “GPU utilization averaged only 42%” or “Throughput saturates near concurrency 16”.
  • Timeline — a waterfall of every request split into client-queue, waiting (TTFT), and decode phases. Growing gray/yellow bands = queueing; long blue = decode-bound.
  • Charts — latency/TTFT histograms, percentile curves, tokens/s over time, in-flight concurrency, GPU utilization & memory, prompt/output length vs latency, and concurrency scaling (on sweeps).

Architecture

cli ─▶ config (pydantic) ─▶ runner ──▶ workload   (synthetic / JSONL prompts)
                              │    ──▶ scheduler  (closed / open-poisson / burst)
                              │    ──▶ collector  (httpx + SSE timing per request)
                              │    ──▶ gpu_monitor (NVML/psutil background sampling)
                              ▼
                    metrics ─▶ timeline ─▶ analysis ─▶ charts ─▶ report.html

See docs/architecture.md. Adding a non-OpenAI backend means implementing one execute_request variant; everything downstream operates on RequestRecords and is backend-agnostic.

Development

git clone https://github.com/SarnadAbhilash/inference_profiler
cd inference-profiler
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest            # no GPU or server required — uses a mock transport
ruff check .
pre-commit install

License

MIT

Metadata

Release files for inference-profiler 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for inference-profiler 0.1.0
File Size Uploaded
inference_profiler-0.1.0.tar.gz 36.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for inference-profiler 0.1.0
File Interpreter ABI Platform
inference_profiler-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 70.5 kB

Release files / inference_profiler-0.1.0.tar.gz

Download URL inference_profiler-0.1.0.tar.gz
Size 36.5 kB
Tags Source
SHA-256 checksum
How to use checksums
a957f7c16449f75328a72f9fa8ac268a1b5974345237f18f8f363f71950c243e
BLAKE2b-256 checksum
How to use checksums
2c604f9800844ec519e31c4ca7b65cc8c04d2e310bc10f03154e3db26e82411a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.11

Release files / inference_profiler-0.1.0-py3-none-any.whl

Download URL inference_profiler-0.1.0-py3-none-any.whl
Size 34.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b3e53b0c15e34133792693aee7ae9c94487248f7b83654ab5a662476dda0658b
BLAKE2b-256 checksum
How to use checksums
90d88bb5f50246c5f2d98a125a70293df250f87141126f9c9ddcda489dfa4789
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.11

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page