vLLM Metrics Monitor (vmm) is a lightweight dashboard that scrapes Prometheus metrics from vLLM, persists to SQLite, and serves a real-time web UI. Zero external dependencies — pure Python standard library.
Monitor multiple vLLM servers at once: pass several metrics endpoints and switch between them (labeled by model name) from the dropdown in the dashboard header.
🚀 Quick Start
# Install with uv
uv tool install vllm-metrics-monitor
# Or install with pip: `pip install vllm-metrics-monitor`
# Launch dashboard
vmm http://your-vllm:8000/metrics
Open http://localhost:8080 — that's it.
📖 Usage
vmm [URL...] [OPTIONS]
Positional:
URL vLLM Prometheus metrics endpoint(s); multiple URLs are
scraped concurrently and switchable in the dashboard
(default: http://localhost:8000/metrics)
Options:
-p, --port PORT Dashboard HTTP port (default: 8080)
-i, --interval SEC Scrape interval in seconds (default: 3)
--retention HOURS Data retention period (default: 720, i.e. 30 days)
--db PATH SQLite database path (default: ~/.vmm/data.db)
--reset Delete existing database and start fresh
--debug Enable debug logging
Examples
vmm http://vllm-server:8000/metrics -p 9090 -i 5
vmm http://vllm-1:8000/metrics http://vllm-2:8000/metrics # multiple servers
vmm http://vllm-server:8000/metrics --reset
vmm http://vllm-server:8000/metrics --retention 72 --db /data/vmm.db
Docker
# Build
docker build -t vmm .
# Run
docker run -d --network host vmm http://localhost:8000/metrics
# Or use docker compose (space-separated for multiple endpoints)
METRICS_URLS="http://192.168.1.100:8000/metrics http://192.168.1.101:8000/metrics" docker compose up -d
🏗️ Architecture
graph LR
A[vLLM /metrics] -->|scrape every 3s| B[vmm]
B --> C[Scraper Thread]
C --> D[(SQLite<br/>~/.vmm/data.db)]
B --> E[HTTP Server]
E -->|JSON API| F[Browser<br/>Chart.js Dashboard]
📈 Monitored Metrics
| Metric | Source | Type |
|---|---|---|
| Running Requests | vllm:num_requests_running |
Gauge |
| Waiting Requests | vllm:num_requests_waiting |
Gauge |
| KV Cache Usage | vllm:kv_cache_usage_perc |
Gauge |
| Cache Hit Rate | prompt_tokens_cached / prompt_tokens |
Derived |
| Requests/s | vllm:request_success_total delta |
Counter rate |
| Output Tokens/s | vllm:generation_tokens_total delta |
Counter rate |
| Input Tokens/s | vllm:prompt_tokens_total delta |
Counter rate |
| TTFT | time_to_first_token_seconds |
Histogram avg |
| ITL | inter_token_latency_seconds |
Histogram avg |
| E2E Latency | e2e_request_latency_seconds |
Histogram avg |
| Uptime | process_start_time_seconds |
Gauge |
License
MIT
Metadata
Release files for vllm-metrics-monitor 0.4.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| vllm_metrics_monitor-0.4.1.tar.gz | 139.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| vllm_metrics_monitor-0.4.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 174.9 kB
Release files / vllm_metrics_monitor-0.4.1.tar.gz
| Download URL | vllm_metrics_monitor-0.4.1.tar.gz |
|---|---|
| Size | 139.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4c504ec5f01eb2bc15588dfff4634e52e1e85794e760e3a62e67175ad813e8d9
|
|
BLAKE2b-256 checksum How to use checksums |
96d9208f0c38586561d95d957f6d258c33a87a81157b28e3eb5844a66739f4f2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 18, 2026.
Transparency logRelease files / vllm_metrics_monitor-0.4.1-py3-none-any.whl
| Download URL | vllm_metrics_monitor-0.4.1-py3-none-any.whl |
|---|---|
| Size | 35.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5dec0fda64ace92bac2c381decfb8b6b63d885d01c140152a93e119ae612c86a
|
|
BLAKE2b-256 checksum How to use checksums |
8f6c16ff6ae4ce2996d72e97a059323d9f3304c9cfb79f26683847e903762568
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 18, 2026.
Transparency log