This release is a pre-release and may not be stable for production use.
vTune
vTune is a local-first experimentation and optimization tool for vLLM serving configurations. Users define the parameters and workloads they care about; vTune manages the server lifecycle, runs repeatable benchmarks, explores the search space, and reports which configurations performed best.
vTune is alpha software targeting Linux with NVIDIA GPUs and Python 3.11–3.12.
Documentation · Quick start · PyPI
The current code is verified with vLLM 0.28.0 and GuideLLM 0.7.3 on WSL2
with an RTX 3080. That host required
VLLM_USE_V2_MODEL_RUNNER: "0" because UVA was unavailable and
VLLM_USE_FLASHINFER_SAMPLER: "0" because the CUDA compiler toolkit was
not installed. Native Linux systems may not require these settings.
Other combinations may work but are not yet verified.
The published py3-none-any wheel installs on Linux and Windows. Configuration
validation and stored-result inspection work on Windows, but starting an
experiment is supported only on Linux because vLLM has no native Windows
runtime.
Installation
Choose the installation that matches what you want to do:
| Goal | Command | Platform |
|---|---|---|
| Run complete experiments | pip install "vtune[runtime]" |
Linux/WSL with NVIDIA GPU |
| Read configs, results, and reports | pip install vtune |
Linux, Windows, or macOS |
The core package intentionally does not install GPU frameworks. The runtime
extra adds vLLM and GuideLLM, which select large PyTorch/CUDA dependencies for
the machine. See the installation guide
for virtual environments, CUDA guidance, and verification commands.
Quick start
Create and activate a Python 3.11 or 3.12 virtual environment on Linux or WSL, then install the complete experiment runtime:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install "vtune[runtime]"
vllm --help
guidellm --help
Create experiment.yaml:
schema_version: 1
experiment:
name: first-run
server:
model: /models/opt-125m
gpu-memory-utilization: 0.8
tune:
max-num-seqs:
values: [8, 16]
benchmark:
engine: guidellm # Default. Use vllm for `vllm bench serve`.
runs:
- name: throughput
profile:
kind: throughput
max_concurrency: 16
constraints:
- kind: max_requests
count: 10
data:
- kind: synthetic_text
prompt_tokens: 32
output_tokens: 16
optimization:
maximize: output_tokens_per_second
sampler: tpe
trials: 2
Run it:
vtune --config experiment.yaml
The short form is vtune -c experiment.yaml. The command validates the file,
runs the experiment, persists results, and generates its exports and report.
vTune binds vLLM to 127.0.0.1 by default. Set server.host explicitly
only when the benchmark server must be reachable from another host.
Fixed vLLM flags go directly under server; tunable flags use top-level
tune. Fixed and tunable environment variables use env and tune_env.
See the configuration guide
for categorical, boolean, integer-range, float-range, list, and environment
examples. The complete YAML
and benchmark guide show
every supported control with copyable examples.
To use vLLM's native benchmark, set benchmark.engine: vllm. Its args
map directly to vllm bench serve flags; vTune supplies the model, server
address, and JSON output path:
benchmark:
engine: vllm
runs:
- name: throughput
args:
dataset-name: random
random-input-len: 32
random-output-len: 16
num-prompts: 100
request-rate: inf
max-concurrency: 16
Terminal output is concise by default. To stream server and benchmark logs:
vtune --config experiment.yaml --verbose
The persistent equivalent uses GuideLLM's logging level names:
logging:
level: DEBUG
Supported levels are DEBUG, INFO, WARNING, ERROR, and CRITICAL.
Full per-trial log files are always saved. --verbose overrides the configured
level with DEBUG for that invocation.
Retry one or more selected trials into a new immutable linked run:
vtune retry --run runs/EXPERIMENT/RUN_ID \
--trial trial-0001 --trial trial-0004
The source run is never modified.
Display every stored vLLM and GuideLLM command for a trial without executing anything:
vtune reproduce --run runs/EXPERIMENT/RUN_ID --trial trial-0001
Each completed run also contains a self-contained report.html decision
dashboard with the best observed configuration, baseline comparison, score
history, throughput/latency tradeoff, and observed parameter effects.
Random and TPE runs never execute the same resolved configuration twice. If
optimization.trials exceeds the unique search space, vTune warns and runs
every unique configuration once.
Multiple independent trials can run on explicitly assigned, non-overlapping GPU sets and ports. Sequential execution remains the default. See parallel trials for the YAML and measurement caveats.
Product documents
- First MVP specification
- Future implementation roadmap
- Architecture overview and early sketch
- Editable Draw.io architecture diagram
- Contributor guide
- Release notes
The MVP specification defines the first releasable version and its acceptance criteria. The roadmap describes capabilities that should be designed for now but implemented after the core experiment loop is reliable.
vTune is available under the MIT License.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file vtune-0.1.0a5.tar.gz.
File metadata
- Download URL: vtune-0.1.0a5.tar.gz
- Upload date:
- Size: 56.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4e4814beb810eea3e7c9f8fb3ffd56cd15963e121068beef8cdcd0c5d28e72ed
|
|
| MD5 |
a6bbe8bce9232d87ae3e2035013720d4
|
|
| BLAKE2b-256 |
08fcf12378907cebd862ffff9b3926a824db311743bb9433c33d6f1e8210ee74
|
Provenance
The following attestation bundles were made for vtune-0.1.0a5.tar.gz:
Publisher:
publish.yml on brtydse100/vTune
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
vtune-0.1.0a5.tar.gz -
Subject digest:
4e4814beb810eea3e7c9f8fb3ffd56cd15963e121068beef8cdcd0c5d28e72ed - Sigstore transparency entry: 2655477650
- Sigstore integration time:
-
Permalink:
brtydse100/vTune@e48003f187cfea5c81ca605d058ea6ebffcb93d7 -
Branch / Tag:
refs/tags/v0.1.0a5 - Owner: https://github.com/brtydse100
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@e48003f187cfea5c81ca605d058ea6ebffcb93d7 -
Trigger Event:
release
-
Statement type:
File details
Details for the file vtune-0.1.0a5-py3-none-any.whl.
File metadata
- Download URL: vtune-0.1.0a5-py3-none-any.whl
- Upload date:
- Size: 79.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f99d3077156c5cdba05aa92a8919706fb6994792ed18fe45934edb96f357a9d8
|
|
| MD5 |
f0b4b7393c01034f07be299d27d78a7c
|
|
| BLAKE2b-256 |
9cf5e70a6ff39ba9b408f11b4b2dcac2f3aee9613f8c09c91aa2c9c60b433440
|
Provenance
The following attestation bundles were made for vtune-0.1.0a5-py3-none-any.whl:
Publisher:
publish.yml on brtydse100/vTune
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
vtune-0.1.0a5-py3-none-any.whl -
Subject digest:
f99d3077156c5cdba05aa92a8919706fb6994792ed18fe45934edb96f357a9d8 - Sigstore transparency entry: 2655477675
- Sigstore integration time:
-
Permalink:
brtydse100/vTune@e48003f187cfea5c81ca605d058ea6ebffcb93d7 -
Branch / Tag:
refs/tags/v0.1.0a5 - Owner: https://github.com/brtydse100
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@e48003f187cfea5c81ca605d058ea6ebffcb93d7 -
Trigger Event:
release
-
Statement type: