Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

vTune

vTune is a local-first experimentation and optimization tool for vLLM serving configurations. Users define the parameters and workloads they care about; vTune manages the server lifecycle, runs repeatable benchmarks, explores the search space, and reports which configurations performed best.

vTune is alpha software targeting Linux with NVIDIA GPUs and Python 3.11–3.12.

Tested integrations:

  • vLLM 0.10.2 with GuideLLM 0.7.3 on WSL2 and an RTX 3080.
  • vLLM 0.28.0 with GuideLLM 0.7.3 on the same system. WSL2 required VLLM_USE_V2_MODEL_RUNNER: "0" because UVA was unavailable and VLLM_USE_FLASHINFER_SAMPLER: "0" because the CUDA compiler toolkit was not installed. Native Linux systems may not require these settings.

Other combinations may work but are not yet verified.

The published py3-none-any wheel installs on Linux and Windows. Configuration validation and stored-result inspection work on Windows, but starting an experiment is supported only on Linux because vLLM has no native Windows runtime.

Quick start

Requirements: Linux, an NVIDIA GPU, a local model directory, and working vllm and guidellm commands. Install this checkout with:

pip install -e .

Create experiment.yaml:

schema_version: 1
experiment:
  name: first-run
model:
  path: /models/opt-125m
server:
  args:
    gpu-memory-utilization: 0.8
  tune:
    max-num-seqs:
      values: [8, 16]
benchmark:
  runs:
    - name: throughput
      profile:
        kind: throughput
        max_concurrency: 16
      constraints:
        - kind: max_requests
          count: 10
      data:
        - kind: synthetic_text
          prompt_tokens: 32
          output_tokens: 16
optimization:
  maximize: output_tokens_per_second
  sampler: tpe
  trials: 2

Run it:

vtune --config experiment.yaml

The short form is vtune -c experiment.yaml. The command validates the file, runs the experiment, persists results, and generates its exports and report. vTune binds vLLM to 127.0.0.1 by default. Set server.args.host explicitly only when the benchmark server must be reachable from another host.

Terminal output is concise by default. To stream vLLM and GuideLLM logs:

vtune --config experiment.yaml --verbose

The persistent equivalent uses GuideLLM's logging level names:

logging:
  level: DEBUG

Supported levels are DEBUG, INFO, WARNING, ERROR, and CRITICAL. Full per-trial log files are always saved. --verbose overrides the configured level with DEBUG for that invocation.

Retry one or more selected trials into a new immutable linked run:

vtune retry --run runs/EXPERIMENT/RUN_ID \
  --trial trial-0001 --trial trial-0004

The source run is never modified.

Display every stored vLLM and GuideLLM command for a trial without executing anything:

vtune reproduce --run runs/EXPERIMENT/RUN_ID --trial trial-0001

Each completed run also contains a self-contained report.html decision dashboard with the best observed configuration, baseline comparison, score history, throughput/latency tradeoff, and observed parameter effects.

Random and TPE runs never execute the same resolved configuration twice. optimization.trials cannot exceed the number of unique configurations in the declared search space.

Product documents

The MVP specification defines the first releasable version and its acceptance criteria. The roadmap describes capabilities that should be designed for now but implemented after the core experiment loop is reliable.

vTune is available under the MIT License.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

vtune-0.1.0a1.tar.gz (43.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

vtune-0.1.0a1-py3-none-any.whl (63.0 kB view details)

Uploaded Python 3

File details

Details for the file vtune-0.1.0a1.tar.gz.

File metadata

  • Download URL: vtune-0.1.0a1.tar.gz
  • Upload date:
  • Size: 43.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for vtune-0.1.0a1.tar.gz
Algorithm Hash digest
SHA256 5bde18fa98035f016d2a0bb5ffdd098d40f6f57d9443ded098e4fd49c4c29c55
MD5 fd7be0b722464eacae4a3d79f805ae86
BLAKE2b-256 8103e93be4c25d2d4bbab6884ea4c607314d5e6ca9cfc93f3690ea6d0f9b5582

See more details on using hashes here.

Provenance

The following attestation bundles were made for vtune-0.1.0a1.tar.gz:

Publisher: publish.yml on brtydse100/vTune

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file vtune-0.1.0a1-py3-none-any.whl.

File metadata

  • Download URL: vtune-0.1.0a1-py3-none-any.whl
  • Upload date:
  • Size: 63.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for vtune-0.1.0a1-py3-none-any.whl
Algorithm Hash digest
SHA256 1c9588855e564fa6d0587061034a549f0a64188e9647512f18bd28c2ea7625d8
MD5 f6c9a5115fb6e5a109d2bfbe6b5e99ab
BLAKE2b-256 db20d7f4c77738fcad22d357cecef33503c92e56fe0c329d42f414645b07d03d

See more details on using hashes here.

Provenance

The following attestation bundles were made for vtune-0.1.0a1-py3-none-any.whl:

Publisher: publish.yml on brtydse100/vTune

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page