Skip to main content

qauvern — IBM Quantum Load Balancer

A Python CLI tool for optimizing quantum allocations across instances to maximize utilization of IBM Quantum resources.

The name comes from QAU (Quantum Allocation Unit) + govern — managing and governing those allocations.

Overview

qauvern helps administrators manage quantum computing allocations efficiently by:

  • Analyzing current instance usage and allocations
  • Identifying underutilized instances
  • Recommending optimal allocation adjustments
  • Automatically applying optimizations to maximize resource utilization
  • Giving temporary boosts to an instance's limit.

Refer to How it works for more information on the algorithm.

Key Concepts

  • Rolling Window: 28-day backward-looking usage period
  • Fairness: Ratio of consumed time to allocated time. Lower fairness = higher priority
  • Allocation: The target consumption for an instance during the rolling window. An instance can exceed its allocation, but its priority will decrease due to the fairness score.
  • Limit: An optional hard cap on instance consumption.
  • Net Grant: A bonus configured in the qauvern config file to temporarily boost an instance's limit. net_grant_seconds is a lifetime budget. Multiple grants stack, and a grant keeps crediting the limit after it expires until the usage it funded rolls out of the 28-day rolling window — see How net grants expire.

Installation

There are two ways to install qauvern:

  • Pex — a single self-contained file you can download and run directly. No virtual environment or dependency management needed; just Python 3.10+ on your system. Best if you want to get running quickly or avoid modifying your Python environment. (What is Pex?)
  • pip — a standard install from source into a virtual environment. Best if you want to pin a version in a requirements file or integrate with an existing Python workflow.

Option 1: Pex (single-file executable)

Download qauvern.pex from the GitHub Releases page. The Pex file only requires Python 3.10+ on your macOS or Linux system — no pip or virtual environment needed.

chmod +x qauvern.pex
./qauvern.pex --help

You can also build the Pex yourself if you have the repo cloned and Just installed (see CONTRIBUTING.md):

just pex
./dist/qauvern.pex --help

Option 2: pip install

Install "qauvern" into a virtual environment:

python3 -m venv .venv
source .venv/bin/activate
pip install qauvern

The qauvern CLI is available while the virtual environment is active:

qauvern --help

Configuration

The program operates on YAML configuration files where you define your account, plan, and instances, such as:

# config.yaml
account_id: "your-ibm-cloud-account-id"
plan: "premium"  # one of: internal, premium, paygo

# Minimum allocation to maintain for each instance.
# `qauvern configure` defaults to 60 seconds.
minimum_allocation_seconds: 60

# Hold back a percentage of account allocation from rebalancing (optional, default: 0)
# allocation_reserve_percent: 20

# Once account usage exceeds this percent of the account's allocation budget, stop
# requiring every instance's allocation to stay >= its 28-day usage (optional, default: 100,
# meaning always enforced).
# usage_floor_relax_above_percent: 90

instances:
  - name: "Quantum Chemistry Research"
    crn: "crn:v1:bluemix:public:quantum-computing:us-east:a/abc123:instance-1::"

    # Optional: Hard limit applied on every optimize run
    # limit_seconds: 216000

    # Optional: Temporary time bonus above limit_seconds (requires limit_seconds).
    # Applies while start_date <= today < end_date, with tz-aware ISO timestamps;
    # end_date defaults to start_date + 28 days, and any length works.
    # net_grant_seconds is a lifetime budget. Multiple grants stack, and a grant keeps crediting
    # the limit until it rolls out of the 28-day window — see "How Net Grants Expire".
    # net_grants:
    #   - start_date: "2026-05-01T00:00:00+00:00"
    #     net_grant_seconds: 360000  # 100 extra hours for a May sprint
    #   - start_date: "2026-06-15T00:00:00+00:00"
    #     end_date: "2026-07-31T00:00:00+00:00"
    #     net_grant_seconds: 180000

  - name: "Quantum Machine Learning"
    crn: "crn:v1:bluemix:public:quantum-computing:us-east:a/abc123:instance-2::"
    limit_seconds: 72000

See examples/config-example.yaml for a complete example. Use qauvern configure to generate your initial file. Use qauvern update to automatically update the config file, such as adding new instances.

You should check in this file to version control.

Usage

Authentication

The tool uses IBM Cloud IAM for authentication. Set your IBM Cloud API key as an environment variable:

export IBMCLOUD_API_KEY="your-ibm-cloud-api-key"

Or pass it directly with the --api-key flag to any command.

Obtaining an IBM Cloud API Key

  1. Log in to IBM Cloud
  2. Go to Manage > Access (IAM) > API keys
  3. Click Create an IBM Cloud API key
  4. Give it a name and description
  5. Copy the API key (you won't be able to see it again)

Commands

Configure (Generate Configuration)

Generate a base configuration file from an existing IBM Cloud account:

qauvern configure --account-id your-account-id --plan premium --output config.yaml

This queries the IBM Quantum API to discover active instances for the given account and plan, then writes a YAML config file you can then edit.

Options:

  • --account-id, -a: IBM Cloud account ID (required)
  • --plan, -p: Plan name — internal, premium, or paygo (required)
  • --api-key, -k: IBM Cloud API key (or use IBMCLOUD_API_KEY env var)
  • --region: Limit to instances in a specific region (e.g., us-east, eu-de). Because all other commands only operate on instances in your config file, you can use this to restrict qauvern to a single region.
  • --output, -o: Output file path (default: config.yaml)

After generating the configuration, optionally make these edits:

  • Set limit_seconds and net_grants per instance to control hard caps and temporary bonuses.
  • Change minimum_allocation_seconds from its default of 60 seconds.
  • Set allocation_reserve_percent from [0, 100) to hold back a buffer as a fraction of the total account budget (e.g. 20 caps total allocation at 80% of budget).
  • Lower usage_floor_relax_above_percent below its default of 100 if your account's plan allows an instance to consume more than its allocation once account usage crosses that percent of the account budget (qauvern otherwise always refuses to let allocation drop below 28-day usage).

Use qauvern update to keep the file in sync as instances are added, removed, or renamed.

Update (Reconcile Configuration)

Reconcile an existing configuration file with the live IBM Quantum API:

qauvern update --config config.yaml
qauvern update --config config.yaml --dry-run   # preview changes only

The update command asks for confirmation before making edits.

Whereas configure generates a fresh file from scratch, update is for ongoing maintenance of an existing config. It performs five reconciliation steps by default:

  • Prune net_grants: drops net_grants entries that have fully rolled out of the 28-day window (end_date more than 28 days ago). Grants that have ended but are still crediting the limit are kept and reported under Notes: with the date they become removable (see How net grants expire).
  • Remove instances: removes entries for archived or missing instances
  • Fix names: updates instance names that have drifted from the live API
  • Add instances: appends newly discovered instances
  • Add missing limits: fills in limit_seconds from the live API for instances that have a live limit but none configured. Existing limit_seconds are never overwritten.

Comments and customizations (limit_seconds, allocation_reserve_percent, custom dates, etc.) are preserved because the file is rewritten in round-trip YAML mode.

Options:

  • --config, -c: Path to the config file (required)
  • --api-key, -k: IBM Cloud API key (or use IBMCLOUD_API_KEY env var)
  • --region: Limit discovery to a specific region, like us-east or eu-de.
  • --dry-run: Print planned changes without writing the file
  • --yes, -y: Skip the confirmation prompt (for automation)
  • --no-net-grants: Skip removing rolled-off net_grants
  • --no-add: Skip adding newly discovered instances
  • --no-names: Skip fixing instance name drift
  • --no-remove: Skip removing archived/missing instances
  • --no-limits: Skip adding limit_seconds for instances that have a live limit but none in the config

Show Current Allocations

Display a summary of your account and instance allocations:

qauvern show --config config.yaml

Output includes:

  • Account summary (total allocation, consumption, utilization)
  • Instance details with fairness values

Instances (Non-Admin View)

Display instance usage summary, without requiring admin privileges:

qauvern instances --config config.yaml

Analyze Allocations

Analyze current allocations and show optimization recommendations, without making changes:

qauvern analyze --config config.yaml

This command identifies underutilized instances, calculates optimal reallocations, and shows what changes would be made.

Scope: analyze only considers instances listed in your config file. Allocation held by unconfigured instances on the same account+plan is left untouched, but it is counted against the account cap so the recommendations never overcommit. The summary block reports it on the Held by unconfigured instances line.

Output formats

Use --format to choose how results are rendered:

  • table (default) — human-readable summary block plus an instance table, followed by a LIMIT BREAKDOWN section for any instance whose limit is currently shaped by a net grant.
  • csv — one row per configured instance, suitable for spreadsheets or quick pipelines. Account-level info is omitted because CSV is a flat row-based format.
  • json — a structured payload for scripts. Includes account-level info, the reserve, validation errors, and per-instance rows with pre-computed allocation and limit deltas.

Both csv and json write data to stdout and logs to stderr, allowing you to pipe the stdout:

qauvern analyze --config config.yaml --format csv  > analysis.csv
qauvern analyze --config config.yaml --format json > analysis.json

json writes validatoin errors to a validation_errors array, whereas csv logs the errors to stderr.

Common notes for both machine formats:

  • All durations are raw integer seconds.
  • The "new" allocation/limit value is always emitted, even if it is the same as the current value. Use the delta fields to quickly determine if there was a change, such as limit_delta_seconds with json.
  • The effective limit is broken into its terms (see How net grants expire) as limit_breakdown in json and as the trailing limit_base, limit_grant_funded, limit_unspent_grant, limit_overage columns in csv — null/blank when the config sets no limit.

Inspect the JSON schema with jq keys and jq '.instances[0] | keys' against a real run.

Previewing a future date

--preview-date YYYY-MM-DD evaluates net-grant activation and rolling-window boundaries as of that date instead of today — e.g. to see how a limit will look once a scheduled grant starts, or after it expires:

qauvern analyze --config config.yaml --preview-date 2026-12-01

This is not a usage forecast: consumed_* figures always reflect real, current usage. If the preview date is far enough ahead that the 28-day rolling window no longer overlaps any real usage, those figures will show near zero — that's the rolling window working as designed. The effective date used is echoed back as preview_date in --format json and as a Preview date: line in --format table; --format csv omits it.

Optimize Allocations

Apply optimization recommendations to update instance allocations:

qauvern optimize --config config.yaml
qauvern optimize --config config.yaml --dry-run   # preview only

This command will:

  1. Determine if there are changes to any instance's limit from setting limit_seconds and net_grants in the config file.
  2. Calculate optimal allocations.
  3. Display proposed changes.
  4. Prompt for confirmation.
  5. Apply allocation and limit updates via API.

Use --dry-run to compute and display changes without applying them. Use --yes / -y to skip the confirmation prompt in automated pipelines.

Scope: Like analyze, optimize only modifies instances listed in your config file. Unconfigured instances keep their existing allocation and limit; their allocation is reserved against the account cap when computing the new distribution. To bring an instance under management, run qauvern update or manually add it to the config.

Create Instance

Provision a new IBM Quantum service instance:

qauvern create my-instance \
  --target us-east \
  --resource-group your-resource-group-id \
  --plan premium \
  --allocation 10h

Options:

  • NAME (positional, required): Name for the new instance
  • --target, -t: Deployment region (required, e.g., us-east, eu-de)
  • --resource-group, -g: IBM Cloud resource group ID (required)
  • --plan, -p: Plan name — internal, premium, or paygo (required)
  • --allocation, -a: Initial allocation (required, e.g., 96000, 10h, 2.5d)
  • --limit, -l: Instance limit (e.g., 9600, 10h)
  • --tag: Tags to apply (repeatable)

Staging Environment

To target the IBM Quantum staging environment (test.cloud.ibm.com) instead of production, use the --staging flag or set the IBMCLOUD_STAGING environment variable:

qauvern --staging analyze --config config-staging.yaml
# or
export IBMCLOUD_STAGING=True

The --staging flag is a global option and applies to all commands.

How It Works

Optimization Algorithm

For each managed instance, qauvern:

  1. Resolves the effective limit from limit_seconds and any net_grants in the config file, including recently expired grants that are still crediting (see How net grants expire). If the config file does not set limit_seconds or net_grants, qauvern uses the live limit in IBM Quantum Platform, if any. qauvern will apply this new effective limit and also use it as the upper bound on the instance's allocation.
  2. Computes an activity score by exponentially weighting recent usage (24h carries 16× the weight of 28d). Instances with no usage across all buckets get score 0 and are classified inactive.

Then, account-wide:

  1. Pins every managed instance to its floor — max(minimum_allocation_seconds, 28-day consumed). Inactive instances stay at the floor.
  2. Builds a redistribution pool from unallocated headroom plus everything managed instances hold above their floor. If allocation_reserve_percent is set, withholds a fixed fraction of the total account budget so total allocation never exceeds budget × (1 − allocation_reserve_percent / 100).
  3. Uses the water-fill algorithm to distribute the pool across active instances proportional to activity score. When an instance hits its effective limit, it drops out and its surplus flows to the rest. If every active instance is capped, leftover capacity stays unallocated rather than being forced onto any instance.

See Design.md for full algorithm details and the invariants the optimizer enforces.

How Net Grants Expire

A grant is active while start_date <= today < end_date, but it does not vanish from the effective limit the moment it ends. IBM Quantum measures usage over a 28-day rolling window, so the minutes a grant paid for stay counted against the instance for up to 28 days after the grant is over — an instance that spent its grant on the last day would otherwise find its whole base limit already consumed. qauvern keeps crediting the usage a grant funded until that usage itself rolls out, which guarantees:

An instance is never worse off after a grant expires than if the grant had never existed.

To uphold this invariant, qauvern spends grant time before base time: each day's usage is charged first to any active grant with remaining budget, and only the shortfall is charged to the base limit. For example, with limit_seconds: 10 and a grant of 100 spent on the grant's last day:

Usage during the grant Paid by grant Paid by base Effective limit at expiry Available
40 (grant under-spent) 40 0 10 + 40 = 50 10 — the whole base limit
100 (the grant, exactly) 100 0 10 + 100 = 110 10 — the whole base limit
105 (5 past the grant) 100 5 10 + 100 = 110 5
110 (grant + base) 100 10 10 + 100 = 110 0 — spent, but not in debt
120 (10 past both) 100 20 10 + 100 = 110 −10 — the grant credits at most 100

A grant is a separate pot: spending it never eats the base limit, and you keep only what you drew from it — hence the first row's unspent 60 is dropped at end_date (limit 50, not 110).

The credit then decays as the funded days age out. For example, with limit_seconds: 10 and a grant of 100 covering March 1–11, with 70s used on March 5 and 50s used on March 8:

Date Grant status Effective limit Usage in window Available
Mar 10 active 110 120 −10
Mar 11 expired, still crediting 110 120 −10
Apr 2 expired, still crediting 110 120 −10
Apr 3 Mar 5 rolled out 40 50 −10
Apr 6 Mar 8 rolled out 10 0 10
Apr 8 fully rolled off 10 0 10

The credit decays the same way even when the base limit only absorbs a little of the overage. For example, imagine the same limit_seconds: 10 and 100 grant covering March 1–11. This time, only 105s is used, all on March 5: the grant pays the first 100, and the base pays the remaining 5.

Date Grant status Effective limit Usage in window Available
Mar 5 active 110 105 5
Mar 11 expired, still crediting 110 105 5
Apr 2 expired, still crediting 110 105 5
Apr 3 Mar 5 rolled out 10 0 10

Once Mar 5 rolls out of the 28-day window, the credit (100) and the usage it funded (105) roll out together, so availability returns to the full base limit of 10 rather than staying negative.

The budget is a lifetime budget

net_grant_seconds covers the grant's whole period. Meanwhile, limit_seconds refreshes as usage rolls out of the 28-day rolling window. A grant is therefore worth its budget once, plus the base limit per rolling window.

For example, take limit_seconds: 100 and a grant of net_grant_seconds: 1000 running March 1 to May 1 (61 days — longer than the 28-day window). Say the instance draws 600 seconds of usage in a single day, March 5, and nothing else for the rest of the grant:

Date What's happening Grant-funded (in window) Unspent grant Effective limit Available
Mar 5 the 600 is drawn and charged to the grant 600 400 100 + 600 + 400 = 1100 500
Apr 2 Mar 5 is still inside the 28-day window 600 400 1100 500
Apr 3 Mar 5 rolls out of the window 0 400 100 + 0 + 400 = 500 500
May 1 (end_date) grant ends; the undrawn 400 is forfeit 0 0 100 100

Two separate numbers move here: grant-funded tracks usage inside today's 28-day window and decays as that usage ages out (0 by Apr 3, since the only usage was Mar 5); unspent grant tracks budget not yet drawn and only changes when the instance spends it or the grant ends. Availability holds steady at 500 across the Apr 3 rolloff because the two exactly offset — the grant-funded usage that ages out is matched by the unspent budget still sitting unclaimed. It only drops, to the base 100, when the grant ends on May 1 and that unspent 400 is forfeit.

Configured vs. Unconfigured Instances

analyze, optimize, show, and instances only operate on instances listed in your config file. Any other instance that exists on the same account and plan is unconfigured and is left exactly as-is — its allocation and limit are never touched.

Unconfigured instances still consume from the account-wide cap, so the optimizer subtracts their allocation before deciding how much to redistribute. Concretely:

raw_pool        = account.allocation_budget
                  − sum(unconfigured allocations)   ← reserved, untouched
                  − sum(floors of configured)       ← max(minimum_allocation_seconds, 28-day usage)

reserve_amount  = account.allocation_budget × (allocation_reserve_percent / 100)
redistributable = max(0, raw_pool − reserve_amount)

This means you can safely manage a subset of an account's instances with qauvern: anything you leave out of the config file is opaque to the optimizer except as a fixed reservation. To bring an instance under management, add it to the config (or regenerate with qauvern configure).

Caveat: the configured instances will absorb all the available account allocation, which leaves no available allocation for the unconfigured instances. For example, if you configure 2 of 10 instances, those 2 will claim every spare second on the account and the remaining 8 are left with no buffer to expand into. Use allocation_reserve_percent to hold back a fraction of the total account budget if you need headroom for unconfigured instances.

Examples

Basic Workflow

# 1. Generate initial configuration from your account
qauvern configure --account-id your-account-id --plan premium --output config.yaml

# 2. Edit config.yaml to customize limits, net_grants, and other config like `allocation_reserve_percent`.

# 3. Check current status
qauvern show --config config.yaml

# 4. Analyze and see recommendations
qauvern analyze --config config.yaml

# 5. Apply optimizations
qauvern optimize --config config.yaml

After the initial setup, run qauvern update --config config.yaml periodically to keep the config in sync with the live API (new instances, renames, archived instances, fully rolled-off net_grants).

Automations

qauvern optimize --yes skips the confirmation prompt, and qauvern analyze --format json emits a structured report — together they make the tool easy to drive from cron, CI, or a monitoring script.

Run the optimizer on a schedule to keep allocations balanced:

# Add to crontab for weekly optimization
0 0 * * 0 qauvern optimize --config /path/to/config.yaml --yes

Capture a periodic analysis snapshot for dashboards or alerting:

qauvern analyze --config /path/to/config.yaml --format json > /var/log/qauvern/$(date +%Y-%m-%d).json

For example, surface instances the optimizer wants to shrink:

qauvern analyze --config config.yaml --format json \
  | jq '.instances[] | select(.allocation_delta_seconds < 0) | {name, allocation_delta_seconds, allocation_change_reason}'

Metadata

Release files for qauvern 0.13.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for qauvern 0.13.0
File Size Uploaded
qauvern-0.13.0.tar.gz 141.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for qauvern 0.13.0
File Interpreter ABI Platform
qauvern-0.13.0-py3-none-any.whl Python 3 none any Details

Total release size: 195.4 kB

Release files / qauvern-0.13.0.tar.gz

Download URL qauvern-0.13.0.tar.gz
Size 141.3 kB
Tags Source
SHA-256 checksum
How to use checksums
44d3c22a18e3e68c5814b31f31a1bf3685548e0c792187c8c78cd29d56cca677
BLAKE2b-256 checksum
How to use checksums
d9e893007e4f3dd4af48e9447cbd1d6f9350887c12b39182c5b2ab6e7edaf52b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 10, 2026.

Transparency log

Release files / qauvern-0.13.0-py3-none-any.whl

Download URL qauvern-0.13.0-py3-none-any.whl
Size 54.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
425da6d31821c1714211af35b46637ad01fa28707b59c84d9f006873a07b6219
BLAKE2b-256 checksum
How to use checksums
7e4c4b1fc08a762bdd0ab2bbbf0dd6451555b068d894c91171ed9143336c064f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 10, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.13.0 This release

2 release files

0.12.0

2 release files

0.10.0

2 release files

0.9.0

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page