qauvern — IBM Quantum Load Balancer
A Python CLI tool for optimizing quantum allocations across instances to maximize utilization of IBM Quantum resources.
The name comes from QAU (Quantum Allocation Unit) + govern — managing and governing those allocations.
Overview
qauvern helps administrators manage quantum computing allocations efficiently by:
- Analyzing current instance usage and allocations
- Identifying underutilized instances
- Recommending optimal allocation adjustments
- Automatically applying optimizations to maximize resource utilization
- Giving temporary boosts to an instance's limit.
Refer to How it works for more information on the algorithm.
Key Concepts
- Rolling Window: 28-day backward-looking usage period
- Fairness: Ratio of consumed time to allocated time. Lower fairness = higher priority
- Allocation: The target consumption for an instance during the rolling window. An instance can exceed its allocation, but its priority will decrease due to the fairness score.
- Limit: An optional hard cap on instance consumption.
- Net Grant: A bonus configured in the
qauvernconfig file to temporarily boost an instance's limit.net_grant_secondsis a lifetime budget. Multiple grants stack, and a grant keeps crediting the limit after it expires until the usage it funded rolls out of the 28-day rolling window — see How net grants expire.
Installation
There are two ways to install qauvern:
- Pex — a single self-contained file you can download and run directly. No virtual environment or dependency management needed; just Python 3.10+ on your system. Best if you want to get running quickly or avoid modifying your Python environment. (What is Pex?)
- pip — a standard install from source into a virtual environment. Best if you want to pin a version in a requirements file or integrate with an existing Python workflow.
Option 1: Pex (single-file executable)
Download qauvern.pex from the GitHub Releases page. The Pex file only requires Python 3.10+ on your macOS or Linux system — no pip or virtual environment needed.
chmod +x qauvern.pex
./qauvern.pex --help
You can also build the Pex yourself if you have the repo cloned and Just installed (see CONTRIBUTING.md):
just pex
./dist/qauvern.pex --help
Option 2: pip install
Install "qauvern" into a virtual environment:
python3 -m venv .venv
source .venv/bin/activate
pip install qauvern
The qauvern CLI is available while the virtual environment is active:
qauvern --help
Configuration
The program operates on YAML configuration files where you define your account, plan, and instances, such as:
# config.yaml
account_id: "your-ibm-cloud-account-id"
plan: "premium" # one of: internal, premium, paygo
# Minimum allocation to maintain for each instance.
# `qauvern configure` defaults to 60 seconds.
minimum_allocation_seconds: 60
# Hold back a percentage of account allocation from rebalancing (optional, default: 0)
# allocation_reserve_percent: 20
# Once account usage exceeds this percent of the account's allocation budget, stop
# requiring every instance's allocation to stay >= its 28-day usage (optional, default: 100,
# meaning always enforced).
# usage_floor_relax_above_percent: 90
instances:
- name: "Quantum Chemistry Research"
crn: "crn:v1:bluemix:public:quantum-computing:us-east:a/abc123:instance-1::"
# Optional: Hard limit applied on every optimize run
# limit_seconds: 216000
# Optional: Temporary time bonus above limit_seconds (requires limit_seconds).
# Applies while start_date <= today < end_date, with tz-aware ISO timestamps;
# end_date defaults to start_date + 28 days, and any length works.
# net_grant_seconds is a lifetime budget. Multiple grants stack, and a grant keeps crediting
# the limit until it rolls out of the 28-day window — see "How Net Grants Expire".
# net_grants:
# - start_date: "2026-05-01T00:00:00+00:00"
# net_grant_seconds: 360000 # 100 extra hours for a May sprint
# - start_date: "2026-06-15T00:00:00+00:00"
# end_date: "2026-07-31T00:00:00+00:00"
# net_grant_seconds: 180000
- name: "Quantum Machine Learning"
crn: "crn:v1:bluemix:public:quantum-computing:us-east:a/abc123:instance-2::"
limit_seconds: 72000
See examples/config-example.yaml for a complete example. Use qauvern configure to generate your initial file. Use qauvern update to automatically update the config file, such as adding new instances.
You should check in this file to version control.
Usage
Authentication
The tool uses IBM Cloud IAM for authentication. Set your IBM Cloud API key as an environment variable:
export IBMCLOUD_API_KEY="your-ibm-cloud-api-key"
Or pass it directly with the --api-key flag to any command.
Obtaining an IBM Cloud API Key
- Log in to IBM Cloud
- Go to Manage > Access (IAM) > API keys
- Click Create an IBM Cloud API key
- Give it a name and description
- Copy the API key (you won't be able to see it again)
Commands
Configure (Generate Configuration)
Generate a base configuration file from an existing IBM Cloud account:
qauvern configure --account-id your-account-id --plan premium --output config.yaml
This queries the IBM Quantum API to discover active instances for the given account and plan, then writes a YAML config file you can then edit.
Options:
--account-id, -a: IBM Cloud account ID (required)--plan, -p: Plan name —internal,premium, orpaygo(required)--api-key, -k: IBM Cloud API key (or useIBMCLOUD_API_KEYenv var)--region: Limit to instances in a specific region (e.g.,us-east,eu-de). Because all other commands only operate on instances in your config file, you can use this to restrictqauvernto a single region.--output, -o: Output file path (default:config.yaml)
After generating the configuration, optionally make these edits:
- Set
limit_secondsandnet_grantsper instance to control hard caps and temporary bonuses. - Change
minimum_allocation_secondsfrom its default of 60 seconds. - Set
allocation_reserve_percentfrom[0, 100)to hold back a buffer as a fraction of the total account budget (e.g.20caps total allocation at 80% of budget). - Lower
usage_floor_relax_above_percentbelow its default of100if your account's plan allows an instance to consume more than its allocation once account usage crosses that percent of the account budget (qauvern otherwise always refuses to let allocation drop below 28-day usage).
Use qauvern update to keep the file in sync as instances are added, removed, or renamed.
Update (Reconcile Configuration)
Reconcile an existing configuration file with the live IBM Quantum API:
qauvern update --config config.yaml
qauvern update --config config.yaml --dry-run # preview changes only
The update command asks for confirmation before making edits.
Whereas configure generates a fresh file from scratch, update is for ongoing maintenance of an existing config. It performs five reconciliation steps by default:
- Prune net_grants: drops
net_grantsentries that have fully rolled out of the 28-day window (end_datemore than 28 days ago). Grants that have ended but are still crediting the limit are kept and reported underNotes:with the date they become removable (see How net grants expire). - Remove instances: removes entries for archived or missing instances
- Fix names: updates instance names that have drifted from the live API
- Add instances: appends newly discovered instances
- Add missing limits: fills in
limit_secondsfrom the live API for instances that have a live limit but none configured. Existinglimit_secondsare never overwritten.
Comments and customizations (limit_seconds, allocation_reserve_percent, custom dates, etc.) are preserved because the file is rewritten in round-trip YAML mode.
Options:
--config, -c: Path to the config file (required)--api-key, -k: IBM Cloud API key (or useIBMCLOUD_API_KEYenv var)--region: Limit discovery to a specific region, likeus-eastoreu-de.--dry-run: Print planned changes without writing the file--yes, -y: Skip the confirmation prompt (for automation)--no-net-grants: Skip removing rolled-offnet_grants--no-add: Skip adding newly discovered instances--no-names: Skip fixing instance name drift--no-remove: Skip removing archived/missing instances--no-limits: Skip addinglimit_secondsfor instances that have a live limit but none in the config
Show Current Allocations
Display a summary of your account and instance allocations:
qauvern show --config config.yaml
Output includes:
- Account summary (total allocation, consumption, utilization)
- Instance details with fairness values
Instances (Non-Admin View)
Display instance usage summary, without requiring admin privileges:
qauvern instances --config config.yaml
Analyze Allocations
Analyze current allocations and show optimization recommendations, without making changes:
qauvern analyze --config config.yaml
This command identifies underutilized instances, calculates optimal reallocations, and shows what changes would be made.
Scope:
analyzeonly considers instances listed in your config file. Allocation held by unconfigured instances on the same account+plan is left untouched, but it is counted against the account cap so the recommendations never overcommit. The summary block reports it on theHeld by unconfigured instancesline.
Output formats
Use --format to choose how results are rendered:
table(default) — human-readable summary block plus an instance table, followed by aLIMIT BREAKDOWNsection for any instance whose limit is currently shaped by a net grant.csv— one row per configured instance, suitable for spreadsheets or quick pipelines. Account-level info is omitted because CSV is a flat row-based format.json— a structured payload for scripts. Includes account-level info, the reserve, validation errors, and per-instance rows with pre-computed allocation and limit deltas.
Both csv and json write data to stdout and logs to stderr, allowing you to pipe the stdout:
qauvern analyze --config config.yaml --format csv > analysis.csv
qauvern analyze --config config.yaml --format json > analysis.json
json writes validatoin errors to a validation_errors array, whereas csv logs the errors to stderr.
Common notes for both machine formats:
- All durations are raw integer seconds.
- The "new" allocation/limit value is always emitted, even if it is the same as the current value. Use the delta fields to quickly determine if there was a change, such as
limit_delta_secondswithjson. - The effective limit is broken into its terms (see How net grants expire) as
limit_breakdowninjsonand as the trailinglimit_base,limit_grant_funded,limit_unspent_grant,limit_overagecolumns incsv—null/blank when the config sets no limit.
Inspect the JSON schema with jq keys and jq '.instances[0] | keys' against a real run.
Previewing a future date
--preview-date YYYY-MM-DD evaluates net-grant activation and rolling-window boundaries as of that
date instead of today — e.g. to see how a limit will look once a scheduled grant starts, or after
it expires:
qauvern analyze --config config.yaml --preview-date 2026-12-01
This is not a usage forecast: consumed_* figures always reflect real, current usage. If the
preview date is far enough ahead that the 28-day rolling window no longer overlaps any real usage,
those figures will show near zero — that's the rolling window working as designed. The
effective date used is echoed back as preview_date in --format json and as a Preview date:
line in --format table; --format csv omits it.
Optimize Allocations
Apply optimization recommendations to update instance allocations:
qauvern optimize --config config.yaml
qauvern optimize --config config.yaml --dry-run # preview only
This command will:
- Determine if there are changes to any instance's limit from setting
limit_secondsandnet_grantsin the config file. - Calculate optimal allocations.
- Display proposed changes.
- Prompt for confirmation.
- Apply allocation and limit updates via API.
Use --dry-run to compute and display changes without applying them. Use --yes / -y to skip the confirmation prompt in automated pipelines.
Scope: Like
analyze,optimizeonly modifies instances listed in your config file. Unconfigured instances keep their existing allocation and limit; their allocation is reserved against the account cap when computing the new distribution. To bring an instance under management, runqauvern updateor manually add it to the config.
Create Instance
Provision a new IBM Quantum service instance:
qauvern create my-instance \
--target us-east \
--resource-group your-resource-group-id \
--plan premium \
--allocation 10h
Options:
NAME(positional, required): Name for the new instance--target, -t: Deployment region (required, e.g.,us-east,eu-de)--resource-group, -g: IBM Cloud resource group ID (required)--plan, -p: Plan name —internal,premium, orpaygo(required)--allocation, -a: Initial allocation (required, e.g.,96000,10h,2.5d)--limit, -l: Instance limit (e.g.,9600,10h)--tag: Tags to apply (repeatable)
Staging Environment
To target the IBM Quantum staging environment (test.cloud.ibm.com) instead of production, use the --staging flag or set the IBMCLOUD_STAGING environment variable:
qauvern --staging analyze --config config-staging.yaml
# or
export IBMCLOUD_STAGING=True
The --staging flag is a global option and applies to all commands.
How It Works
Optimization Algorithm
For each managed instance, qauvern:
- Resolves the effective limit from
limit_secondsand anynet_grantsin the config file, including recently expired grants that are still crediting (see How net grants expire). If the config file does not setlimit_secondsornet_grants,qauvernuses the live limit in IBM Quantum Platform, if any.qauvernwill apply this new effective limit and also use it as the upper bound on the instance's allocation. - Computes an activity score by exponentially weighting recent usage (24h carries 16× the weight of 28d). Instances with no usage across all buckets get score 0 and are classified inactive.
Then, account-wide:
- Pins every managed instance to its floor —
max(minimum_allocation_seconds, 28-day consumed). Inactive instances stay at the floor. - Builds a redistribution pool from unallocated headroom plus everything managed instances hold above their floor. If
allocation_reserve_percentis set, withholds a fixed fraction of the total account budget so total allocation never exceedsbudget × (1 − allocation_reserve_percent / 100). - Uses the water-fill algorithm to distribute the pool across active instances proportional to activity score. When an instance hits its effective limit, it drops out and its surplus flows to the rest. If every active instance is capped, leftover capacity stays unallocated rather than being forced onto any instance.
See Design.md for full algorithm details and the invariants the optimizer enforces.
How Net Grants Expire
A grant is active while start_date <= today < end_date, but it does not vanish from the effective limit the moment it ends. IBM Quantum measures usage over a 28-day rolling window, so the minutes a grant paid for stay counted against the instance for up to 28 days after the grant is over — an instance that spent its grant on the last day would otherwise find its whole base limit already consumed. qauvern keeps crediting the usage a grant funded until that usage itself rolls out, which guarantees:
An instance is never worse off after a grant expires than if the grant had never existed.
To uphold this invariant, qauvern spends grant time before base time: each day's usage is charged first to any active grant with remaining budget, and only the shortfall is charged to the base limit. For example, with limit_seconds: 10 and a grant of 100 spent on the grant's last day:
| Usage during the grant | Paid by grant | Paid by base | Effective limit at expiry | Available |
|---|---|---|---|---|
| 40 (grant under-spent) | 40 | 0 | 10 + 40 = 50 | 10 — the whole base limit |
| 100 (the grant, exactly) | 100 | 0 | 10 + 100 = 110 | 10 — the whole base limit |
| 105 (5 past the grant) | 100 | 5 | 10 + 100 = 110 | 5 |
| 110 (grant + base) | 100 | 10 | 10 + 100 = 110 | 0 — spent, but not in debt |
| 120 (10 past both) | 100 | 20 | 10 + 100 = 110 | −10 — the grant credits at most 100 |
A grant is a separate pot: spending it never eats the base limit, and you keep only what you drew from it — hence the first row's unspent 60 is dropped at end_date (limit 50, not 110).
The credit then decays as the funded days age out. For example, with limit_seconds: 10 and a grant of 100 covering March 1–11, with 70s used on March 5 and 50s used on March 8:
| Date | Grant status | Effective limit | Usage in window | Available |
|---|---|---|---|---|
| Mar 10 | active | 110 | 120 | −10 |
| Mar 11 | expired, still crediting | 110 | 120 | −10 |
| Apr 2 | expired, still crediting | 110 | 120 | −10 |
| Apr 3 | Mar 5 rolled out | 40 | 50 | −10 |
| Apr 6 | Mar 8 rolled out | 10 | 0 | 10 |
| Apr 8 | fully rolled off | 10 | 0 | 10 |
The credit decays the same way even when the base limit only absorbs a little of the overage. For example, imagine the same limit_seconds: 10 and 100 grant covering March 1–11. This time, only 105s is used, all on March 5: the grant pays the first 100, and the base pays the remaining 5.
| Date | Grant status | Effective limit | Usage in window | Available |
|---|---|---|---|---|
| Mar 5 | active | 110 | 105 | 5 |
| Mar 11 | expired, still crediting | 110 | 105 | 5 |
| Apr 2 | expired, still crediting | 110 | 105 | 5 |
| Apr 3 | Mar 5 rolled out | 10 | 0 | 10 |
Once Mar 5 rolls out of the 28-day window, the credit (100) and the usage it funded (105) roll out together, so availability returns to the full base limit of 10 rather than staying negative.
The budget is a lifetime budget
net_grant_seconds covers the grant's whole period. Meanwhile, limit_seconds refreshes as usage rolls out of the 28-day rolling window. A grant is therefore worth its budget once, plus the base limit per rolling window.
For example, take limit_seconds: 100 and a grant of net_grant_seconds: 1000 running March 1 to May 1 (61 days — longer than the 28-day window). Say the instance draws 600 seconds of usage in a single day, March 5, and nothing else for the rest of the grant:
| Date | What's happening | Grant-funded (in window) | Unspent grant | Effective limit | Available |
|---|---|---|---|---|---|
| Mar 5 | the 600 is drawn and charged to the grant | 600 | 400 | 100 + 600 + 400 = 1100 | 500 |
| Apr 2 | Mar 5 is still inside the 28-day window | 600 | 400 | 1100 | 500 |
| Apr 3 | Mar 5 rolls out of the window | 0 | 400 | 100 + 0 + 400 = 500 | 500 |
May 1 (end_date) |
grant ends; the undrawn 400 is forfeit | 0 | 0 | 100 | 100 |
Two separate numbers move here: grant-funded tracks usage inside today's 28-day window and decays as that usage ages out (0 by Apr 3, since the only usage was Mar 5); unspent grant tracks budget not yet drawn and only changes when the instance spends it or the grant ends. Availability holds steady at 500 across the Apr 3 rolloff because the two exactly offset — the grant-funded usage that ages out is matched by the unspent budget still sitting unclaimed. It only drops, to the base 100, when the grant ends on May 1 and that unspent 400 is forfeit.
Configured vs. Unconfigured Instances
analyze, optimize, show, and instances only operate on instances listed in your config file. Any other instance that exists on the same account and plan is unconfigured and is left exactly as-is — its allocation and limit are never touched.
Unconfigured instances still consume from the account-wide cap, so the optimizer subtracts their allocation before deciding how much to redistribute. Concretely:
raw_pool = account.allocation_budget
− sum(unconfigured allocations) ← reserved, untouched
− sum(floors of configured) ← max(minimum_allocation_seconds, 28-day usage)
reserve_amount = account.allocation_budget × (allocation_reserve_percent / 100)
redistributable = max(0, raw_pool − reserve_amount)
This means you can safely manage a subset of an account's instances with qauvern: anything you leave out of the config file is opaque to the optimizer except as a fixed reservation. To bring an instance under management, add it to the config (or regenerate with qauvern configure).
Caveat: the configured instances will absorb all the available account allocation, which leaves no available allocation for the unconfigured instances. For example, if you configure 2 of 10 instances, those 2 will claim every spare second on the account and the remaining 8 are left with no buffer to expand into. Use allocation_reserve_percent to hold back a fraction of the total account budget if you need headroom for unconfigured instances.
Examples
Basic Workflow
# 1. Generate initial configuration from your account
qauvern configure --account-id your-account-id --plan premium --output config.yaml
# 2. Edit config.yaml to customize limits, net_grants, and other config like `allocation_reserve_percent`.
# 3. Check current status
qauvern show --config config.yaml
# 4. Analyze and see recommendations
qauvern analyze --config config.yaml
# 5. Apply optimizations
qauvern optimize --config config.yaml
After the initial setup, run qauvern update --config config.yaml periodically to keep the config in sync with the live API (new instances, renames, archived instances, fully rolled-off net_grants).
Automations
qauvern optimize --yes skips the confirmation prompt, and qauvern analyze --format json emits a structured report — together they make the tool easy to drive from cron, CI, or a monitoring script.
Run the optimizer on a schedule to keep allocations balanced:
# Add to crontab for weekly optimization
0 0 * * 0 qauvern optimize --config /path/to/config.yaml --yes
Capture a periodic analysis snapshot for dashboards or alerting:
qauvern analyze --config /path/to/config.yaml --format json > /var/log/qauvern/$(date +%Y-%m-%d).json
For example, surface instances the optimizer wants to shrink:
qauvern analyze --config config.yaml --format json \
| jq '.instances[] | select(.allocation_delta_seconds < 0) | {name, allocation_delta_seconds, allocation_change_reason}'
Metadata
Release files for qauvern 0.13.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| qauvern-0.13.0.tar.gz | 141.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| qauvern-0.13.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 195.4 kB
Release files / qauvern-0.13.0.tar.gz
| Download URL | qauvern-0.13.0.tar.gz |
|---|---|
| Size | 141.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
44d3c22a18e3e68c5814b31f31a1bf3685548e0c792187c8c78cd29d56cca677
|
|
BLAKE2b-256 checksum How to use checksums |
d9e893007e4f3dd4af48e9447cbd1d6f9350887c12b39182c5b2ab6e7edaf52b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 10, 2026.
Transparency logRelease files / qauvern-0.13.0-py3-none-any.whl
| Download URL | qauvern-0.13.0-py3-none-any.whl |
|---|---|
| Size | 54.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
425da6d31821c1714211af35b46637ad01fa28707b59c84d9f006873a07b6219
|
|
BLAKE2b-256 checksum How to use checksums |
7e4c4b1fc08a762bdd0ab2bbbf0dd6451555b068d894c91171ed9143336c064f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 10, 2026.
Transparency log