Skip to main content

urun CLI

Deploy Python apps to urun from your terminal.

PyPI Python

Install

uv tool install urun-cli
# or
pip install urun-cli

The package installs the urun command:

urun --version

For one-off uvx usage:

uvx --from urun-cli urun --version
# or the package-matching command alias
uvx urun-cli --version

Quick start

Save an org-scoped deploy API key locally with urun login:

urun login

If URUN_API_KEY is not already set, urun login opens the urun console API keys page, prompts you to paste the generated urun_sk_... key, verifies it with the urun API, and stores credentials for later commands.

If you already have a key, you can pass it directly:

urun login --api-key urun_sk_<secret>

For CI or one-off commands, you can still use the environment variable:

export URUN_API_KEY=urun_sk_<secret>

Create app.py:

import urun
from urun import App

app = App("hello-h100")


@app.function(gpus="h100:1")
def hello(ctx: urun.Context):
    print(f"running on {ctx.device}")
    return {"device": str(ctx.device)}

Run it:

urun run app.py

In this release, urun run uses the same deploy pipeline as urun deploy. deploy remains available as the lower-level command while the full deploy/run/monitor workflow is being built.

Launch an official browser demo

For official urun-examples apps, the shortest browser stream path is:

cd ~/workspace/urun-examples
urun deploy prompt_canary/backend/app.py --name prompt-canary --timeout 900
urun demo prompt-canary --examples-root ~/workspace/urun-examples

The positional argument is the deployed app slug (as shown by urun list apps). Nothing is declared in a manifest — every value is discovered:

value where it comes from
frontend directory the one <example>/frontend under the examples root whose name matches the app slug (hyphens and underscores are equivalent, and a --dev suffix is ignored)
pnpm package that frontend's own package.json "name"
local port the --port N in that package's scripts.dev, else 3000
function read from the deployed app itself

urun demo uses your stored API key server-side to request a short-lived, function-scoped, origin-pinned browser JWT, sets the frontend's NEXT_PUBLIC_* environment variables, starts the local Next.js frontend, and opens it in your browser. The API key itself never reaches the page.

Discovery is never a guess: if no example directory matches the app — or if more than one does — the command fails and names what it found. Point it at the right directory yourself with --frontend-dir (relative to the examples root):

urun demo test-matrix-game --frontend-dir matrix-game-3/frontend

There is no curated list of demoable examples

urun demo will launch any <example>/frontend it can match, and it does not know which examples actually work. Earlier releases read a manifest that carried a supported flag with a reason, and refused the examples marked unsupported; that allowlist is gone along with the manifest.

The practical consequence: a frontend that is a stub, half-built, or known broken now starts a dev server and opens a browser instead of failing up front, and it fails on its own terms once loaded. notebook-playground is the clearest case — its frontend/ is a deliberate loud-failing stub. If a demo comes up wrong, check the example's own README before suspecting the CLI.

This was an accepted trade for deleting the manifest: which examples work is not derivable from the filesystem or the control plane, so the CLI cannot know it without a second declared source that would drift exactly as the manifest did.

Override the examples checkout with --examples-root or URUN_EXAMPLES_ROOT; override the session API with --session-url or URUN_SESSION_BASE_URL; pin the function with --function and the environment with --environment.

Serve a model from the catalog

urun serve deploys a model straight from the urun model catalog — no app code required. It is urun deploy with a serve app templated from the resolved catalog row (engine, HF repo, GPU, engine args).

List the enumerable model matrix (models, variants, the GPU each fits, engine shape, and whether the placement is dev-testable or prod-only):

urun serve catalog
urun serve catalog --json

Serve a model. With no variant the model's default (first) variant is used; with no --gpu the variant's first placement is used:

urun serve qwen-coder                    # default variant + first placement
urun serve glm-5.2:UD-IQ2_M              # explicit <id>:<variant>
urun serve qwen-coder-480b:fp8 --gpu b200:8

The catalog is anon-readable reference data published by a urun-infra migration. Reads go over PostgREST on the same control-plane host as the API URL; supply the public Supabase anon key via --anon-key or URUN_CATALOG_ANON_KEY (URUN_SUPABASE_ANON_KEY is also honored). Override the PostgREST base with --catalog-url / URUN_CATALOG_URL if needed.

By default urun serve <id> renders a templated serve app and deploys it. If you maintain your own _serve app, point --entrypoint at it; the CLI deploys that app with URUN_SERVE_CONFIG and URUN_SERVE_CATALOG_ROW set to the resolved row instead of rendering one.

Inspect apps

List every app deployed in your org and its current status:

urun list apps

Sample output:

APP                    FUNCTION        COMPUTE  STATUS         RELEASE       DETAIL
causal-forcing-stream  generate_video  h100:1   ready          fa61f31b0961  0/1 GPU units in use
queued-app             warmup          a10:1    provisioning   000000000000  building

The STATUS column is one of provisioning, ready, pending, paused, or failed. It is derived from three backend signals reporting on sequential lifecycle phases:

Build (S3 status) Promotion (app_deployments) Capacity (function_ready) STATUS
queued | building (no row yet) - provisioning
failed (no row yet) - failed
ready active false pending
ready active true ready
ready paused (irrelevant) paused
ready failed (irrelevant) failed

The DETAIL column carries the disambiguating signal (raw build state, error message, ready_reason, or in-use GPU counts). Pass --json for the raw payload.

This command is experimental and requires the server-side GET /apps endpoint, which is in development.

View logs

Use urun logs for build logs or runtime logs scoped to an app, a function within an app, or a session:

urun list logs
urun logs build <build-id>
urun logs app <app>
urun logs function <app> <function>
urun logs session <session>

Runtime commands accept --follow/-f for live output and --json for JSON (one-shot) or NDJSON (followed) output. Use urun list logs without filters to list all accessible logs, or add the same app, function, session, level, text, and time filters. urun tail logs remains available for the legacy follow workflow. The current logs API does not filter by environment.

Inspect sessions

List live and historical sessions in your org. Newest sessions appear at the bottom of the table so the command works well with tail:

urun list sessions
urun list sessions | tail -20
urun list sessions --state failed
urun list sessions --limit 500

Sample output:

ID            APP                FUNCTION        SHAPE    STARTED               DURATION  STATE       DETAIL
2cc8a91f4b3d  helios             world_gen       h100:4   2026-06-02 10:55 UTC       44s  failed      no_capacity
3f0017daee01  helios             world_gen       h100:1   2026-06-02 11:08 UTC    18m43s  completed   client_disconnect
4a1c886e2d0a  causal-forcing-…   generate_video  h100:1   2026-06-02 14:21 UTC     3m12s  live        -

The STATE column maps the raw backend status to a user-friendly label:

Backend status STATE
allocated starting
connected live
closed completed
failed failed
cancelled cancelled

DURATION is computed from allocated_at to closed_at for terminal sessions, or allocated_at to now for live ones. DETAIL carries close_reason when present. Pass --json for the raw payload (full IDs, ISO timestamps, all fields).

Pass --limit to control how many rows are fetched (default 100).

This command is experimental and requires the server-side GET /sessions endpoint, which is in development.

Inspect active compute

List the compute slices your org currently has provisioned:

urun list compute

Sample output:

APP                    FUNCTION        SHAPE   INSTANCES  GPU UNITS  SESSIONS  AGE
causal-forcing-stream  generate_video  h100:1  1/2        1/2        1         12s
helios                 world_gen       h100:4  0/1        0/4        0         3m

Each row is one actively provisioned (app, function, compute_shape) slice. INSTANCES and GPU UNITS show <allocated>/<provisioned> — a row with 0/1 is an idle warm runtime with no active sessions on it. SESSIONS is the live session count. AGE is how stale the capacity snapshot is; very old ages may indicate the runtime is no longer reporting.

Slices with no provisioned capacity are omitted, so this command answers "what is running right now". For the full deployment catalogue (including paused / failed / unprovisioned apps) use urun list apps; for historical or in-flight sessions use urun list sessions.

Pass --limit to control how many rows are fetched (default 100).

This command is experimental and requires the server-side GET /compute endpoint, which is in development.

Manage apps

Manage the lifecycle of a single deployed app. The app is addressed by its slug (the name shown under APP in urun list apps); every operation is org-scoped via your API key.

Show detailed status for one app (the single-app complement to list apps):

urun app status lingbot
            App: lingbot
           Name: LingBot
    Environment: prod
     App status: active
     Deployment: active
Desired replicas: 2
       Function: handle_lingbot_runtime
        Compute: b200:4
            GPU: 4 x b200
        Release: 1c6d6287abcd
  Live sessions: 1

Scale an app's runtime replica count (the backend's scaling knob; the control plane turns it into the runtime StatefulSet replica count):

urun app scale lingbot --replicas 3
urun app scale lingbot --replicas 0   # drain to zero without retiring

GPU count and compute shape are fixed at deploy time per release (set via @app.function), so scale intentionally exposes only --replicas.

Disable an app so the control plane stops running it (drives the deployment to paused and the app to disabled, so the materializer stops recreating its runtime). This is the clean, reversible, API-driven alternative to a manual database edit:

urun app disable lingbot-handle        # prompts for confirmation
urun app disable lingbot-handle --yes  # skip confirmation

Enable a disabled app and bring it back online:

urun app enable lingbot-handle

Stop a queued, starting, or live session:

urun stop session sess_123 --yes

All app subcommands accept --environment (default prod) and --json. These commands are experimental and require the server-side app lifecycle endpoint, which is in development.

Scratch instances (rung-4 verification, dev-only)

The verification ladder for platform changes is: unit tests -> CPU harness -> kind hop -> scratch pod -> full deploy. Rungs 1-3 are one command each; urun scratch makes rung 4 one command too — a real GPU pod running the real render, without the 30-90 min deploy loop and without touching the live app:

# Reproduce a broken app in isolation with debug env, candidate wheel overlaid:
urun scratch gemma-voice dg-brain --shape rtx6000:1 \
  --env VLLM_LOGGING_LEVEL=DEBUG --env CUDA_LAUNCH_BLOCKING=1 \
  --env PYTHONFAULTHANDLER=1 --name dg-brain-noble-repro

urun scratch ls                      # list scratch instances + expiry
urun scratch rm dg-brain-noble-repro # tear down
urun scratch rm --expired            # reap everything past its TTL

What it does:

  • Clones the app's rendered StatefulSet into an isolated instance with every platform label detached: the materializer never adopts it, the capacity reconciler never counts it, and no session is ever routed to it. The clone boots even when the source app is crashlooping or scaled to 0.
  • Overlays candidate artifacts: --wheel <req|url|path-on-storage> installs --no-deps into a hardlink COPY of the deps venv (the shared venv is never modified); --env KEY=VAL adds debug vars; hot reload is always pinned off so the instance stays on the currently-rendered release (--release <hash> asserts which one that is).
  • Takes GPUs explicitly: --shape is required and must match the source render; the target karpenter nodepool is checked against its GPU limit and the command refuses at capacity unless --allow-contention. Scratch never silently steals demo GPUs.
  • Never leaks: every instance carries a TTL annotation (default 4h, --ttl 6h/90m) consumed by the dev scratch reaper; urun scratch rm tears down sooner. Boot/self-test logs stream to your terminal (Ctrl-C detaches without tearing down).

Dev-only (v1): talks to the dev cluster via your kubeconfig (--context, default dev-usw2) and refuses prod contexts.

What gets deployed

urun deploy creates a source manifest from your Python entrypoint:

Entrypoint Included source
urun deploy app.py app.py and local Python files it imports

Dependencies are declared in your urun app code. Project-level files such as pyproject.toml and requirements.txt are not uploaded as dependency declarations by the CLI.

Generated/cache content such as .git, dotfiles, __pycache__, and .pyc files is excluded. Add .urunignore to exclude additional paths.

Shipping extra files: Dependencies(files=[...])

The collector's heuristic is Python-first: imported .py source, plus the non-Python files sitting next to that source. Anything it does not reach — a prompt tensor in its own directory, an asset an .urunignore pattern drops — is declared explicitly on your app's dependencies:

from urun import App, Dependencies

app = App("flashvsr-superres")


@app.function(
    deps=Dependencies(
        python=["torch"],
        files=["flashvsr_utils/prompt_tensor/posi_prompt.pth"],
    )
)
def superres(): ...
  • Paths are relative to the entrypoint file (app.py).
  • Declared files ship regardless of .urunignore and the built-in exclusions — that is the point of declaring them.
  • files= is purely additive. With no files=, collection is exactly what it was before.
  • Entries are individual files, not directories or globs. List each file.
  • A declared path that does not exist, escapes the app directory (.., a symlink pointing outside), or is absolute fails the deploy immediately, naming the path and the directory it was looked up under. A declared file is never silently skipped.

Common options

Shared by run and deploy:

Option Description
--name Override the derived app name.
--api-url Override the API URL; defaults to URUN_API_URL, saved login credentials, or https://api.urun.sh/v1.
--api-key Deploy API key; defaults to URUN_API_KEY or saved login credentials.
--no-wait Finalize but do not poll for readiness.
--poll-interval, --timeout Control readiness polling.

Troubleshooting

Error Fix
missing API key Run urun login, set URUN_API_KEY, or pass --api-key.
invalid API key format Use urun_<32 lowercase hex chars>.
entrypoint not found Run from the project root or pass the entrypoint path.
path is outside the project root Move the file under the project before deploying.
Expected files are missing Import local Python files from app.py. Non-Python assets are collected only when they sit next to collected source — declare anything else with Dependencies(files=[...]).

Development

Contributing and test instructions are in CONTRIBUTING.md.

License

MIT.

Development environment

This repo has a Nix/direnv/devcontainer baseline:

direnv allow
just sync
just check

Use VS Code Dev Containers to open the repository with the same toolchain in a container. Copy devcontainer.env.example to .devcontainer.env if you need to pass local git identity or other non-secret development settings into the container.

Release files for urun-cli 0.5.24

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for urun-cli 0.5.24
File Size Uploaded
urun_cli-0.5.24.tar.gz 238.1 kB Details

Built distributions (wheels)

Table of built distributions (wheels) for urun-cli 0.5.24
File
urun_cli-0.5.24-py3-none-manylinux_2_39_x86_64.whl Python 3 none Linux glibc 2.39+ x86-64 Details
urun_cli-0.5.24-py3-none-manylinux_2_39_aarch64.whl Python 3 none Linux glibc 2.39+ ARM64 Details
urun_cli-0.5.24-py3-none-macosx_15_0_arm64.whl Python 3 none macOS 15.0+ ARM64 Details
urun_cli-0.5.24-py3-none-any.whl Python 3 none any Details

Total release size: 102.2 MB

Release files / urun_cli-0.5.24.tar.gz

Download URL urun_cli-0.5.24.tar.gz
Size 238.1 kB
Tags Source
SHA-256 checksum
How to use checksums
a174789e50cf7439132504fb60db86f931878e517ec41955092984386d36c2ef
BLAKE2b-256 checksum
How to use checksums
43d569017e954a688e425b64e8ce208855083ead9865e7a1ed42cf337d083c3b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 5, 2026.

Transparency log

Release files / urun_cli-0.5.24-py3-none-manylinux_2_39_x86_64.whl

Download URL urun_cli-0.5.24-py3-none-manylinux_2_39_x86_64.whl
Size 39.7 MB
Tags Linux glibc 2.39+ x86-64 Python 3
SHA-256 checksum
How to use checksums
32065cf06e133b1e807ce16e6c77bf03c857cd59ca9c99577b31826dc041ea33
BLAKE2b-256 checksum
How to use checksums
bfd8bdbf68146b4a766725d8632e47f05a22f677561da80b616695d8e21bdc2b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 5, 2026.

Transparency log

Release files / urun_cli-0.5.24-py3-none-manylinux_2_39_aarch64.whl

Download URL urun_cli-0.5.24-py3-none-manylinux_2_39_aarch64.whl
Size 39.3 MB
Tags Linux glibc 2.39+ ARM64 Python 3
SHA-256 checksum
How to use checksums
fba25de5699429fb4e4198db3fdaf67fe362cd684d5eba151995a264613bb4d6
BLAKE2b-256 checksum
How to use checksums
80d1b4f92543225dddb31b872fb0f2a519f7c93ea90122d6c62f97b6451f17b7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 5, 2026.

Transparency log

Release files / urun_cli-0.5.24-py3-none-macosx_15_0_arm64.whl

Download URL urun_cli-0.5.24-py3-none-macosx_15_0_arm64.whl
Size 22.7 MB
Tags Python 3 macOS 15.0+ ARM64
SHA-256 checksum
How to use checksums
102bffe0fca2fcf1d496fb2bfc45ad15dce96ae85c0e95d60da2bca6c18e4b94
BLAKE2b-256 checksum
How to use checksums
edc89b388877e28b1e36c8b1040cae1dd0806d59e0098186cbeb23654e012414
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 5, 2026.

Transparency log

Release files / urun_cli-0.5.24-py3-none-any.whl

Download URL urun_cli-0.5.24-py3-none-any.whl
Size 246.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
817425ff15c6478f0bfcef709cf3de776b8b3c43c36c2d608c34c6b245e7056e
BLAKE2b-256 checksum
How to use checksums
444cd8f9bee595efffb24a6ce92e0cce769c15a16121963c626787a432b95bbd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 5, 2026.

Transparency log

Release history Release notifications | RSS feed

0.7.3

5 release files

0.7.2

5 release files

0.7.1

5 release files

0.7.0

5 release files

0.6.3

5 release files

0.6.2

5 release files

0.6.1

5 release files

0.6.0

5 release files

0.5.32

5 release files

0.5.31

5 release files

0.5.30

5 release files

0.5.29

5 release files

0.5.28

5 release files

0.5.27

5 release files

0.5.26

5 release files

0.5.25

5 release files

This release

0.5.24 This release

5 release files

0.5.22

2 release files

0.5.21

2 release files

0.5.20

2 release files

0.5.19

2 release files

0.5.18

2 release files

0.5.17

2 release files

0.5.16

2 release files

0.5.15

2 release files

0.5.14

2 release files

0.5.13

2 release files

0.5.12

2 release files

0.5.11

2 release files

0.5.10

2 release files

0.5.9

2 release files

0.5.8

2 release files

0.5.7

2 release files

0.5.6

2 release files

0.5.5

2 release files

0.5.4

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page