Skip to main content

iocloud

CI Python License Status

Python SDK and CLI for io.net compute: run containers, serve models behind an OpenAI-compatible API, fine-tune them, and open a GPU notebook — on a decentralised GPU network, from one tool.

There is exactly one primitive underneath: a CaaS deployment, a prepaid, time-boxed cluster running N identical replicas of one container image. Every product in this SDK is that primitive with a job-shaped wrapper around it.

Product What you get CLI Python
Endpoints vLLM/SGLang behind an OpenAI-compatible API, authenticated by default iocloud endpoint client.endpoints
Fine-tuning A TRL job (SFT/DPO/GRPO) with an S3 data plane iocloud finetune client.finetune
Notebooks Jupyter on a GPU box, token generated locally iocloud notebook client.notebooks
Recipes Curated, ready-to-deploy model configs (consumer GGUF to datacenter) iocloud recipe iocloud.recipes
RLVR Colocated GRPO with verifiable rewards iocloud rlvr client.rlvr
Raw deployments Any container you can push to a registry iocloud deploy client.deployments

Install

Python 3.10 or newer.

pip install ionet-cloud          # the package is 'ionet-cloud'; you import 'iocloud'

Or from source:

git clone https://github.com/ionet-official/iocloud.git
cd iocloud
pip install -e .

Optional extras: ionet-cloud[openai] (Endpoint.openai_client()), ionet-cloud[data] (iocloud data push/pull, fine-tune result verification), ionet-cloud[tui], or ionet-cloud[all].

Sixty seconds to a running model

iocloud auth login                    # paste an API token from the io.net console

# Price it before you spend: every mutating command takes --dry-run.
iocloud endpoint deploy --model Qwen/Qwen3-8B --gpu H100:8 --hours 1 --dry-run

# Deploy for real. This blocks until the server answers its health path,
# not merely until the platform says "running".
iocloud endpoint deploy --model Qwen/Qwen3-8B --gpu H100:8 --hours 1

iocloud endpoint list
iocloud endpoint call <endpoint-id> --prompt "Write a haiku about GPUs."
iocloud endpoint delete <endpoint-id>

iocloud dash                          # live dashboard, with pip install 'ionet-cloud[tui]'

--hours is prepaid: you are billed up front and deleting early refunds nothing, it only frees the capacity. That is why --dry-run prints a live quote before anything is charged.

The same thing from Python:

import iocloud

with iocloud.Client() as io:                      # IOCLOUD_API_KEY, or ~/.iocloud
    endpoint = io.endpoints.create(model="Qwen/Qwen3-8B", gpu="H100:8", hours=1)
    endpoint.wait()                             # running + URL + health 200
    print(endpoint.url)

    client = endpoint.openai_client()           # needs ionet-cloud[openai]
    answer = client.chat.completions.create(
        model="Qwen/Qwen3-8B",
        messages=[{"role": "user", "content": "Write a haiku about GPUs."}],
    )
    print(answer.choices[0].message.content)

    endpoint.delete()

Deploy your own container instead:

import iocloud

io = iocloud.Client()
deployment = io.deployments.create(
    name="web",
    image="nginx:latest",
    gpu="H100:8",
    hours=1,
    port=80,
)
deployment.wait()
print(deployment.url())
for line in deployment.logs(follow=True):
    print(line.text)

iocloud.AsyncClient is the 1:1 async twin — same names, same arguments, await in front.

Recipes

A recipe is a curated, ready-to-deploy serving config for a model: it pins the inference server, the concrete GPU hardware, the GPU count, and the server flags that make that model fast and fit in memory. Deploy one instead of remembering the right --gpu, --max-model-len and quantization per model.

iocloud recipe list                        # all recipes (bundled + your own)
iocloud recipe list --task coding          # filter by use-case
iocloud recipe show qwen3.6-27b            # the full config + rendered flags

iocloud endpoint deploy --recipe qwen3.6-27b       # deploy from a recipe
iocloud endpoint deploy --model Qwen/Qwen3.6-27B   # resolve that model's default recipe

The bundled catalogue spans consumer GPUs (quantized GGUF via llama.cpp) up to datacenter cards (vLLM on H100/H200/B200), across chat, coding, vision, reasoning and video tasks, and is updated as new models ship.

A model can carry several variants (name@variant) for different hardware tiers, and every recipe field is just a default — any flag you pass wins:

iocloud endpoint deploy --recipe qwen3.6-27b@h100    # a specific tier
iocloud endpoint deploy --recipe qwen3.6-27b --gpu H100:1   # consumer recipe on a bigger card

Overriding --gpu swaps only the hardware; the image and server stay put, so a consumer (GGUF) recipe deploys unchanged on a larger card — the extra VRAM is headroom.

From Python, and to add your own recipes:

import iocloud

io = iocloud.Client()
endpoint = io.endpoints.create(recipe="qwen3.6-27b")   # overrides win, e.g. gpu="H100:1"
endpoint.wait()
print(endpoint.url)

Drop a YAML file in ~/.iocloud/recipes/ to add a recipe or shadow a bundled one of the same name[@variant]; iocloud recipe list marks yours SOURCE = user. See Recipes for the schema and every field.

Documentation

Full docs live in docs/ and build into a site with make docs-serve (needs pip install -e ".[docs]").

Runnable examples are in examples/.

Project status

Pre-1.0, and honest about what that means:

  • The SDK and CLI are feature-complete for the surface documented above; anything not in the docs does not exist yet.
  • API models are hand-written against the live backend, not generated — the bundled OpenAPI schema is stale. Server-side changes can therefore drift; the SDK tolerates unknown enum values rather than crashing, and an optional nightly live-smoke workflow is the only drift detector.
  • Product bookkeeping (which deployment is an "endpoint") lives in a local state file, because the API has no tags. It is best-effort and documented as such: iocloud deployment list always shows the ground truth.
  • Breaking changes are possible before 1.0.

Not supported, by platform constraint rather than by omission: volumes and FUSE mounts, more than one port per deployment, autoscaling or scale-to-zero, and east-west networking between deployments. See Limits.

Development

make dev          # pip install -e ".[dev,tui]"
make check        # ruff + mypy + pytest with the coverage gate
make test         # just the offline suite
make docs-serve   # preview the documentation site

Tests are offline by default: a FakeCaaS fixture simulates the deployment lifecycle, so the suite never touches the network or your credentials. The opt-in live smoke tests (make test-live) need a real token.

Contributions are welcome — small, reviewable pull requests with tests, please. Every code sample in this README and in docs/ is compiled by tests/docs/test_examples.py, so keep them real.

License

Released under the MIT License.

Metadata

Release files for ionet-cloud 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ionet-cloud 0.3.0
File Size Uploaded
ionet_cloud-0.3.0.tar.gz 176.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ionet-cloud 0.3.0
File Interpreter ABI Platform
ionet_cloud-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 388.8 kB

Release files / ionet_cloud-0.3.0.tar.gz

Download URL ionet_cloud-0.3.0.tar.gz
Size 176.7 kB
Tags Source
SHA-256 checksum
How to use checksums
4ce5c28354f722c9e9991d212e04914e2301fa5ba5387216eddf09e02d3609b0
BLAKE2b-256 checksum
How to use checksums
773e9fb22366cbc0242706d4d4bf0dc2de3ee14a9f3e9f940f72a98fb549d574
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / ionet_cloud-0.3.0-py3-none-any.whl

Download URL ionet_cloud-0.3.0-py3-none-any.whl
Size 212.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d22dc9fc00794c6218f9c3c5cf8318707a313c66e7d49ec67e487474c5f1e1bd
BLAKE2b-256 checksum
How to use checksums
9656abcf938a224ab8073d5aa546ee80ddf07afa55482a9b518ce62831f32f15
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

0.4.1

2 release files

0.4.0

2 release files

This release

0.3.0 This release

2 release files

0.2.1

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page