Skip to main content

Verda IO

Python worker framework for the Verda Inference Orchestrator.

Installation

pip install verda-io

Or with uv:

uv add verda-io

Quick Start

Create a worker with initialization and prediction functions:

import json
import verda_io

@verda_io.initialize
def setup():
    """Called once at startup. Load your model here."""
    global model
    model = load_my_model()

@verda_io.predict
def predict(payload: bytes) -> verda_io.Response:
    """Called for each inference request."""
    result = model(payload)
    return verda_io.Response(
        data=json.dumps(result).encode(),
        headers={"Content-Type": "application/json"},
    )

Run it with the Verda server:

verda-io run server -c config.yaml

Or run the worker directly:

python -m verda_io main.py --control-socket /tmp/verda-io.ctrl.sock --data-socket /tmp/verda-io.data.sock --name worker-1

Available flags:

  • --control-socket — Path to the control Unix socket
  • --data-socket — Path to the data Unix socket
  • --name — Worker identity name
  • -c, --config — Path to a YAML config file (alternative to flags)
  • --debug — Enable debug logging

Ordered Initialization

When your startup sequence requires specific ordering, use @verda_io.initialize with an order parameter. Functions run in ascending order:

import verda_io

@verda_io.initialize(0)
def load_tokenizer():
    global tokenizer
    tokenizer = AutoTokenizer.from_pretrained("model-name")

@verda_io.initialize(1)
def load_model():
    global model
    model = AutoModelForCausalLM.from_pretrained("model-name")

@verda_io.initialize(2)
def warm_up():
    model.generate(tokenizer("warm up", return_tensors="pt").input_ids)

Using @verda_io.initialize without an argument defaults to order 0. Multiple functions with the same order run in registration order.

Response Types

The @verda_io.predict function can return:

  • verda_io.Response — includes response data, HTTP headers, and status code (recommended)
  • bytes — sent as-is with 200 OK and application/octet-stream (backward compatible)
  • generator yielding bytes — streamed as SSE chunks
# Return with headers and status code
@verda_io.predict
def predict(payload: bytes) -> verda_io.Response:
    try:
        data = json.loads(payload)
        result = model(data)
        return verda_io.Response(
            data=json.dumps(result).encode(),
            headers={"Content-Type": "application/json"},
        )
    except ValueError as e:
        return verda_io.Response(
            data=json.dumps({"error": str(e)}).encode(),
            headers={"Content-Type": "application/json"},
            status_code=422,
        )

Streaming Responses

Return a generator from your predict function to stream results:

@verda_io.predict
def predict(payload: bytes):
    for token in model.generate_stream(payload):
        yield token.encode()

Hardware Health

Register a function to report hardware status via the /hardware-health endpoint:

@verda_io.hardware_health
def check_hardware() -> dict:
    import torch
    if not torch.cuda.is_available():
        return {"status": "no cuda device detected"}
    free, total = torch.cuda.mem_get_info(0)
    return {
        "status": "ok",
        "device": torch.cuda.get_device_properties(0).name,
        "total_memory_mb": total // (1024 * 1024),
        "free_memory_mb": free // (1024 * 1024),
    }

If no @verda_io.hardware_health function is registered, the endpoint returns {"status": "not enabled"}.

API

  • @verda_io.initialize - Register a startup function (called once, supports ordering)
  • @verda_io.predict - Register the prediction function (called per request, exactly one required)
  • @verda_io.hardware_health - Register a hardware health check function (optional)
  • verda_io.Response - Wrap response data with HTTP headers and status code
  • verda_io.run() - Start the worker loop programmatically

How It Works

The verda_io package connects to the Verda server via Unix domain sockets using a binary protocol (MessagePack + 12-byte headers). Initialization functions run once at startup in order, and @verda_io.predict is called for each inference request routed by the server. When a predict function returns verda_io.Response, the worker sends the data with an Envelope data type that carries HTTP headers and status code back to the server.

Requirements

  • Python 3.10+
  • Linux or macOS (Unix domain sockets required)

License

Apache License 2.0

Metadata

Release files for verda-io 0.4.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for verda-io 0.4.2
File Size Uploaded
verda_io-0.4.2.tar.gz 17.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for verda-io 0.4.2
File Interpreter ABI Platform
verda_io-0.4.2-py3-none-any.whl Python 3 none any Details

Total release size: 38.1 kB

Release files / verda_io-0.4.2.tar.gz

Download URL verda_io-0.4.2.tar.gz
Size 17.7 kB
Tags Source
SHA-256 checksum
How to use checksums
20f30320ad886b2ca53d7e2a81aade2beb4d1e0bf528ddc1c7bff05ce5413931
BLAKE2b-256 checksum
How to use checksums
ff8fd0b401116b03264afb9a3f8b76af678529867a7cf87327c3aef835a03522
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.8

Release files / verda_io-0.4.2-py3-none-any.whl

Download URL verda_io-0.4.2-py3-none-any.whl
Size 20.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d2889a89a9782fe16840060b2d0c40cac7101f9d7ff2e81b168be4299e5b8f36
BLAKE2b-256 checksum
How to use checksums
2579294303ea23b16f3d29e2bb0ea6d2df25af2304295901d9718053c9281c57
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.8

Release history Release notifications | RSS feed

0.4.4

2 release files

0.4.3

2 release files

This release

0.4.2 This release

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page