Skip to main content

Verda IO

Python worker framework for the Verda Inference Orchestrator.

Installation

pip install verda-io

Or with uv:

uv add verda-io

Quick Start

Create a worker with initialization and prediction functions:

import json
import verda_io

@verda_io.initialize
def setup():
    """Called once at startup. Load your model here."""
    global model
    model = load_my_model()

@verda_io.predict
def predict(payload: bytes) -> verda_io.Response:
    """Called for each inference request."""
    result = model(payload)
    return verda_io.Response(
        data=json.dumps(result).encode(),
        headers={"Content-Type": "application/json"},
    )

Run it with the Verda server:

verda-io run server -c config.yaml

Or run the worker directly:

python -m verda_io main.py --control-socket /tmp/verda-io.ctrl.sock --data-socket /tmp/verda-io.data.sock --name worker-1

Available flags:

  • --control-socket — Path to the control Unix socket
  • --data-socket — Path to the data Unix socket
  • --name — Worker identity name
  • -c, --config — Path to a YAML config file (alternative to flags)
  • --debug — Enable debug logging

Ordered Initialization

When your startup sequence requires specific ordering, use @verda_io.initialize with an order parameter. Functions run in ascending order:

import verda_io

@verda_io.initialize(0)
def load_tokenizer():
    global tokenizer
    tokenizer = AutoTokenizer.from_pretrained("model-name")

@verda_io.initialize(1)
def load_model():
    global model
    model = AutoModelForCausalLM.from_pretrained("model-name")

@verda_io.initialize(2)
def warm_up():
    model.generate(tokenizer("warm up", return_tensors="pt").input_ids)

Using @verda_io.initialize without an argument defaults to order 0. Multiple functions with the same order run in registration order.

Response Types

The @verda_io.predict function can return:

  • verda_io.Response — includes response data, HTTP headers, and status code (recommended)
  • bytes — sent as-is with 200 OK and application/octet-stream (backward compatible)
  • generator yielding bytes — streamed as SSE chunks
# Return with headers and status code
@verda_io.predict
def predict(payload: bytes) -> verda_io.Response:
    try:
        data = json.loads(payload)
        result = model(data)
        return verda_io.Response(
            data=json.dumps(result).encode(),
            headers={"Content-Type": "application/json"},
        )
    except ValueError as e:
        return verda_io.Response(
            data=json.dumps({"error": str(e)}).encode(),
            headers={"Content-Type": "application/json"},
            status_code=422,
        )

Streaming Responses

Return a generator from your predict function to stream results:

@verda_io.predict
def predict(payload: bytes):
    for token in model.generate_stream(payload):
        yield token.encode()

Hardware Health

Register a function to report hardware status via the /hardware-health endpoint:

@verda_io.hardware_health
def check_hardware() -> dict:
    import torch
    if not torch.cuda.is_available():
        return {"status": "no cuda device detected"}
    free, total = torch.cuda.mem_get_info(0)
    return {
        "status": "ok",
        "device": torch.cuda.get_device_properties(0).name,
        "total_memory_mb": total // (1024 * 1024),
        "free_memory_mb": free // (1024 * 1024),
    }

If no @verda_io.hardware_health function is registered, the endpoint returns {"status": "not enabled"}.

API

  • @verda_io.initialize - Register a startup function (called once, supports ordering)
  • @verda_io.predict - Register the prediction function (called per request, exactly one required)
  • @verda_io.hardware_health - Register a hardware health check function (optional)
  • verda_io.Response - Wrap response data with HTTP headers and status code
  • verda_io.run() - Start the worker loop programmatically

How It Works

The verda_io package connects to the Verda server via Unix domain sockets using a binary protocol (MessagePack + 12-byte headers). Initialization functions run once at startup in order, and @verda_io.predict is called for each inference request routed by the server. When a predict function returns verda_io.Response, the worker sends the data with an Envelope data type that carries HTTP headers and status code back to the server.

Requirements

  • Python 3.10+
  • Linux or macOS (Unix domain sockets required)

License

Apache License 2.0

Metadata

Release files for verda-io 0.4.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for verda-io 0.4.4
File Size Uploaded
verda_io-0.4.4.tar.gz 17.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for verda-io 0.4.4
File Interpreter ABI Platform
verda_io-0.4.4-py3-none-any.whl Python 3 none any Details

Total release size: 38.4 kB

Release files / verda_io-0.4.4.tar.gz

Download URL verda_io-0.4.4.tar.gz
Size 17.8 kB
Tags Source
SHA-256 checksum
How to use checksums
263927953d3b701d40eb96f626a746de58ef7a918d7cbc20fc6fd745f8a3fbf2
BLAKE2b-256 checksum
How to use checksums
83734ffe165022639bb920c9500d9f10333ea8459b94722f4ff110dcf118c15b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.8

Release files / verda_io-0.4.4-py3-none-any.whl

Download URL verda_io-0.4.4-py3-none-any.whl
Size 20.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3db5774bee868d07406f7a57678570cd19dcb405ce61c33f06c5361254b58524
BLAKE2b-256 checksum
How to use checksums
efeaafb09e4abc8ecdd1617ec7f5f9c051ef603bfcc727d46fd9f96dc950b613
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.8

Release history Release notifications | RSS feed

This release

0.4.4 This release

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page