Skip to main content

Verda IO

Python worker framework for the Verda Inference Orchestrator.

Installation

pip install verda-io

Or with uv:

uv add verda-io

Quick Start

Create a worker with initialization and prediction functions:

import json
import verda_io

@verda_io.initialize
def setup():
    """Called once at startup. Load your model here."""
    global model
    model = load_my_model()

@verda_io.predict
def predict(payload: bytes) -> verda_io.Response:
    """Called for each inference request."""
    result = model(payload)
    return verda_io.Response(
        data=json.dumps(result).encode(),
        headers={"Content-Type": "application/json"},
    )

Run it with the Verda server:

verda-io run server -c config.yaml

Or run the worker directly:

python -m verda_io main.py --control-socket /tmp/verda-io.ctrl.sock --data-socket /tmp/verda-io.data.sock --name worker-1

Available flags:

  • --control-socket — Path to the control Unix socket
  • --data-socket — Path to the data Unix socket
  • --name — Worker identity name
  • -c, --config — Path to a YAML config file (alternative to flags)
  • --debug — Enable debug logging

Ordered Initialization

When your startup sequence requires specific ordering, use @verda_io.initialize with an order parameter. Functions run in ascending order:

import verda_io

@verda_io.initialize(0)
def load_tokenizer():
    global tokenizer
    tokenizer = AutoTokenizer.from_pretrained("model-name")

@verda_io.initialize(1)
def load_model():
    global model
    model = AutoModelForCausalLM.from_pretrained("model-name")

@verda_io.initialize(2)
def warm_up():
    model.generate(tokenizer("warm up", return_tensors="pt").input_ids)

Using @verda_io.initialize without an argument defaults to order 0. Multiple functions with the same order run in registration order.

Response Types

The @verda_io.predict function can return:

  • verda_io.Response — includes response data, HTTP headers, and status code (recommended)
  • bytes — sent as-is with 200 OK and application/octet-stream (backward compatible)
  • generator yielding bytes — streamed as SSE chunks
# Return with headers and status code
@verda_io.predict
def predict(payload: bytes) -> verda_io.Response:
    try:
        data = json.loads(payload)
        result = model(data)
        return verda_io.Response(
            data=json.dumps(result).encode(),
            headers={"Content-Type": "application/json"},
        )
    except ValueError as e:
        return verda_io.Response(
            data=json.dumps({"error": str(e)}).encode(),
            headers={"Content-Type": "application/json"},
            status_code=422,
        )

Streaming Responses

Return a generator from your predict function to stream results:

@verda_io.predict
def predict(payload: bytes):
    for token in model.generate_stream(payload):
        yield token.encode()

Hardware Health

Register a function to report hardware status via the /hardware-health endpoint:

@verda_io.hardware_health
def check_hardware() -> dict:
    import torch
    if not torch.cuda.is_available():
        return {"status": "no cuda device detected"}
    free, total = torch.cuda.mem_get_info(0)
    return {
        "status": "ok",
        "device": torch.cuda.get_device_properties(0).name,
        "total_memory_mb": total // (1024 * 1024),
        "free_memory_mb": free // (1024 * 1024),
    }

If no @verda_io.hardware_health function is registered, the endpoint returns {"status": "not enabled"}.

API

  • @verda_io.initialize - Register a startup function (called once, supports ordering)
  • @verda_io.predict - Register the prediction function (called per request, exactly one required)
  • @verda_io.hardware_health - Register a hardware health check function (optional)
  • verda_io.Response - Wrap response data with HTTP headers and status code
  • verda_io.run() - Start the worker loop programmatically

How It Works

The verda_io package connects to the Verda server via Unix domain sockets using a binary protocol (MessagePack + 12-byte headers). Initialization functions run once at startup in order, and @verda_io.predict is called for each inference request routed by the server. When a predict function returns verda_io.Response, the worker sends the data with an Envelope data type that carries HTTP headers and status code back to the server.

Requirements

  • Python 3.10+
  • Linux or macOS (Unix domain sockets required)

License

Apache License 2.0

Metadata

Release files for verda-io 0.4.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for verda-io 0.4.3
File Size Uploaded
verda_io-0.4.3.tar.gz 17.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for verda-io 0.4.3
File Interpreter ABI Platform
verda_io-0.4.3-py3-none-any.whl Python 3 none any Details

Total release size: 38.1 kB

Release files / verda_io-0.4.3.tar.gz

Download URL verda_io-0.4.3.tar.gz
Size 17.7 kB
Tags Source
SHA-256 checksum
How to use checksums
1e811bf63164ca0608aea9959cec9bbfe72608223fa22794cf0a23e410523a80
BLAKE2b-256 checksum
How to use checksums
3585791eb2682a058c7e729a0f92132a6241051e7869a89205bf754a041898c5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.8

Release files / verda_io-0.4.3-py3-none-any.whl

Download URL verda_io-0.4.3-py3-none-any.whl
Size 20.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5eb005275d6c3a94d857d4d4ed40579eb30b1c1a0d1092508e216880e221bc02
BLAKE2b-256 checksum
How to use checksums
3fee2197ac44d8f5a7117e53dabb09ea59a09ef13d2c13b4bb4608f77be21aca
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.8

Release history Release notifications | RSS feed

0.4.4

2 release files

This release

0.4.3 This release

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page