Skip to main content

Verda IO

Python worker framework for the Verda Inference Orchestrator.

Installation

pip install verda-io

Or with uv:

uv add verda-io

Quick Start

Create a worker with initialization and prediction functions:

import json
import verda_io

@verda_io.initialize
def setup():
    """Called once at startup. Load your model here."""
    global model
    model = load_my_model()

@verda_io.predict
def predict(payload: bytes) -> verda_io.Response:
    """Called for each inference request."""
    result = model(payload)
    return verda_io.Response(
        data=json.dumps(result).encode(),
        headers={"Content-Type": "application/json"},
    )

Run it with the Verda server:

verda-io run server -c config.yaml

Or run the worker directly:

python -m verda_io main.py --control-socket /tmp/verda-io.ctrl.sock --data-socket /tmp/verda-io.data.sock --name worker-1

Available flags:

  • --control-socket — Path to the control Unix socket
  • --data-socket — Path to the data Unix socket
  • --name — Worker identity name
  • -c, --config — Path to a YAML config file (alternative to flags)
  • --debug — Enable debug logging

Ordered Initialization

When your startup sequence requires specific ordering, use @verda_io.initialize with an order parameter. Functions run in ascending order:

import verda_io

@verda_io.initialize(0)
def load_tokenizer():
    global tokenizer
    tokenizer = AutoTokenizer.from_pretrained("model-name")

@verda_io.initialize(1)
def load_model():
    global model
    model = AutoModelForCausalLM.from_pretrained("model-name")

@verda_io.initialize(2)
def warm_up():
    model.generate(tokenizer("warm up", return_tensors="pt").input_ids)

Using @verda_io.initialize without an argument defaults to order 0. Multiple functions with the same order run in registration order.

Response Types

The @verda_io.predict function can return:

  • verda_io.Response — includes response data, HTTP headers, and status code (recommended)
  • bytes — sent as-is with 200 OK and application/octet-stream (backward compatible)
  • generator yielding bytes — streamed as SSE chunks
# Return with headers and status code
@verda_io.predict
def predict(payload: bytes) -> verda_io.Response:
    try:
        data = json.loads(payload)
        result = model(data)
        return verda_io.Response(
            data=json.dumps(result).encode(),
            headers={"Content-Type": "application/json"},
        )
    except ValueError as e:
        return verda_io.Response(
            data=json.dumps({"error": str(e)}).encode(),
            headers={"Content-Type": "application/json"},
            status_code=422,
        )

Streaming Responses

Return a generator from your predict function to stream results:

@verda_io.predict
def predict(payload: bytes):
    for token in model.generate_stream(payload):
        yield token.encode()

Hardware Health

Register a function to report hardware status via the /hardware-health endpoint:

@verda_io.hardware_health
def check_hardware() -> dict:
    import torch
    if not torch.cuda.is_available():
        return {"status": "no cuda device detected"}
    free, total = torch.cuda.mem_get_info(0)
    return {
        "status": "ok",
        "device": torch.cuda.get_device_properties(0).name,
        "total_memory_mb": total // (1024 * 1024),
        "free_memory_mb": free // (1024 * 1024),
    }

If no @verda_io.hardware_health function is registered, the endpoint returns {"status": "not enabled"}.

API

  • @verda_io.initialize - Register a startup function (called once, supports ordering)
  • @verda_io.predict - Register the prediction function (called per request, exactly one required)
  • @verda_io.hardware_health - Register a hardware health check function (optional)
  • verda_io.Response - Wrap response data with HTTP headers and status code
  • verda_io.run() - Start the worker loop programmatically

How It Works

The verda_io package connects to the Verda server via Unix domain sockets using a binary protocol (MessagePack + 12-byte headers). Initialization functions run once at startup in order, and @verda_io.predict is called for each inference request routed by the server. When a predict function returns verda_io.Response, the worker sends the data with an Envelope data type that carries HTTP headers and status code back to the server.

Requirements

  • Python 3.10+
  • Linux or macOS (Unix domain sockets required)

License

Apache License 2.0

Metadata

Release files for verda-io 0.4.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for verda-io 0.4.1
File Size Uploaded
verda_io-0.4.1.tar.gz 17.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for verda-io 0.4.1
File Interpreter ABI Platform
verda_io-0.4.1-py3-none-any.whl Python 3 none any Details

Total release size: 38.0 kB

Release files / verda_io-0.4.1.tar.gz

Download URL verda_io-0.4.1.tar.gz
Size 17.7 kB
Tags Source
SHA-256 checksum
How to use checksums
e5e450b03cdd694356a030c0c5386a684c026852a39b6661fc526741e093833a
BLAKE2b-256 checksum
How to use checksums
311c42b349ef2809fa345dde5ad58ddac07abae9225a4c338960407e9e2e609f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.8

Release files / verda_io-0.4.1-py3-none-any.whl

Download URL verda_io-0.4.1-py3-none-any.whl
Size 20.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1d314daeeb86efd10dbba5590bf592d571c69b3c9c9258455b9c2aa20a160619
BLAKE2b-256 checksum
How to use checksums
6d13b8396e414a8942fb4aacadb95a982ebdfe9b6198e07578214752b2a748f0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.8

Release history Release notifications | RSS feed

0.4.4

2 release files

0.4.3

2 release files

0.4.2

2 release files

This release

0.4.1 This release

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page