Verda IO
Python worker framework for the Verda Inference Orchestrator.
Installation
pip install verda-io
Or with uv:
uv add verda-io
Quick Start
Create a worker with initialization and prediction functions:
import json
import verda_io
@verda_io.initialize
def setup():
"""Called once at startup. Load your model here."""
global model
model = load_my_model()
@verda_io.predict
def predict(payload: bytes) -> verda_io.Response:
"""Called for each inference request."""
result = model(payload)
return verda_io.Response(
data=json.dumps(result).encode(),
headers={"Content-Type": "application/json"},
)
Run it with the Verda server:
verda-io run server -c config.yaml
Or run the worker directly:
python -m verda_io main.py --control-socket /tmp/verda-io.ctrl.sock --data-socket /tmp/verda-io.data.sock --name worker-1
Available flags:
--control-socket— Path to the control Unix socket--data-socket— Path to the data Unix socket--name— Worker identity name-c, --config— Path to a YAML config file (alternative to flags)--debug— Enable debug logging
Ordered Initialization
When your startup sequence requires specific ordering, use @verda_io.initialize with an order parameter. Functions run in ascending order:
import verda_io
@verda_io.initialize(0)
def load_tokenizer():
global tokenizer
tokenizer = AutoTokenizer.from_pretrained("model-name")
@verda_io.initialize(1)
def load_model():
global model
model = AutoModelForCausalLM.from_pretrained("model-name")
@verda_io.initialize(2)
def warm_up():
model.generate(tokenizer("warm up", return_tensors="pt").input_ids)
Using @verda_io.initialize without an argument defaults to order 0. Multiple functions with the same order run in registration order.
Response Types
The @verda_io.predict function can return:
verda_io.Response— includes response data, HTTP headers, and status code (recommended)bytes— sent as-is with200 OKandapplication/octet-stream(backward compatible)- generator yielding
bytes— streamed as SSE chunks
# Return with headers and status code
@verda_io.predict
def predict(payload: bytes) -> verda_io.Response:
try:
data = json.loads(payload)
result = model(data)
return verda_io.Response(
data=json.dumps(result).encode(),
headers={"Content-Type": "application/json"},
)
except ValueError as e:
return verda_io.Response(
data=json.dumps({"error": str(e)}).encode(),
headers={"Content-Type": "application/json"},
status_code=422,
)
Streaming Responses
Return a generator from your predict function to stream results:
@verda_io.predict
def predict(payload: bytes):
for token in model.generate_stream(payload):
yield token.encode()
Hardware Health
Register a function to report hardware status via the /hardware-health endpoint:
@verda_io.hardware_health
def check_hardware() -> dict:
import torch
if not torch.cuda.is_available():
return {"status": "no cuda device detected"}
free, total = torch.cuda.mem_get_info(0)
return {
"status": "ok",
"device": torch.cuda.get_device_properties(0).name,
"total_memory_mb": total // (1024 * 1024),
"free_memory_mb": free // (1024 * 1024),
}
If no @verda_io.hardware_health function is registered, the endpoint returns {"status": "not enabled"}.
API
@verda_io.initialize- Register a startup function (called once, supports ordering)@verda_io.predict- Register the prediction function (called per request, exactly one required)@verda_io.hardware_health- Register a hardware health check function (optional)verda_io.Response- Wrap response data with HTTP headers and status codeverda_io.run()- Start the worker loop programmatically
How It Works
The verda_io package connects to the Verda server via Unix domain sockets using a binary protocol (MessagePack + 12-byte headers). Initialization functions run once at startup in order, and @verda_io.predict is called for each inference request routed by the server. When a predict function returns verda_io.Response, the worker sends the data with an Envelope data type that carries HTTP headers and status code back to the server.
Requirements
- Python 3.10+
- Linux or macOS (Unix domain sockets required)
License
Apache License 2.0
Metadata
Release files for verda-io 0.4.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| verda_io-0.4.2.tar.gz | 17.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| verda_io-0.4.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 38.1 kB
Release files / verda_io-0.4.2.tar.gz
| Download URL | verda_io-0.4.2.tar.gz |
|---|---|
| Size | 17.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
20f30320ad886b2ca53d7e2a81aade2beb4d1e0bf528ddc1c7bff05ce5413931
|
|
BLAKE2b-256 checksum How to use checksums |
ff8fd0b401116b03264afb9a3f8b76af678529867a7cf87327c3aef835a03522
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.8
|
Release files / verda_io-0.4.2-py3-none-any.whl
| Download URL | verda_io-0.4.2-py3-none-any.whl |
|---|---|
| Size | 20.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
d2889a89a9782fe16840060b2d0c40cac7101f9d7ff2e81b168be4299e5b8f36
|
|
BLAKE2b-256 checksum How to use checksums |
2579294303ea23b16f3d29e2bb0ea6d2df25af2304295901d9718053c9281c57
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.8
|