Verda IO
Python worker framework for the Verda Inference Orchestrator.
Installation
pip install verda-io
Or with uv:
uv add verda-io
Quick Start
Create a worker with initialization and prediction functions:
import json
import verda_io
@verda_io.initialize
def setup():
"""Called once at startup. Load your model here."""
global model
model = load_my_model()
@verda_io.predict
def predict(payload: bytes) -> verda_io.Response:
"""Called for each inference request."""
result = model(payload)
return verda_io.Response(
data=json.dumps(result).encode(),
headers={"Content-Type": "application/json"},
)
Run it with the Verda server:
verda-io run server -c config.yaml
Or run the worker directly:
python -m verda_io main.py --control-socket /tmp/verda-io.ctrl.sock --data-socket /tmp/verda-io.data.sock --name worker-1
Available flags:
--control-socket— Path to the control Unix socket--data-socket— Path to the data Unix socket--name— Worker identity name-c, --config— Path to a YAML config file (alternative to flags)--debug— Enable debug logging
Ordered Initialization
When your startup sequence requires specific ordering, use @verda_io.initialize with an order parameter. Functions run in ascending order:
import verda_io
@verda_io.initialize(0)
def load_tokenizer():
global tokenizer
tokenizer = AutoTokenizer.from_pretrained("model-name")
@verda_io.initialize(1)
def load_model():
global model
model = AutoModelForCausalLM.from_pretrained("model-name")
@verda_io.initialize(2)
def warm_up():
model.generate(tokenizer("warm up", return_tensors="pt").input_ids)
Using @verda_io.initialize without an argument defaults to order 0. Multiple functions with the same order run in registration order.
Response Types
The @verda_io.predict function can return:
verda_io.Response— includes response data, HTTP headers, and status code (recommended)bytes— sent as-is with200 OKandapplication/octet-stream(backward compatible)- generator yielding
bytes— streamed as SSE chunks
# Return with headers and status code
@verda_io.predict
def predict(payload: bytes) -> verda_io.Response:
try:
data = json.loads(payload)
result = model(data)
return verda_io.Response(
data=json.dumps(result).encode(),
headers={"Content-Type": "application/json"},
)
except ValueError as e:
return verda_io.Response(
data=json.dumps({"error": str(e)}).encode(),
headers={"Content-Type": "application/json"},
status_code=422,
)
Streaming Responses
Return a generator from your predict function to stream results:
@verda_io.predict
def predict(payload: bytes):
for token in model.generate_stream(payload):
yield token.encode()
Hardware Health
Register a function to report hardware status via the /hardware-health endpoint:
@verda_io.hardware_health
def check_hardware() -> dict:
import torch
if not torch.cuda.is_available():
return {"status": "no cuda device detected"}
free, total = torch.cuda.mem_get_info(0)
return {
"status": "ok",
"device": torch.cuda.get_device_properties(0).name,
"total_memory_mb": total // (1024 * 1024),
"free_memory_mb": free // (1024 * 1024),
}
If no @verda_io.hardware_health function is registered, the endpoint returns {"status": "not enabled"}.
API
@verda_io.initialize- Register a startup function (called once, supports ordering)@verda_io.predict- Register the prediction function (called per request, exactly one required)@verda_io.hardware_health- Register a hardware health check function (optional)verda_io.Response- Wrap response data with HTTP headers and status codeverda_io.run()- Start the worker loop programmatically
How It Works
The verda_io package connects to the Verda server via Unix domain sockets using a binary protocol (MessagePack + 12-byte headers). Initialization functions run once at startup in order, and @verda_io.predict is called for each inference request routed by the server. When a predict function returns verda_io.Response, the worker sends the data with an Envelope data type that carries HTTP headers and status code back to the server.
Requirements
- Python 3.10+
- Linux or macOS (Unix domain sockets required)
License
Apache License 2.0
Metadata
Release files for verda-io 0.4.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| verda_io-0.4.3.tar.gz | 17.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| verda_io-0.4.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 38.1 kB
Release files / verda_io-0.4.3.tar.gz
| Download URL | verda_io-0.4.3.tar.gz |
|---|---|
| Size | 17.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1e811bf63164ca0608aea9959cec9bbfe72608223fa22794cf0a23e410523a80
|
|
BLAKE2b-256 checksum How to use checksums |
3585791eb2682a058c7e729a0f92132a6241051e7869a89205bf754a041898c5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.8
|
Release files / verda_io-0.4.3-py3-none-any.whl
| Download URL | verda_io-0.4.3-py3-none-any.whl |
|---|---|
| Size | 20.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5eb005275d6c3a94d857d4d4ed40579eb30b1c1a0d1092508e216880e221bc02
|
|
BLAKE2b-256 checksum How to use checksums |
3fee2197ac44d8f5a7117e53dabb09ea59a09ef13d2c13b4bb4608f77be21aca
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.8
|