Python client for MoDIS Service Nodes
Project description
MoDIS Python Client
modis-client is the Python client package for MoDIS Service Nodes (MSNs).
It exposes MoDIS-native request and response objects for direct Python callers.
The PyPI distribution name is modis-client. pymodis and modis are already
used by unrelated packages, while modis-client currently appears available.
Install
From a checkout:
python -m pip install -e .
The import package is modis_client:
from modis_client import SingleNodeClient
with SingleNodeClient(base_url="http://localhost:8000") as client:
response = client.chat_completion({
"model": "qwen3.6-27b",
"messages": [{"role": "user", "content": "Hello MoDIS."}],
})
print(response.content.content)
The package supports sync and async single-node clients, routing clients, batch requests, explicit streaming methods, MoDIS tool-call fields, and Pydantic v2 models with MoDIS wire aliases.
The default client timeout is 900 seconds because a MoDIS request may include a
cold model load before inference begins. Pass timeout= to any client
constructor when a shorter or longer budget is required.
Async Client
from modis_client import AsyncSingleNodeClient
async with AsyncSingleNodeClient(base_url="http://localhost:8000") as client:
response = await client.responses({
"model": "qwen3.6-27b",
"input": "Summarize MoDIS in one sentence.",
})
Batch Requests
responses = client.chat_completion([
{"model": "qwen3.6-27b", "messages": [{"role": "user", "content": "one"}]},
{"model": "qwen3.6-27b", "messages": [{"role": "user", "content": "two"}]},
])
Streaming
for event in client.chat_completion_stream({
"model": "qwen3.6-27b",
"messages": [{"role": "user", "content": "Count to three."}],
}):
if event.event == "delta":
print(event.payload)
Image streams preserve progress and final image payload events:
for event in client.image_generation_stream({
"model": "qwen-image",
"prompt": "A clean product photo of a compact AI workstation.",
"numOutputs": 1,
}):
if event.event == "progress":
print(event.payload["step"], event.payload.get("total"))
if event.event == "completed":
images = event.payload["images"]
Text To Speech
Non-streaming speech requests return a Pydantic response with audio_b64,
audio_url, and mime_type:
import base64
response = client.text_to_speech({
"model": "qwen-tts-custom-voice",
"text": "Hello from MoDIS.",
"speaker": "Vivian",
"responseFormat": "wav",
})
audio_bytes = base64.b64decode(response.audio_b64)
PCM streaming is exposed separately because /audio/speech returns raw audio
bytes, not SSE events, when stream: true is used:
with open("speech.pcm", "wb") as output:
for chunk in client.text_to_speech_pcm_stream({
"model": "qwen-tts-custom-voice",
"text": "Hello from MoDIS.",
"speaker": "Vivian",
}):
output.write(chunk)
Reference voices can be passed per request with referenceAudio and
referenceText. Use hosted URLs when the vLLM-Omni server can fetch the audio,
or inline data URLs/base64 when the reference is not hosted.
Speech To Text
Non-streaming ASR requests call /audio/transcriptions and return a transcript
in response.text:
response = client.speech_to_text({
"model": "qwen3-asr",
"audioUrl": "https://example.com/audio.wav",
"language": "en",
})
print(response.text)
The typed SpeechToTextRequest model also accepts audioB64, audioPath,
audio, url, prompt, responseFormat, and timeoutSeconds. The current
MoDIS ASR route is non-streaming; uploaded-audio output streaming and realtime
WebSocket transcription are service follow-ups.
Structured Output
response = client.chat_completion({
"model": "qwen3.6-27b",
"messages": [{"role": "user", "content": "Return one rating."}],
"format": {
"title": "Rating",
"type": "object",
"properties": {"score": {"type": "integer"}},
"required": ["score"],
},
})
format: "json" requests JSON-object output. A JSON Schema object is forwarded
to MoDIS text runtimes as a structured response schema.
Reasoning Output
response = client.chat_completion({
"model": "qwen3.6-27b",
"messages": [{"role": "user", "content": "Solve 25^(1/2)."}],
"includeReasoning": True,
})
print(response.content.reasoning)
print(response.content.content)
includeReasoning asks MoDIS to include parsed runtime reasoning when the text
backend exposes it separately from the final assistant message.
Tool Calls
response = client.chat_completion({
"model": "qwen3.6-27b",
"messages": [{"role": "user", "content": "What time is it?"}],
"tools": [
{
"type": "function",
"function": {
"name": "get_time",
"description": "Return the current local time.",
"parameters": {"type": "object", "properties": {}},
},
}
],
"toolChoice": "auto",
"parallelToolCalls": True,
})
Routing
from modis_client import RoutingClient
client = RoutingClient(
nodes=[
"http://msn-a:8000",
{"base_url": "http://msn-b:8000", "priority": 5, "node_id": "msn-b"},
],
strategy="balanced",
)
Pydantic Models
Responses are returned as Pydantic models by default. Use to_wire() or
model_dump(by_alias=True, exclude_none=True) when you need the MoDIS JSON
shape:
wire_payload = response.to_wire()
ModisClientError is raised for transport failures, typed MoDIS errors, and
stream protocol errors.
Development
python -m pip install -e ".[dev]"
pytest
ruff check .
ruff format --check .
mypy src
python -m build
Live Service Tests
Live tests are opt-in and require a running modis-service or compatible
MoDIS endpoint. They are skipped unless MODIS_LIVE_SERVICE_URL is set.
MODIS_LIVE_SERVICE_URL=http://localhost:8000 \
MODIS_LIVE_TEXT_MODEL=qwen3.6-27b \
MODIS_LIVE_IMAGE_MODEL=qwen-image \
.venv/bin/python -m pytest -m live
Optional variables:
| Variable | Default | Purpose |
|---|---|---|
MODIS_LIVE_TIMEOUT |
900 |
Per-request timeout in seconds. |
MODIS_LIVE_CONCURRENCY |
4 |
Concurrent async text requests. |
MODIS_LIVE_IMAGE_WIDTH |
1024 |
Image test width. |
MODIS_LIVE_IMAGE_HEIGHT |
1024 |
Image test height. |
MODIS_LIVE_IMAGE_STEPS |
20 |
Image inference steps. |
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file modis_client-0.1.4.tar.gz.
File metadata
- Download URL: modis_client-0.1.4.tar.gz
- Upload date:
- Size: 32.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1bbeb08d4ee4833c88b1db1a347f2a5c07388ecf673872cec69e095cdad05745
|
|
| MD5 |
4e31fff044139041bcb038234b560feb
|
|
| BLAKE2b-256 |
e47e17fe175db5cbff7c9c682812402815858fad3a0701fc6a5d8d15fd7eae78
|
File details
Details for the file modis_client-0.1.4-py3-none-any.whl.
File metadata
- Download URL: modis_client-0.1.4-py3-none-any.whl
- Upload date:
- Size: 32.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
683050173404019601166e127a26e8d268ebf0f6aedb794633b097690082ac2c
|
|
| MD5 |
e491976abcb6a7ad5827c61098ccbb60
|
|
| BLAKE2b-256 |
d5ec085168d4c66d8f2fba58d6a66a28575a903a90ead6074875b56c041039b7
|