Skip to main content

Python client for MoDIS Service Nodes

Project description

MoDIS Python Client

modis-client is the Python client package for MoDIS Service Nodes (MSNs). It exposes MoDIS-native request and response objects for direct Python callers.

The PyPI distribution name is modis-client. pymodis and modis are already used by unrelated packages, while modis-client currently appears available.

Install

From a checkout:

python -m pip install -e .

The import package is modis_client:

from modis_client import SingleNodeClient

with SingleNodeClient(base_url="http://localhost:8000") as client:
    response = client.chat_completion({
        "model": "qwen3.6-27b",
        "messages": [{"role": "user", "content": "Hello MoDIS."}],
    })

print(response.content.content)

The package supports sync and async single-node clients, routing clients, batch requests, explicit streaming methods, MoDIS tool-call fields, and Pydantic v2 models with MoDIS wire aliases.

The default client timeout is 900 seconds because a MoDIS request may include a cold model load before inference begins. Pass timeout= to any client constructor when a shorter or longer budget is required.

Async Client

from modis_client import AsyncSingleNodeClient

async with AsyncSingleNodeClient(base_url="http://localhost:8000") as client:
    response = await client.responses({
        "model": "qwen3.6-27b",
        "input": "Summarize MoDIS in one sentence.",
    })

Batch Requests

responses = client.chat_completion([
    {"model": "qwen3.6-27b", "messages": [{"role": "user", "content": "one"}]},
    {"model": "qwen3.6-27b", "messages": [{"role": "user", "content": "two"}]},
])

Streaming

for event in client.chat_completion_stream({
    "model": "qwen3.6-27b",
    "messages": [{"role": "user", "content": "Count to three."}],
}):
    if event.event == "delta":
        print(event.payload)

Image streams preserve progress and final image payload events:

for event in client.image_generation_stream({
    "model": "qwen-image",
    "prompt": "A clean product photo of a compact AI workstation.",
    "numOutputs": 1,
}):
    if event.event == "progress":
        print(event.payload["step"], event.payload.get("total"))
    if event.event == "completed":
        images = event.payload["images"]

Text To Speech

Non-streaming speech requests return a Pydantic response with audio_b64, audio_url, and mime_type:

import base64

response = client.text_to_speech({
    "model": "qwen-tts-custom-voice",
    "text": "Hello from MoDIS.",
    "speaker": "Vivian",
    "responseFormat": "wav",
})

audio_bytes = base64.b64decode(response.audio_b64)

PCM streaming is exposed separately because /audio/speech returns raw audio bytes, not SSE events, when stream: true is used:

with open("speech.pcm", "wb") as output:
    for chunk in client.text_to_speech_pcm_stream({
        "model": "qwen-tts-custom-voice",
        "text": "Hello from MoDIS.",
        "speaker": "Vivian",
    }):
        output.write(chunk)

Reference voices can be passed per request with referenceAudio and referenceText. Use hosted URLs when the vLLM-Omni server can fetch the audio, or inline data URLs/base64 when the reference is not hosted.

Speech To Text

Non-streaming ASR requests call /audio/transcriptions and return a transcript in response.text:

response = client.speech_to_text({
    "model": "qwen3-asr",
    "audioUrl": "https://example.com/audio.wav",
    "language": "en",
})

print(response.text)

The typed SpeechToTextRequest model also accepts audioB64, audioPath, audio, url, prompt, responseFormat, and timeoutSeconds. The current MoDIS ASR route is non-streaming; uploaded-audio output streaming and realtime WebSocket transcription are service follow-ups.

Structured Output

response = client.chat_completion({
    "model": "qwen3.6-27b",
    "messages": [{"role": "user", "content": "Return one rating."}],
    "format": {
        "title": "Rating",
        "type": "object",
        "properties": {"score": {"type": "integer"}},
        "required": ["score"],
    },
})

format: "json" requests JSON-object output. A JSON Schema object is forwarded to MoDIS text runtimes as a structured response schema.

Reasoning Output

response = client.chat_completion({
    "model": "qwen3.6-27b",
    "messages": [{"role": "user", "content": "Solve 25^(1/2)."}],
    "includeReasoning": True,
})

print(response.content.reasoning)
print(response.content.content)

includeReasoning asks MoDIS to include parsed runtime reasoning when the text backend exposes it separately from the final assistant message.

Tool Calls

response = client.chat_completion({
    "model": "qwen3.6-27b",
    "messages": [{"role": "user", "content": "What time is it?"}],
    "tools": [
        {
            "type": "function",
            "function": {
                "name": "get_time",
                "description": "Return the current local time.",
                "parameters": {"type": "object", "properties": {}},
            },
        }
    ],
    "toolChoice": "auto",
    "parallelToolCalls": True,
})

Routing

from modis_client import RoutingClient

client = RoutingClient(
    nodes=[
        "http://msn-a:8000",
        {"base_url": "http://msn-b:8000", "priority": 5, "node_id": "msn-b"},
    ],
    strategy="balanced",
)

Pydantic Models

Responses are returned as Pydantic models by default. Use to_wire() or model_dump(by_alias=True, exclude_none=True) when you need the MoDIS JSON shape:

wire_payload = response.to_wire()

ModisClientError is raised for transport failures, typed MoDIS errors, and stream protocol errors.

Development

python -m pip install -e ".[dev]"
pytest
ruff check .
ruff format --check .
mypy src
python -m build

Live Service Tests

Live tests are opt-in and require a running modis-service or compatible MoDIS endpoint. They are skipped unless MODIS_LIVE_SERVICE_URL is set.

MODIS_LIVE_SERVICE_URL=http://localhost:8000 \
MODIS_LIVE_TEXT_MODEL=qwen3.6-27b \
MODIS_LIVE_IMAGE_MODEL=qwen-image \
.venv/bin/python -m pytest -m live

Optional variables:

Variable Default Purpose
MODIS_LIVE_TIMEOUT 900 Per-request timeout in seconds.
MODIS_LIVE_CONCURRENCY 4 Concurrent async text requests.
MODIS_LIVE_IMAGE_WIDTH 1024 Image test width.
MODIS_LIVE_IMAGE_HEIGHT 1024 Image test height.
MODIS_LIVE_IMAGE_STEPS 20 Image inference steps.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

modis_client-0.1.4.tar.gz (32.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

modis_client-0.1.4-py3-none-any.whl (32.2 kB view details)

Uploaded Python 3

File details

Details for the file modis_client-0.1.4.tar.gz.

File metadata

  • Download URL: modis_client-0.1.4.tar.gz
  • Upload date:
  • Size: 32.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.9

File hashes

Hashes for modis_client-0.1.4.tar.gz
Algorithm Hash digest
SHA256 1bbeb08d4ee4833c88b1db1a347f2a5c07388ecf673872cec69e095cdad05745
MD5 4e31fff044139041bcb038234b560feb
BLAKE2b-256 e47e17fe175db5cbff7c9c682812402815858fad3a0701fc6a5d8d15fd7eae78

See more details on using hashes here.

File details

Details for the file modis_client-0.1.4-py3-none-any.whl.

File metadata

  • Download URL: modis_client-0.1.4-py3-none-any.whl
  • Upload date:
  • Size: 32.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.9

File hashes

Hashes for modis_client-0.1.4-py3-none-any.whl
Algorithm Hash digest
SHA256 683050173404019601166e127a26e8d268ebf0f6aedb794633b097690082ac2c
MD5 e491976abcb6a7ad5827c61098ccbb60
BLAKE2b-256 d5ec085168d4c66d8f2fba58d6a66a28575a903a90ead6074875b56c041039b7

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page