Skip to main content

Python client for MoDIS Service Nodes

Project description

MoDIS Python Client

modis-client is the Python client package for MoDIS Service Nodes (MSNs). It exposes MoDIS-native request and response objects for direct Python callers.

The PyPI distribution name is modis-client. pymodis and modis are already used by unrelated packages, while modis-client currently appears available.

Install

From a checkout:

python -m pip install -e .

The import package is modis_client:

from modis_client import SingleNodeClient

with SingleNodeClient(base_url="http://localhost:8000") as client:
    response = client.chat_completion({
        "model": "qwen3.6-27b",
        "messages": [{"role": "user", "content": "Hello MoDIS."}],
    })

print(response.content.content)

The package supports sync and async single-node clients, routing clients, batch requests, explicit streaming methods, MoDIS tool-call fields, and Pydantic v2 models with MoDIS wire aliases.

The default client timeout is 900 seconds because a MoDIS request may include a cold model load before inference begins. Pass timeout= to any client constructor when a shorter or longer budget is required.

Async Client

from modis_client import AsyncSingleNodeClient

async with AsyncSingleNodeClient(base_url="http://localhost:8000") as client:
    response = await client.responses({
        "model": "qwen3.6-27b",
        "input": "Summarize MoDIS in one sentence.",
    })

Batch Requests

responses = client.chat_completion([
    {"model": "qwen3.6-27b", "messages": [{"role": "user", "content": "one"}]},
    {"model": "qwen3.6-27b", "messages": [{"role": "user", "content": "two"}]},
])

Streaming

for event in client.chat_completion_stream({
    "model": "qwen3.6-27b",
    "messages": [{"role": "user", "content": "Count to three."}],
}):
    if event.event == "delta":
        print(event.payload)

Image streams preserve progress and final image payload events:

for event in client.image_generation_stream({
    "model": "qwen-image",
    "prompt": "A clean product photo of a compact AI workstation.",
    "numOutputs": 1,
}):
    if event.event == "progress":
        print(event.payload["step"], event.payload.get("total"))
    if event.event == "completed":
        images = event.payload["images"]

Text To Speech

Non-streaming speech requests return a Pydantic response with audio_b64, audio_url, and mime_type:

import base64

response = client.text_to_speech({
    "model": "qwen-tts-custom-voice",
    "text": "Hello from MoDIS.",
    "speaker": "Vivian",
    "responseFormat": "wav",
})

audio_bytes = base64.b64decode(response.audio_b64)

PCM streaming is exposed separately because /audio/speech returns raw audio bytes, not SSE events, when stream: true is used:

with open("speech.pcm", "wb") as output:
    for chunk in client.text_to_speech_pcm_stream({
        "model": "qwen-tts-custom-voice",
        "text": "Hello from MoDIS.",
        "speaker": "Vivian",
    }):
        output.write(chunk)

Reference voices can be passed per request with referenceAudio and referenceText. Use hosted URLs when the vLLM-Omni server can fetch the audio, or inline data URLs/base64 when the reference is not hosted.

Structured Output

response = client.chat_completion({
    "model": "qwen3.6-27b",
    "messages": [{"role": "user", "content": "Return one rating."}],
    "format": {
        "title": "Rating",
        "type": "object",
        "properties": {"score": {"type": "integer"}},
        "required": ["score"],
    },
})

format: "json" requests JSON-object output. A JSON Schema object is forwarded to MoDIS text runtimes as a structured response schema.

Reasoning Output

response = client.chat_completion({
    "model": "qwen3.6-27b",
    "messages": [{"role": "user", "content": "Solve 25^(1/2)."}],
    "includeReasoning": True,
})

print(response.content.reasoning)
print(response.content.content)

includeReasoning asks MoDIS to include parsed runtime reasoning when the text backend exposes it separately from the final assistant message.

Tool Calls

response = client.chat_completion({
    "model": "qwen3.6-27b",
    "messages": [{"role": "user", "content": "What time is it?"}],
    "tools": [
        {
            "type": "function",
            "function": {
                "name": "get_time",
                "description": "Return the current local time.",
                "parameters": {"type": "object", "properties": {}},
            },
        }
    ],
    "toolChoice": "auto",
    "parallelToolCalls": True,
})

Routing

from modis_client import RoutingClient

client = RoutingClient(
    nodes=[
        "http://msn-a:8000",
        {"base_url": "http://msn-b:8000", "priority": 5, "node_id": "msn-b"},
    ],
    strategy="balanced",
)

Pydantic Models

Responses are returned as Pydantic models by default. Use to_wire() or model_dump(by_alias=True, exclude_none=True) when you need the MoDIS JSON shape:

wire_payload = response.to_wire()

ModisClientError is raised for transport failures, typed MoDIS errors, and stream protocol errors.

Development

python -m pip install -e ".[dev]"
pytest
ruff check .
ruff format --check .
mypy src
python -m build

Live Service Tests

Live tests are opt-in and require a running modis-service or compatible MoDIS endpoint. They are skipped unless MODIS_LIVE_SERVICE_URL is set.

MODIS_LIVE_SERVICE_URL=http://localhost:8000 \
MODIS_LIVE_TEXT_MODEL=qwen3.6-27b \
MODIS_LIVE_IMAGE_MODEL=qwen-image \
.venv/bin/python -m pytest -m live

Optional variables:

Variable Default Purpose
MODIS_LIVE_TIMEOUT 900 Per-request timeout in seconds.
MODIS_LIVE_CONCURRENCY 4 Concurrent async text requests.
MODIS_LIVE_IMAGE_WIDTH 1024 Image test width.
MODIS_LIVE_IMAGE_HEIGHT 1024 Image test height.
MODIS_LIVE_IMAGE_STEPS 20 Image inference steps.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

modis_client-0.1.2.tar.gz (27.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

modis_client-0.1.2-py3-none-any.whl (30.9 kB view details)

Uploaded Python 3

File details

Details for the file modis_client-0.1.2.tar.gz.

File metadata

  • Download URL: modis_client-0.1.2.tar.gz
  • Upload date:
  • Size: 27.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.9

File hashes

Hashes for modis_client-0.1.2.tar.gz
Algorithm Hash digest
SHA256 03efd62c4a24086f4741aec5b351b5c9ab9d77791cfa8118cd334d206747209a
MD5 8b1c933f5e22e813f7389dfa68a4edf5
BLAKE2b-256 3c5398c0e6ba2612151da3b8ef4184134b55a401a725fcafa453bab0a7ed2a03

See more details on using hashes here.

File details

Details for the file modis_client-0.1.2-py3-none-any.whl.

File metadata

  • Download URL: modis_client-0.1.2-py3-none-any.whl
  • Upload date:
  • Size: 30.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.9

File hashes

Hashes for modis_client-0.1.2-py3-none-any.whl
Algorithm Hash digest
SHA256 ddfb11607f1b4e9180d0ae8f2478130499b7a2491a6b3434e5dbeac7d5047104
MD5 ff487e0a64862084c92ef6d690c05a1d
BLAKE2b-256 3a4d851615c1315033a0fc7cd397d696307ddc6fd5be558f1c8ff669583c6d65

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page