Skip to main content

MMSP Python Implementation

This document demonstrates how to use AutoLLMClient for unified LLM interactions in MMSP.

Building

make install  # Install dependencies
make build    # Build Python package
make lint     # Run ruff linter
make test     # Run tests

AutoLLMClient Overview

AutoLLMClient is a stateful client that automatically routes requests to the appropriate model-specific implementation. It maintains conversation history and provides a unified interface for different LLM providers.

Initialization

Create a client by specifying the model name:

from mmsp import AutoLLMClient

# The official OpenAI client, named by the model id's family
client = AutoLLMClient(model="gpt-5.5")

# The same, spelled out, with the key given in code
client = AutoLLMClient(model="gpt-5.5", client_type="openai-official", api_key="your-openai-api-key")

# A compatible client, for any endpoint that serves OpenAI Chat Completions
client = AutoLLMClient(
    model="custom-model", client_type="openai-chat", base_url="http://127.0.0.1:8000/v1/", api_key="none"
)

# Gemini on Google Vertex AI: the service-account JSON key is the API key
client = AutoLLMClient(model="gemini-3.8-flash", api_key=open("service-account.json").read())

client_type names one of the official clients (openai-official, anthropic-official, gemini-official, zai-official, moonshot-official, deepseek-official, minimax-official) or one of the compatible clients (openai-responses, openai-chat, openai-chat-vllm-adapter, openai-embedding, ant-messages, gemini-generate-content). It may be omitted for a model id that begins with a known family (gpt-, text-embedding-, claude-, gemini-, glm-, kimi-, deepseek-, minimax-), which names its official client; any other id raises and asks for one.

A Vertex AI service-account key is served through generateContent, because Vertex AI's Interactions endpoint serves none of the Gemini models; any other Gemini key uses the Interactions API. client_type="gemini-generate-content" names generateContent explicitly, for gateways that proxy it.

Core Methods

streaming_response

Stateless method that requires passing the full message history on each call:

import asyncio
from mmsp import AutoLLMClient


async def main():
    client = AutoLLMClient(model="gpt-5.5")

    async for event in client.streaming_response(
        messages=[{"role": "user", "content_items": [{"type": "text.done", "text": "Hello!"}]}], config={}
    ):
        print(event)


asyncio.run(main())

Both streaming methods yield delta events, each carrying exactly one content item, then exactly one stop event, always last, carrying usage_metadata and finish_reason. Each item streams as one or more .delta fragments (text.delta, tool_call.delta, …) followed by its complete .done item (text.done, tool_call.done, …); items never interleave.

streaming_response_stateful

Stateful method that maintains conversation history internally:

import asyncio
from mmsp import AutoLLMClient


async def main():
    client = AutoLLMClient(model="gpt-5.5")

    # First message
    async for event in client.streaming_response_stateful(
        message={"role": "user", "content_items": [{"type": "text.done", "text": "My name is Alice"}]}, config={}
    ):
        print(event)

    # Second message - history is maintained automatically
    async for event in client.streaming_response_stateful(
        message={"role": "user", "content_items": [{"type": "text.done", "text": "What's my name?"}]}, config={}
    ):
        print(event)


asyncio.run(main())

get_history

Retrieve the conversation history:

# Get all messages in the conversation
history = client.get_history()
print(f"Total messages: {len(history)}")

for msg in history:
    print(f"Role: {msg['role']}")
    print(f"Content: {msg['content_items']}")

clear_history

Clear the conversation history:

# Clear all conversation history
client.clear_history()

# Verify history is empty
assert len(client.get_history()) == 0

set_history

Replace the conversation history with a copy of the provided list:

# Save current history
saved_history = client.get_history()

# ... do other things, then restore
client.set_history(saved_history)

# Verify history was replaced
assert len(client.get_history()) == len(saved_history)

Tool Calling

When using tools, you must handle tool_call_id correctly:

import asyncio
import json
from mmsp import AutoLLMClient


def get_weather(location: str) -> str:
    """Mock function to get weather."""
    return f"Temperature in {location}: 22°C"


async def main():
    # Define tool
    weather_function = {
        "name": "get_weather",
        "description": "Gets the current weather for a given location.",
        "parameters": {
            "type": "object",
            "properties": {"location": {"type": "string", "description": "The city name"}},
            "required": ["location"],
        },
    }

    client = AutoLLMClient(model="gpt-5.5")
    config = {"tools": [weather_function]}

    # User asks about weather
    events = []
    async for event in client.streaming_response_stateful(
        message={"role": "user", "content_items": [{"type": "text.done", "text": "What's the weather in London?"}]},
        config=config,
    ):
        events.append(event)

    # Read the complete call from its tool_call.done item; tool_call.delta items are fragments
    tool_call = None
    for event in events:
        for item in event["content_items"]:
            if item["type"] == "tool_call.done":
                tool_call = item
                break

        if tool_call:
            break

    # Execute function and send result back with tool_call_id
    if tool_call:
        result = get_weather(**tool_call["arguments"])

        # IMPORTANT: Include tool_call_id in the tool response
        async for event in client.streaming_response_stateful(
            message={
                "role": "user",
                "content_items": [
                    {
                        "type": "tool_result.done",
                        "text": result,
                        "tool_call_id": tool_call["tool_call_id"],  # Required for tool responses
                    }
                ],
            },
            config=config,
        ):
            print(event)


asyncio.run(main())

Message Format

UniMessage Structure

{
    "role": "user" | "assistant",
    "content_items": [
        {"type": "text.done", "text": "Hello"},
        {"type": "image_url.done", "image_url": "https://..."},
        {
            "type": "tool_call.done",
            "name": "get_weather",
            "arguments": {"location": "London"},
            "tool_call_id": "call_abc123",
        },
    ],
}

Messages hold complete items only, typed with a .done suffix. Item types without the suffix, saved before 0.5.0, are still accepted with a deprecation warning until 0.6.0; normalize_legacy_messages(messages) converts stored messages.

Tool Response with tool_call_id

When responding to a tool call, include the tool_call_id in the result content item:

{
    "role": "user",
    "content_items": [
        {
            "type": "tool_result.done",
            "text": "London is 22°C today.",
            "tool_call_id": "call_abc123",  # From the tool_call.done item
        }
    ],
}

Configuration Options

from mmsp import PromptCaching, ThinkingLevel

config = {
    "max_tokens": 500,
    "temperature": 1.0,
    "tools": [tool_definition],
    "thinking_summary": True,
    "thinking_level": ThinkingLevel.HIGH,
    "tool_choice": "auto",  # "auto", "required", "none", or ["tool_name"]
    "system_prompt": "You are a helpful assistant",
    "prompt_caching": PromptCaching.ENABLE,
    "trace_id": "agent1/conversation_001",  # Optional: save conversation trace
}

Conversation Tracing

MMSP provides a built-in Tracer to save and browse conversation history. When you specify a trace_id in the config, conversations are automatically saved to both JSON and TXT formats.

Basic Usage

from mmsp import AutoLLMClient

client = AutoLLMClient(model="gpt-5.5")

# Add trace_id to config
config = {"trace_id": "agent1/conversation_001"}

async for event in client.streaming_response_stateful(
    message={"role": "user", "content_items": [{"type": "text.done", "text": "Hello"}]}, config=config
):
    pass  # Conversation is automatically saved

The default cache directory is cache, you can change it by setting MMSP_CACHE_DIR environment variable.

This creates two files in the cache directory:

  • cache/agent1/conversation_001.json - Structured data with full history and config
  • cache/agent1/conversation_001.txt - Human-readable conversation format

Browsing Traces with Web Interface

Start a web server to browse and view saved conversations:

from mmsp.integration.tracer import Tracer

# Start web server
Tracer("path/to/cache").start_web_server(host="127.0.0.1", port=25750)

Or use the CLI:

python -m mmsp.integration.tracer --cache_dir ./cache --host 127.0.0.1 --port 25750

Then visit http://127.0.0.1:25750 in your browser to browse saved conversations.

Test with Playground

Start a web server to test with the playground:

from mmsp.integration.playground import start_playground_server

start_playground_server()

Or use the CLI:

python -m mmsp.integration.playground --host 127.0.0.1 --port 25751

Then visit http://127.0.0.1:25751 in your browser to test with the playground. The integrated tracer is available at http://127.0.0.1:25751/tracer/.

Metadata

Release files for mmsp 0.5.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mmsp 0.5.0
File Size Uploaded
mmsp-0.5.0.tar.gz 113.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for mmsp 0.5.0
File Interpreter ABI Platform
mmsp-0.5.0-py3-none-any.whl Python 3 none any Details

Total release size: 266.6 kB

Release files / mmsp-0.5.0.tar.gz

Download URL mmsp-0.5.0.tar.gz
Size 113.6 kB
Tags Source
SHA-256 checksum
How to use checksums
9fba4ccfdb92f0d36093f917e7a7aed36b4c0fb7a9d491e461ce7d81dc1cf129
BLAKE2b-256 checksum
How to use checksums
950e1996bfae0dcdd5e1a43a4359746c2bc5df6b3a8d74dd48352f57ba0ad14b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.9.11 {"installer":{"name":"uv","version":"0.9.11"},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / mmsp-0.5.0-py3-none-any.whl

Download URL mmsp-0.5.0-py3-none-any.whl
Size 153.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1525c28890549570b7d3febbe03f9008d9187d6f5a0ef710bbefe7d5d00ded42
BLAKE2b-256 checksum
How to use checksums
7c6e0eca5b568647ffe40df802389f4be192577cffd6c35961e0b979e873510f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.9.11 {"installer":{"name":"uv","version":"0.9.11"},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

0.5.1

2 release files

This release

0.5.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page