MMSP Python Implementation
This document demonstrates how to use AutoLLMClient for unified LLM interactions in MMSP.
Building
make install # Install dependencies
make build # Build Python package
make lint # Run ruff linter
make test # Run tests
AutoLLMClient Overview
AutoLLMClient is a stateful client that automatically routes requests to the appropriate model-specific implementation. It maintains conversation history and provides a unified interface for different LLM providers.
Initialization
Create a client by specifying the model name:
from mmsp import AutoLLMClient
# The official OpenAI client, named by the model id's family
client = AutoLLMClient(model="gpt-5.5")
# The same, spelled out, with the key given in code
client = AutoLLMClient(model="gpt-5.5", client_type="openai-official", api_key="your-openai-api-key")
# A compatible client, for any endpoint that serves OpenAI Chat Completions
client = AutoLLMClient(
model="custom-model", client_type="openai-chat", base_url="http://127.0.0.1:8000/v1/", api_key="none"
)
# Gemini on Google Vertex AI: the service-account JSON key is the API key
client = AutoLLMClient(model="gemini-3.8-flash", api_key=open("service-account.json").read())
client_type names one of the official clients (openai-official, anthropic-official, gemini-official, zai-official, moonshot-official, deepseek-official, minimax-official) or one of the compatible clients (openai-responses, openai-chat, openai-chat-vllm-adapter, openai-embedding, ant-messages, gemini-generate-content). It may be omitted for a model id that begins with a known family (gpt-, text-embedding-, claude-, gemini-, glm-, kimi-, deepseek-, minimax-), which names its official client; any other id raises and asks for one.
A Vertex AI service-account key is served through generateContent, because Vertex AI's Interactions endpoint serves none of the Gemini models; any other Gemini key uses the Interactions API. client_type="gemini-generate-content" names generateContent explicitly, for gateways that proxy it.
Core Methods
streaming_response
Stateless method that requires passing the full message history on each call:
import asyncio
from mmsp import AutoLLMClient
async def main():
client = AutoLLMClient(model="gpt-5.5")
async for event in client.streaming_response(
messages=[{"role": "user", "content_items": [{"type": "text.done", "text": "Hello!"}]}], config={}
):
print(event)
asyncio.run(main())
Both streaming methods yield delta events, each carrying exactly one content item, then exactly one stop event, always last, carrying usage_metadata and finish_reason. Each item streams as one or more .delta fragments (text.delta, tool_call.delta, …) followed by its complete .done item (text.done, tool_call.done, …); items never interleave.
streaming_response_stateful
Stateful method that maintains conversation history internally:
import asyncio
from mmsp import AutoLLMClient
async def main():
client = AutoLLMClient(model="gpt-5.5")
# First message
async for event in client.streaming_response_stateful(
message={"role": "user", "content_items": [{"type": "text.done", "text": "My name is Alice"}]}, config={}
):
print(event)
# Second message - history is maintained automatically
async for event in client.streaming_response_stateful(
message={"role": "user", "content_items": [{"type": "text.done", "text": "What's my name?"}]}, config={}
):
print(event)
asyncio.run(main())
get_history
Retrieve the conversation history:
# Get all messages in the conversation
history = client.get_history()
print(f"Total messages: {len(history)}")
for msg in history:
print(f"Role: {msg['role']}")
print(f"Content: {msg['content_items']}")
clear_history
Clear the conversation history:
# Clear all conversation history
client.clear_history()
# Verify history is empty
assert len(client.get_history()) == 0
set_history
Replace the conversation history with a copy of the provided list:
# Save current history
saved_history = client.get_history()
# ... do other things, then restore
client.set_history(saved_history)
# Verify history was replaced
assert len(client.get_history()) == len(saved_history)
Tool Calling
When using tools, you must handle tool_call_id correctly:
import asyncio
import json
from mmsp import AutoLLMClient
def get_weather(location: str) -> str:
"""Mock function to get weather."""
return f"Temperature in {location}: 22°C"
async def main():
# Define tool
weather_function = {
"name": "get_weather",
"description": "Gets the current weather for a given location.",
"parameters": {
"type": "object",
"properties": {"location": {"type": "string", "description": "The city name"}},
"required": ["location"],
},
}
client = AutoLLMClient(model="gpt-5.5")
config = {"tools": [weather_function]}
# User asks about weather
events = []
async for event in client.streaming_response_stateful(
message={"role": "user", "content_items": [{"type": "text.done", "text": "What's the weather in London?"}]},
config=config,
):
events.append(event)
# Read the complete call from its tool_call.done item; tool_call.delta items are fragments
tool_call = None
for event in events:
for item in event["content_items"]:
if item["type"] == "tool_call.done":
tool_call = item
break
if tool_call:
break
# Execute function and send result back with tool_call_id
if tool_call:
result = get_weather(**tool_call["arguments"])
# IMPORTANT: Include tool_call_id in the tool response
async for event in client.streaming_response_stateful(
message={
"role": "user",
"content_items": [
{
"type": "tool_result.done",
"text": result,
"tool_call_id": tool_call["tool_call_id"], # Required for tool responses
}
],
},
config=config,
):
print(event)
asyncio.run(main())
Message Format
UniMessage Structure
{
"role": "user" | "assistant",
"content_items": [
{"type": "text.done", "text": "Hello"},
{"type": "image_url.done", "image_url": "https://..."},
{
"type": "tool_call.done",
"name": "get_weather",
"arguments": {"location": "London"},
"tool_call_id": "call_abc123",
},
],
}
Messages hold complete items only, typed with a .done suffix. Item types without the suffix, saved before 0.5.0, are still accepted with a deprecation warning until 0.6.0; normalize_legacy_messages(messages) converts stored messages.
Tool Response with tool_call_id
When responding to a tool call, include the tool_call_id in the result content item:
{
"role": "user",
"content_items": [
{
"type": "tool_result.done",
"text": "London is 22°C today.",
"tool_call_id": "call_abc123", # From the tool_call.done item
}
],
}
Configuration Options
from mmsp import PromptCaching, ThinkingLevel
config = {
"max_tokens": 500,
"temperature": 1.0,
"tools": [tool_definition],
"thinking_summary": True,
"thinking_level": ThinkingLevel.HIGH,
"tool_choice": "auto", # "auto", "required", "none", or ["tool_name"]
"system_prompt": "You are a helpful assistant",
"prompt_caching": PromptCaching.ENABLE,
"trace_id": "agent1/conversation_001", # Optional: save conversation trace
}
Conversation Tracing
MMSP provides a built-in Tracer to save and browse conversation history. When you specify a trace_id in the config, conversations are automatically saved to both JSON and TXT formats.
Basic Usage
from mmsp import AutoLLMClient
client = AutoLLMClient(model="gpt-5.5")
# Add trace_id to config
config = {"trace_id": "agent1/conversation_001"}
async for event in client.streaming_response_stateful(
message={"role": "user", "content_items": [{"type": "text.done", "text": "Hello"}]}, config=config
):
pass # Conversation is automatically saved
The default cache directory is cache, you can change it by setting MMSP_CACHE_DIR environment variable.
This creates two files in the cache directory:
cache/agent1/conversation_001.json- Structured data with full history and configcache/agent1/conversation_001.txt- Human-readable conversation format
Browsing Traces with Web Interface
Start a web server to browse and view saved conversations:
from mmsp.integration.tracer import Tracer
# Start web server
Tracer("path/to/cache").start_web_server(host="127.0.0.1", port=25750)
Or use the CLI:
python -m mmsp.integration.tracer --cache_dir ./cache --host 127.0.0.1 --port 25750
Then visit http://127.0.0.1:25750 in your browser to browse saved conversations.
Test with Playground
Start a web server to test with the playground:
from mmsp.integration.playground import start_playground_server
start_playground_server()
Or use the CLI:
python -m mmsp.integration.playground --host 127.0.0.1 --port 25751
Then visit http://127.0.0.1:25751 in your browser to test with the playground.
The integrated tracer is available at http://127.0.0.1:25751/tracer/.
Metadata
Release files for mmsp 0.5.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| mmsp-0.5.0.tar.gz | 113.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| mmsp-0.5.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 266.6 kB
Release files / mmsp-0.5.0.tar.gz
| Download URL | mmsp-0.5.0.tar.gz |
|---|---|
| Size | 113.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
9fba4ccfdb92f0d36093f917e7a7aed36b4c0fb7a9d491e461ce7d81dc1cf129
|
|
BLAKE2b-256 checksum How to use checksums |
950e1996bfae0dcdd5e1a43a4359746c2bc5df6b3a8d74dd48352f57ba0ad14b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.9.11 {"installer":{"name":"uv","version":"0.9.11"},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / mmsp-0.5.0-py3-none-any.whl
| Download URL | mmsp-0.5.0-py3-none-any.whl |
|---|---|
| Size | 153.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
1525c28890549570b7d3febbe03f9008d9187d6f5a0ef710bbefe7d5d00ded42
|
|
BLAKE2b-256 checksum How to use checksums |
7c6e0eca5b568647ffe40df802389f4be192577cffd6c35961e0b979e873510f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.9.11 {"installer":{"name":"uv","version":"0.9.11"},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|