Skip to main content

model-compose: Declarative AI Workflow Orchestrator

Project description


model-compose

Deploy production-ready AI services in minutes.

One YAML file. Any model. Any protocol. Any runtime. Build chat APIs, RAG pipelines, autonomous agents, and MCP servers without writing application code — then deploy the same file anywhere, like docker-compose.

AI systems should not be locked into a single provider, runtime, or cloud. model-compose is built on four principles:

  • Composable — Models, agents, workflows, tools, memory, and protocols are interchangeable building blocks.
  • Portable — Define your AI system once, deploy anywhere without re-engineering.
  • Hybrid-First — Bridge cloud APIs and local models on your own terms.
  • Stream-Native — Data flows through workflows as it arrives — tokens, audio, frames, and events as first-class values.

Quick Start

Install with pip:

pip install model-compose

Or with uv:

uv pip install model-compose

Create model-compose.yml:

controller:
  adapter:
    type: http-server
    port: 8080
  webui:
    port: 8081

workflow:
  job:
    component: chatgpt
    input:
      prompt: ${input.prompt}

component:
  id: chatgpt
  type: http-client
  base_url: https://api.openai.com/v1
  action:
    path: /chat/completions
    method: POST
    headers:
      Authorization: Bearer ${env.OPENAI_API_KEY}
    body:
      model: gpt-4o
      messages:
        - role: user
          content: ${input.prompt}

Run it:

export OPENAI_API_KEY=your-key
model-compose up

That's it. You're serving GPT-4o at http://localhost:8080 with a web UI at http://localhost:8081. No application code. No framework boilerplate. Same file runs locally, in Docker, or in production.


What You Can Build

Here's what a single YAML file can serve today — just a few examples.

🤖 Autonomous Agents

Build a ReAct agent that plans, uses tools, and completes multi-step tasks — declaratively.

component:
  id: research-agent
  type: agent
  tools: [search-web, fetch-page]
  max_iteration_count: 10
  action:
    model:
      component: chatgpt
    system_prompt: You are a web research assistant.
    user_prompt: ${input.question}

See simple agents like a code reviewer, a RAG assistant, and a web researcher in agents/.

🔍 RAG Pipelines

Compose embedding, vector search, and generation into a single workflow — no glue code.

workflow:
  jobs:
    - id: embed
      component: embedder
      input: { text: ${input.query} }

    - id: retrieve
      component: knowledge
      action: search
      input: { vector: ${jobs.embed.output} }

    - id: answer
      component: chatgpt
      input:
        context: ${jobs.retrieve.output}
        question: ${input.query}

Native drivers ship for Chroma, Milvus, Qdrant, FAISS, Neo4j, ArangoDB, and Redis.

🌐 MCP Servers

Turn any workflow into an MCP server that Claude, ChatGPT, or Cursor can use — one line change.

controller:
  adapter:
    type: mcp-server   # ← was: http-server
    port: 8080

Full examples live in mcp-servers/, including a Slack bot MCP.

⚡ Streaming Multi-Modal Workflows

Stream tokens, audio chunks, and video frames end-to-end — first-class across every stage.

workflow:
  job:
    component: chatgpt
    output: ${output as sse-text}

component:
  id: chatgpt
  type: http-client
  action:
    body: { stream: true, ... }
    stream_format: json
    output: ${response[].choices[0].delta.content}

Real-time TTS, video-to-frames, and live chat examples live under data-streaming/ and showcase/.


From Development to Production

The same YAML that runs on your laptop scales without a rewrite.

1. Develop locally

model-compose up

Runs on your machine with a Gradio web UI at :8081 — perfect for iteration.

2. Deploy as a container

Add a runtime: block. Same file, same behavior:

controller:
  runtime:
    type: docker
    image: my-ai-service:latest
    ports: [ "8080:8080" ]

3. Scale horizontally

Add a queue. Dispatchers accept jobs, subscribers process them across N machines:

controller:
  adapter: { type: http-server, port: 8080 }
  queue:
    driver: redis
    host: redis.internal
    name: my-queue

No shared filesystem. No code changes. Just add more subscribers to scale.


Why model-compose?

model-compose Managed APIs (OpenAI, etc.) Code Frameworks (LangChain, etc.)
Time to first API Minutes (one YAML) Hours (SDK + server code) Days (framework + integration)
Provider Coupling Multi-provider via config Single provider per SDK Multi-provider via abstractions
Code Coupling Declarative YAML — no application code Application code required Framework-specific code required
Infrastructure Control Full Sovereignty Provider-controlled Heavy Abstraction
Runtime Flexibility Hybrid-First (Local + Cloud) Cloud Only Complex to customize
Protocol Support HTTP / WebSocket / MCP Provider-specific Limited
Data Streaming First-class across all stages Response-only (SSE tokens) Framework-wrapped generators
Deployment Docker / Native / Virtualenv / Process Provider-managed Manual integration

Highlights

  • Any model, anywhere — HuggingFace, vLLM, llama.cpp locally, or OpenAI/Anthropic/Google/xAI via HTTP
  • Agents in YAML — ReAct loops, tool use, multi-step reasoning — no code
  • Human-in-the-loop — pause workflows for approval, resume from CLI/UI/API
  • 20+ components — models, agents, HTTP/WebSocket clients, vector/graph stores, shell, browsers, and more
  • Any protocol — HTTP REST, WebSocket, or MCP with one line
  • Any runtime — Docker, native, virtualenv, process, embedded — switch in one line
  • Distributed — Redis queue dispatch for horizontal scaling
  • Instant Web UI — Gradio-powered UI in 2 lines of YAML
  • Streaming everywhere — SSE, WebSocket, and inter-job streams as first-class values

Examples

Browse examples by category:

Category What's inside
agents/ Code reviewer, RAG assistant, Web researcher, Web page analyzer, ...
showcase/ End-to-end pipelines: disk analysis, face-based scene search, real-time TTS
model-providers/ OpenAI, Anthropic, xAI, Google, ElevenLabs, vLLM
model-tasks/ Local chat, embedding, TTS, VLM, face embedding, ...
mcp-servers/ Build MCP servers exposed to Claude, Cursor, ChatGPT
workflow-queue/ Redis-backed distributed dispatch (streaming + non-streaming)
data-streaming/ Video-to-frames, YouTube live chat, streaming inputs
integrations/ Vector/graph/KV stores, search engines, channels, tunnels

Browse the full catalog in examples/README.md.


Architecture

Protocol adapters → Composition engine → Runtime executors

Architecture Diagram


Contributing

We welcome all contributions — bug fixes, docs improvements, new examples.

git clone https://github.com/hanyeol/model-compose.git
cd model-compose
pip install -e .

See CONTRIBUTING if available, or open a PR directly.


License

MIT License © 2025-2026 Hanyeol Cho.


Contact

Have questions, ideas, or feedback? Open an issue or start a discussion on GitHub Discussions.

Project details


Release history Release notifications | RSS feed

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

model_compose-0.4.82.tar.gz (445.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

model_compose-0.4.82-py3-none-any.whl (881.9 kB view details)

Uploaded Python 3

File details

Details for the file model_compose-0.4.82.tar.gz.

File metadata

  • Download URL: model_compose-0.4.82.tar.gz
  • Upload date:
  • Size: 445.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.11.4

File hashes

Hashes for model_compose-0.4.82.tar.gz
Algorithm Hash digest
SHA256 5ea3380b89881deae2031323e9ddd36a9f2752903949a82a002a6c5cae531a58
MD5 17180c0e0f811907b8d0c5c1b8f63130
BLAKE2b-256 e9ef10d6cd896b55bfb80d5cce548a294a6eeb2ae96f202b0d0e16b133be3000

See more details on using hashes here.

File details

Details for the file model_compose-0.4.82-py3-none-any.whl.

File metadata

  • Download URL: model_compose-0.4.82-py3-none-any.whl
  • Upload date:
  • Size: 881.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.11.4

File hashes

Hashes for model_compose-0.4.82-py3-none-any.whl
Algorithm Hash digest
SHA256 6e8438543cde066ecb6ffb09f543f3183be9173fb76976ea99790686889ea75a
MD5 0acd364c6e72a7a5b5e634aa31bf276a
BLAKE2b-256 dcc76ae442326c095600b2173a1f9f01cb2aa8e27f7dd7b32dabda0ad4f6513a

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page