Skip to main content

SMG Logo

Shepherd Model Gateway

Release Docker PyPI License Docs Discord Slack Ask DeepWiki PyTorch Blog

Engine-agnostic, high-performance model-routing gateway for large-scale LLM deployments. SMG centralizes worker lifecycle management, balances traffic across self-hosted engines and cloud providers, and gives you enterprise-grade control over multi-tenancy, chat-history storage, MCP tooling, and observability — behind one unified endpoint.

SMG architecture: clients flow through the gateway layer and router layer to gRPC workers, HTTP workers, and external APIs

Why SMG?

🚀 Maximize GPU Utilization Cache-aware routing tracks each worker's KV-cache state in radix trees to reuse prefixes across SGLang, vLLM, TensorRT-LLM, TokenSpeed, and MLX — with load modeling that accounts for queued token work and KV pressure.
🔌 One API, Any Backend Route to self-hosted engines over HTTP or gRPC, or to OpenAI, Anthropic, Gemini, and xAI — plus any OpenAI-compatible endpoint — through a single unified gateway.
⚡ Built for Speed Native Rust with streaming gRPC pipelines, cached tokenization with zero-copy cache hits, prefill/decode disaggregation (including a separate encode stage for vision), and DP-aware routing for data-parallel engines.
🔒 Enterprise Control Priority admission scheduling with preemption and per-tenant controls, API-key auth with OIDC on the control plane, WebAssembly plugins for custom logic, and chat history that never leaves your infrastructure.
📊 Full Observability 90+ Prometheus metrics, OpenTelemetry tracing with W3C trace context propagated into the engines over both HTTP and gRPC, and structured JSON logs with request correlation.

API Coverage: OpenAI Chat Completions, Completions, Embeddings, Rerank, and Classify; Responses and Conversations APIs for agents; Anthropic Messages; Gemini Interactions; Realtime over WebSocket and WebRTC; audio transcription; tokenize/detokenize; and MCP tool execution with approval policies in the Responses and Messages APIs.

Quick Start

Install — pick your preferred method:

# Docker
docker pull lightseekorg/smg:latest

# Kubernetes (Helm)
helm install smg oci://ghcr.io/smg-project/charts/smg

# Python
pip install smg

# Rust (needs protoc)
cargo install smg

Run — point SMG at your inference workers:

# Single worker
smg launch --worker-urls http://localhost:8000

# Multiple workers with cache-aware routing
smg launch --worker-urls http://gpu1:8000 http://gpu2:8000 --policy cache_aware

# With high availability mesh
smg launch --worker-urls http://gpu1:8000 --enable-mesh \
  --mesh-advertise-host 10.0.0.1 --mesh-peer-urls 10.0.0.2:39527

Use — send requests to the gateway:

curl http://localhost:30000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "llama3", "messages": [{"role": "user", "content": "Hello!"}]}'

That's it. SMG is now load-balancing requests across your workers.

Supported Backends

Self-Hosted Engines vLLM · SGLang · TokenSpeed · TensorRT-LLM · MLX (Apple Silicon) · any OpenAI-compatible server (e.g. Ollama)
Cloud Providers OpenAI · Anthropic · Google Gemini · xAI · OCI Generative AI · AWS Bedrock · Azure OpenAI · any OpenAI-compatible provider (Groq, Together, …)

Features

Feature Description
10 Routing Policies cache_aware, least_load, power_of_two, consistent_hashing, prefix_hash, bucket, round_robin, random, manual, passthrough
gRPC Pipeline Native streaming gRPC to the engines with prefill/decode and encode disaggregation and DP-aware routing
Kubernetes Discovery Native pod watchers with label selectors, per-role prefill/decode/encode selectors, and router peer discovery
Model Parsers 21 tool-call parsers and 16 reasoning parsers with automatic model detection — DeepSeek, Qwen, Kimi, GLM, Llama, Mistral, Command, Nemotron, and more
MCP Integration Tool discovery and execution over stdio, SSE, and streamable HTTP, with approval policies and audit logging
High Availability Mesh networking with SWIM gossip and CRDT-replicated state for multi-node deployments
Chat History Pluggable storage with schema migrations: PostgreSQL, Oracle, Redis, or in-memory
WASM Plugins Extend request and response handling with custom WebAssembly middleware
Resilience Circuit breakers, retries with backoff and jitter, rate limiting, and priority admission scheduling

Documentation

Full documentation lives at lightseek.org/smg.

Getting Started Installation and first steps
Architecture How SMG works
Configuration CLI reference and options
API Reference OpenAI-compatible endpoints
Kubernetes Setup In-cluster discovery and production setup

Contributing

We welcome contributions! See the Contributing Guide for details.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

tokenspeed_smg-1.9.0.post20260805.tar.gz (3.3 MB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

tokenspeed_smg-1.9.0.post20260805-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (31.2 MB view details)

Uploaded CPython 3.8+manylinux: glibc 2.17+ x86-64

tokenspeed_smg-1.9.0.post20260805-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (32.8 MB view details)

Uploaded CPython 3.8+manylinux: glibc 2.17+ ARM64

File details

Details for the file tokenspeed_smg-1.9.0.post20260805.tar.gz.

File metadata

File hashes

Hashes for tokenspeed_smg-1.9.0.post20260805.tar.gz
Algorithm Hash digest
SHA256 416a062c31a0295a708cd5c244252e2e1e771bb6c368752811b9b1b3d348b5d0
MD5 bb63a22e8acbfab8ec8c0ee8f87862d9
BLAKE2b-256 9898cb54f3b60d14469dfd01003a41547a7a44160788fd0c66c314305397b63e

See more details on using hashes here.

Provenance

The following attestation bundles were made for tokenspeed_smg-1.9.0.post20260805.tar.gz:

Publisher: tokenspeed-smg.yml on lightseekorg/tokenspeed-third-party

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file tokenspeed_smg-1.9.0.post20260805-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for tokenspeed_smg-1.9.0.post20260805-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 3f2e104007fd979ad0b50ff9282eca3a164093cbd099e8682af7107b6617da5d
MD5 3cf1a97afef98c54b7b0913ffa07758f
BLAKE2b-256 a71845b0de850802a925076eefc8e452d4446912d5b7fa6b8d5424082c270188

See more details on using hashes here.

Provenance

The following attestation bundles were made for tokenspeed_smg-1.9.0.post20260805-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: tokenspeed-smg.yml on lightseekorg/tokenspeed-third-party

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file tokenspeed_smg-1.9.0.post20260805-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for tokenspeed_smg-1.9.0.post20260805-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 3460960c04afa30b1e551d69ace86c6fe24c8256138fbceb892b2c2f6fdd56d5
MD5 a486e5b52f58109fd8d6e1fb030b9680
BLAKE2b-256 03bc2f871024fa847eb5a706fb97be0479701960eee17b3600a07dabfa1f2f07

See more details on using hashes here.

Provenance

The following attestation bundles were made for tokenspeed_smg-1.9.0.post20260805-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:

Publisher: tokenspeed-smg.yml on lightseekorg/tokenspeed-third-party

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page