DistLLM — Distributed Inference Across All Your Devices
Pool GPUs from every device you own to run models no single machine can handle.
You have a gaming PC with an RTX 4090. Your laptop has an RTX 4060. Your friend has a desktop with an RTX 3080. None of you can run Llama 3.1 70B alone. Together, you can.
DistLLM splits large language models across all your devices using pipeline parallelism. Each device runs a fraction of the model layers. Automatic discovery. Auto-partitioning. Works over LAN, WiFi, or internet.
Your Laptop (RTX 4060) ────┐
│
Your Gaming PC (RTX 4090) ──┼──► DistLLM Cluster ──► Run 70B models
│
Friend's PC (RTX 3080) ────┘
Why DistLLM?
| Problem | Solution |
|---|---|
| One GPU can't run today's best models | Split across all your devices |
| Cloud inference costs thousands/month | Use the GPUs you already own |
| Data privacy concerns with cloud APIs | Your data stays on your devices |
| Slow single-device inference | Pipeline parallelism = faster generation |
| Setting up distributed systems is hard | One command to start, one to join |
Quick Start
pip install distllm
# On your main machine — start a cluster
distllm cluster start --model meta-llama/Llama-3.2-7B
# On every other machine — join the cluster
distllm cluster join
How It Works
DistLLM uses pipeline parallelism — the model is split across devices by layers:
Device 1 (Laptop): Layers 0-5 ──→ Device 2 (Desktop): Layers 6-11 ──→ Device 3 (Friend's PC): Layers 12-17
Each device runs ~6 layers → fits in 6-8GB VRAM
Combined pool → runs models up to 70B parameters
Key capabilities:
- Auto-discovery: devices find each other on the same network automatically
- Auto-partitioning: automatically assigns layers based on each device's GPU
- Node recovery: if a device disconnects, remaining nodes take over
- Straggler detection: slow nodes are detected and worked around
- WAN optimization: token accumulation for low-latency cross-internet inference
- Privacy-first: keep sensitive layers on your own devices
Installation
# Core package
pip install distllm
# With vLLM backend (recommended for NVIDIA GPUs)
pip install "distllm[vllm]"
# With llama.cpp backend (CPU, AMD, Apple Silicon)
pip install "distllm[llamacpp]"
# Development
pip install -e ".[dev]"
Key Features
- Pipeline parallelism — split any HuggingFace model across N devices
- Auto-discovery — mDNS/zeroconf device finding on LAN
- 6 backends — vLLM, llama.cpp, TensorRT-LLM, ExLlamaV2, ONNX, PyTorch
- Auth plugin — JWT authentication + RBAC role-based access control
- Health watchdog — Continuous node health monitoring with auto circuit-breaking
- Semantic caching — Deduplicate repeated prompts with embedding similarity
- Token streaming —
generate_stream()for real-time token-by-token responses - Config validation — Cross-field validation catches invalid combinations at load time
- Circuit breaker — Graduated backpressure for load shedding
- Auto-partitioning — hardware-aware DP solver for optimal layer assignment
- Node recovery — checkpoint-based recovery when nodes disconnect
- Straggler detection — statistical outlier detection for slow nodes
- Dynamic rebalancing — redistribute layers when nodes join/leave
- P2P KV cache gossip — CRDT-based cache sharing between nodes
- Wide-area support — token accumulation for internet-scale inference
- Quantization — 4-bit/8-bit to fit larger models on consumer GPUs
- OpenAI-compatible API — use any OpenAI client to send requests
- Full observability — Prometheus metrics, OTel tracing, structured logging
- Interactive chat —
distllm chatfor CLI-based interaction - Auth plugin — JWT authentication with RBAC role-based access control
- Health watchdog — continuous node health monitoring with automatic failover
- Semantic cache — caching plugin with deduplication for repeated prompts
- Token streaming —
generate_stream()SDK method for token-by-token responses - Health endpoints —
/healthz(liveness) and/readyz(readiness) for Kubernetes probes
CLI Commands
# Start a coordinator node
distllm-coordinator --model meta-llama/Llama-3.2-1B --local --chat
# Start distributed coordinator
distllm-coordinator --model meta-llama/Llama-3.2-7B \
--nodes laptop:50051:0:5 desktop:50052:6:11 friend:50053:12:17
# Start a worker node
distllm-node --node-id laptop --model meta-llama/Llama-3.2-7B \
--start-layer 0 --end-layer 5 --total-layers 18 \
--coordinator-host 192.168.1.100 --coordinator-port 50050
# Start the REST API server
distllm-api --model meta-llama/Llama-3.2-1B --local
# System diagnostics (check Python, CUDA, GPU, network, ports)
distllm doctor
# Quantize a model for smaller footprint
distllm tune quantize --model meta-llama/Llama-3.2-7B --bits 4
# Batch inference over a dataset
distllm tune batch --input prompts.jsonl --output results.jsonl
# Warm the semantic cache from prior prompts
distllm tune cache --preload cache_seed.jsonl
Documentation
- Architecture — Pipeline parallelism, node topology, KV cache management
- Deployment — Local, Docker, and multi-machine deployment
- API Reference — OpenAI-compatible API docs
Project Structure
src/distllm/
├── core/ # Coordinator, request pipeline, batch scheduler
├── dist/ # Distributed inference engine
│ ├── pipeline.py # PipelineOrchestrator — multi-node execution
│ ├── worker.py # WorkerNode — per-device model subset
│ ├── recovery.py # NodeRecoveryManager — failure handling
│ ├── straggler.py # StragglerDetector — slow node detection
│ ├── rebalancer.py # Dynamic pipeline rebalancing
│ ├── wide_area.py # WAN-optimized inference
│ ├── parallel.py # Hybrid parallelism auto-selector
│ ├── p2p/ # P2P gossip protocol & discovery
│ └── partition/ # Hardware-aware auto-partitioner
├── api/ # OpenAI-compatible REST API
├── models/ # Model partitioning & loading
├── cli/ # CLI tool
├── sdk/ # Python client
├── observability/ # Metrics, tracing, logging
└── dashboard/ # Web dashboard
Supported Model Architectures
GPT-2, GPT-Neo, Llama 2/3, Mistral, Mixtral, Qwen2.5, Phi, DeepSeek, StableLM, Pythia, Baichuan, ChatGLM, InternLM, and more via HuggingFace AutoModel.
Roadmap
- Current: Pipeline parallelism across LAN devices, manual node configuration
- Q3 2026: Auto-discovery (mDNS), auto-partitioning, node recovery, GUI dashboard
- Q4 2026: NAT traversal (cross-internet), P2P model distribution, GPU reputation system
- 2027: Federated clusters, speculative parallelism, privacy-preserving split, incentive system
Contributing
See CONTRIBUTING.md.
License
Apache 2.0. See LICENSE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file distributed_llm-0.4.1.tar.gz.
File metadata
- Download URL: distributed_llm-0.4.1.tar.gz
- Upload date:
- Size: 2.4 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.14.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
aaffe3e52262bea1fa346479400b61277a8023d1765fd02e3327a363e6f2eb0a
|
|
| MD5 |
e1d654d2e1b3f1771edd5faf040aefb5
|
|
| BLAKE2b-256 |
d7444cec00c321973c6109f64e11a7c1234f6319fe4da271e93e00688c50c4cd
|
File details
Details for the file distributed_llm-0.4.1-py3-none-any.whl.
File metadata
- Download URL: distributed_llm-0.4.1-py3-none-any.whl
- Upload date:
- Size: 2.8 MB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.14.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
180d267eac4307173f74ba3557340168f840c91530861d7fbba2f59b08aa0253
|
|
| MD5 |
b4c700832dcc6382d45a8ad248f0c981
|
|
| BLAKE2b-256 |
c2eabf51a2ac4f27cf94b30b1713d9489507c3a999dd1de18983ce32ab87f581
|