moe-l2
Run large MoE models on consumer GPUs. A transparent proxy that predicts which experts your prompt needs, preloads them into a shared-memory LRU cache, so you can run 16 GB+ models on 8 GB GPUs.
Quick start
pip install moe-l2 # keyword-only predictor (zero extra deps)
pip install moe-l2[predictor] # hybrid: keyword + semantic embedding
moe-l2 start --model model.gguf --l2-size 4GB
Your tools (curl, Open WebUI, LangChain) connect to localhost:11435 — no client changes needed.
How it works
MoE models have many "experts" but only activate a few per token. moe-l2 predicts your prompt's domain (codegen, math, chinese_tech, etc.) and preloads the relevant experts into an mmap'd LRU cache before they're needed.
user → moe-l2 proxy (localhost:11435)
├── predict domain
├── preload domain experts → /dev/shm/moe_l2/
└── forward to ollama (localhost:11434)
Real-world benchmark
| Your GPU | Normally fits | With moe-l2 |
|---|---|---|
| 4 GB | — | DeepSeek-V2-Lite (16B MoE) ✅ |
| 8 GB | 7B dense | Qwen2.5-32B-A3B (32B MoE) ✅ |
| 12 GB | 13B dense | DeepSeek-V2 (236B MoE) ✅ |
| 24 GB | 34B dense | DeepSeek-V2 (236B MoE) ✅ |
Without moe-l2, an 8 GB card cannot load these models at all — it OOMs immediately. With moe-l2, a 32B MoE fits in ~2.7 GB VRAM (cache=0.5 on DS-V2-Lite).
Benchmarked on RTX 4090
| Mode | GPU VRAM | Speed | What it means |
|---|---|---|---|
| Standard (all experts on GPU) | 23.3 GB | 65 t/s | Needs a 24 GB card |
| moe-l2 (hot-cached experts) | 2.7 GB | ~7 t/s gen · 103 t/s prompt | Fits in 4 GB cards |
| Savings | 88% less | 11% speed | ~20 GB freed for other work |
We benchmarked Qwen3.6-A3B (32B MoE) and DeepSeek-V2-Lite (16B MoE, 64 experts) on RTX 4090. GPU LRU expert cache (Phase 3) is now stable — 0 crashes across 7 cache levels × 3 conversation types. Gen speed is CPU-bound (~5-7 t/s), but followup prompt processing gets 10× faster (~80-103 t/s) from cache hits. Full reports: Qwen3.6 · DS-V2-Lite
Usage
1. L2 proxy (recommended)
Start the transparent proxy — sits between your client and ollama:
moe-l2 start --model /models/DeepSeek-V2-Lite.Q4_K_M.gguf --l2-size 4GB
All default ollama tools work through it (curl, open-webui, langchain):
# streaming
curl http://localhost:11435/api/chat -d '{
"model":"qwen3:4b",
"messages":[{"role":"user","content":"write a Python script"}],
"stream":true
}'
# blocking
curl http://localhost:11435/api/chat -d '{
"model":"qwen3:4b",
"messages":[{"role":"user","content":"hello"}],
"stream":false
}'
2. Monitor cache stats
moe-l2 stats --port 11435
Example output:
moe-l2 cache stats
requests: 47
hits: 42 (89.4%)
misses: 5
slots_used: 32/48 (66.7%)
memory: 456 MB (68.3% of 668 MB)
3. Use as a library
from moe_l2 import predict, predict_hybrid, domain_to_expert_ids
from moe_l2.cache import L2Cache
# Predict domain (zero-dependency mode)
domain = predict("print hello world") # → "codegen"
# Or use the hybrid semantic predictor
domain = predict_hybrid("how do I sort a list?") # → "codegen"
# Preload experts
cache = L2Cache(model_path="model.gguf", l2_size="4GB")
cache.preload(domain_to_expert_ids[domain])
Architecture
┌────────────────────────────────────────────────────────────┐
│ HTTP client │
│ curl / open-webui / langchain / any OpenAI client │
└──────────┬─────────────────────────────────────────────────┘
│ POST /api/chat
▼
┌────────────────────────────────────────────────────────────┐
│ moe-l2 Proxy (port 11435) │
│ │
│ ┌─────────────────────────────────────────────────────┐ │
│ │ Domain Predictor │ │
│ │ - Keyword mode: zero deps, instant classification │ │
│ │ - Hybrid mode: +sentence-transformers for context │ │
│ └─────────────────────┬───────────────────────────────┘ │
│ │ domain │
│ ┌─────────────────────▼───────────────────────────────┐ │
│ │ L2 Cache (mmap'd shared memory) │ │
│ │ - LRU eviction policy │ │
│ │ - Async preload: next-prediction prefetch │ │
│ │ - Thread-safe concurrent access │ │
│ │ - Zero-copy mmap from SSD → RAM │ │
│ └─────────────────────┬───────────────────────────────┘ │
│ │ forward request │
└────────────────────────┼───────────────────────────────────┘
▼
┌────────────────────────────────────────────────────────────┐
│ ollama / llama.cpp (port 11434) │
│ Hot experts in GPU VRAM │
│ Cold experts loaded from RAM/SSD via mmap │
└────────────────────────────────────────────────────────────┘
CLI reference
| Command | Description |
|---|---|
moe-l2 start --model <path> --l2-size <size> |
Start proxy + cache |
moe-l2 start --model <path> --gpu |
Start with GPU-accelerated llama-server |
moe-l2 stats --port <port> |
Show live cache stats |
moe-l2 download-bins [--release TAG] |
Download pre-built GPU binaries from GitHub |
moe-l2 collect --model <path> |
Collect MoE routing data → ~/.moe-l2/maps/domain_expert_map.json |
moe-l2 stop --port <port> |
Stop proxy |
Options:
--model auto: scan/opt/data/models/*.gguf--l2-size 4GB/--l2-size 512MB: target cache size--port 11435(default)--gpu: enable GPU mode (requires CUDA + NVIDIA GPU)
GPU binaries: Not tracked in git (~530 MB). Fetched at runtime via
moe-l2 download-bins. The repo does not track 500MB+ .so files. When youpip install moe-l2, binaries are included. For git-clone users, runmoe-l2 download-binsto fetch them from GitHub Release.
Platform requirements
- Linux x86_64 only — pre-built binaries target Linux AMD64 (CUDA
.so+llama-server) - macOS, Windows, and ARM Linux are not supported
- NVMe SSD strongly recommended
- NVIDIA GPU required for
--gpumode
More data
| Metric | Standard | With moe-l2 |
|---|---|---|
| Prompt processing | 110 t/s | 110 t/s |
| Generation speed | 65 t/s | ~5-7 t/s |
| VRAM used | 23.3 GB | 2.7 GB |
| Model size / VRAM ratio | 0.26× | 2.2× |
The speed tradeoff is intentional: experts load from system RAM via PCIe. Phase 3 GPU LRU cache is stable — gen speed is CPU-bound (~5-7 t/s), but followup prompt processing reaches ~80-103 t/s from cache hits.
A3 Expert Cache (llama.cpp)
Beyond the proxy layer, moe-l2 ships an A3 (Attention-Aware Expert Cache) patch for llama.cpp that compiles expert LRU caching directly into the CUDA backend — no proxy needed. Run any llama.cpp binary with --cpu-moe --expert-cache <fraction>.
./llama-batched -m DeepSeek-V2-Lite.Q2_K.gguf \
-p "prompt" -n 128 -ngl 99 \
--cpu-moe --expert-cache 0.25
Real benchmark (RTX 4090 · DeepSeek-V2-Lite Q2_K 6.4 GB):
| Mode | VRAM used | Speed | Savings |
|---|---|---|---|
| OG (---no-mmap, full GPU) | 6,635 MB | 126.64 t/s | baseline |
| A3 (--expert-cache 0.25) | 1,175 MB | 8.22 t/s | 5,460 MB (5.64×) |
A3 caches 25% of the most-recently-used experts on GPU and swaps inactive ones in from CPU RAM on demand. Good for single-user chat: acceptable latency (~7-8 t/s gen) with >80% VRAM savings. For batch / high-throughput scenarios, increase the cache fraction or omit --cpu-moe.
When the cache helps (and when it doesn't)
The expert cache only pays off when experts actually live in CPU RAM (default mmap mode). Verified on RTX 4090 with Mixtral 8x7B:
| Mode | Expert location | Gen speed | Cache value |
|---|---|---|---|
--no-mmap (all weights on GPU) |
GPU VRAM | 3.7 t/s | ❌ cache is a no-op layer |
--no-mmap + cache |
GPU VRAM | 3.4-3.5 t/s | ❌ slower (cache forces the generic expert path, ~3 ms/token fixed overhead regardless of hit rate) |
default mmap + cache |
CPU RAM | ~1 t/s (scheduler copy bound) | ⚠️ not recommended |
default mmap (no cache) |
CPU RAM | 0.9 t/s | baseline |
Key findings (2026-08-01, per-segment CUDA timing):
- With
--no-mmap, experts are already fully resident in GPU VRAM (copies are device-to-device, 100% on-GPU) — the cache adds nothing but an extra layer. - Turning the cache on forces the generic
MUL_MAT_IDpipeline (host-side id sorting + two stream syncs ≈ 3.1 ms/token) independent of hit rate — hit rate 27% vs 49% both showed identical ~3.1 ms. The fast expert path only runs with the cache off. - Conclusion: use the cache only when experts are CPU-hosted (mmap). The bundled CLI defaults to this configuration (
mmapdefault, no--no-mmapflag).
Run the demo yourself:
bash examples/demo_a3_compression.sh(edit paths first).
Related work
TencentYoutuResearch/Palm-Infra / mollm is a C++ engine from Tencent for MoE models with SSD expert offload on Apple Silicon / ARM Linux (16.22 t/s, 122B MoE, 16 GB peak RSS).
| Dimension | mollm (Tencent) | moe-l2 |
|---|---|---|
| Platform | Apple Silicon / ARM Linux | Linux x86_64 + GPU (NVIDIA) |
| Install | Build from source (CMake + C++) | pip install moe-l2 |
| Model support | Qwen-series only | Any llama.cpp MoE (DeepSeek, Qwen, Mixtral...) |
| Backend | Custom C++ engine | llama.cpp proxy — zero migration |
| GPU acceleration | CPU only (NEON) | CUDA + GPU VRAM |
| Target user | Mobile / edge developers | Desktop homelab users |
Project status
- ✅ Domain predictor (keyword + optional semantic)
- ✅ L2 cache (mmap LRU, thread-safe, async preload)
- ✅ Transparent proxy (HTTP/SSE forwarding)
- ✅ CLI with auto model detection, GPU mode, and
collect(routing data → expert map) - ✅ GPU mode verified on RTX 4090 (DS-V2-Lite, ~1.6 GiB VRAM, 95% savings)
- ✅ GPU LRU expert cache (verified: Qwen3.6 + DS-V2-Lite, 7 levels × 3 types, 0 crashes)
- ✅ Expert cache boundary verified on Mixtral 8x7B / RTX 4090: under
--no-mmapexperts are already fully resident in VRAM, so the cache is a no-op layer that adds ~3 ms/token overhead (3.7 → 3.4 t/s); it only pays off when experts are CPU-hosted (mmap). CLI defaults to the mmap configuration. - ✅ PyPI package (
moe-l2)
License
Apache 2.0. See LICENSE for details.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file moe_l2-0.4.0.tar.gz.
File metadata
- Download URL: moe_l2-0.4.0.tar.gz
- Upload date:
- Size: 77.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.11.6 {"installer":{"name":"uv","version":"0.11.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ca059bdb4e3f3ff470b25ff3bfc1dccebd90b374d66b5dd557ddd2f1f4e3e969
|
|
| MD5 |
e3715c80376f69cf0020947270181b98
|
|
| BLAKE2b-256 |
a7c70eb6c48654fed07595042560a12d7d2462710564744fa61dd1ab4186820f
|
File details
Details for the file moe_l2-0.4.0-py3-none-any.whl.
File metadata
- Download URL: moe_l2-0.4.0-py3-none-any.whl
- Upload date:
- Size: 79.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.11.6 {"installer":{"name":"uv","version":"0.11.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d03aaca3b3ddf030ed2acb06719cc38928ed0d58698e51e92916750e59cc827f
|
|
| MD5 |
2ba42075106552b5831dcafc970f7454
|
|
| BLAKE2b-256 |
07627ac480b692828212cb78b0debd2bba16b4ff1558647bebadec5eaba19c8b
|