Skip to main content

moe-l2

Run large MoE models on consumer GPUs. A transparent proxy that predicts which experts your prompt needs, preloads them into a shared-memory LRU cache, so you can run 16 GB+ models on 8 GB GPUs.

Quick start

pip install moe-l2                   # keyword-only predictor (zero extra deps)
pip install moe-l2[predictor]        # hybrid: keyword + semantic embedding
moe-l2 start --model model.gguf --l2-size 4GB

Your tools (curl, Open WebUI, LangChain) connect to localhost:11435 — no client changes needed.

How it works

MoE models have many "experts" but only activate a few per token. moe-l2 predicts your prompt's domain (codegen, math, chinese_tech, etc.) and preloads the relevant experts into an mmap'd LRU cache before they're needed.

user → moe-l2 proxy (localhost:11435)
    ├── predict domain
    ├── preload domain experts → /dev/shm/moe_l2/
    └── forward to ollama (localhost:11434)

Real-world benchmark

Your GPU Normally fits With moe-l2
4 GB DeepSeek-V2-Lite (16B MoE) ✅
8 GB 7B dense Qwen2.5-32B-A3B (32B MoE) ✅
12 GB 13B dense DeepSeek-V2 (236B MoE) ✅
24 GB 34B dense DeepSeek-V2 (236B MoE) ✅

Without moe-l2, an 8 GB card cannot load these models at all — it OOMs immediately. With moe-l2, a 32B MoE fits in ~2.7 GB VRAM (cache=0.5 on DS-V2-Lite).

Benchmarked on RTX 4090

Mode GPU VRAM Speed What it means
Standard (all experts on GPU) 23.3 GB 65 t/s Needs a 24 GB card
moe-l2 (hot-cached experts) 2.7 GB ~7 t/s gen · 103 t/s prompt Fits in 4 GB cards
Savings 88% less 11% speed ~20 GB freed for other work

We benchmarked Qwen3.6-A3B (32B MoE) and DeepSeek-V2-Lite (16B MoE, 64 experts) on RTX 4090. GPU LRU expert cache (Phase 3) is now stable — 0 crashes across 7 cache levels × 3 conversation types. Gen speed is CPU-bound (~5-7 t/s), but followup prompt processing gets 10× faster (~80-103 t/s) from cache hits. Full reports: Qwen3.6 · DS-V2-Lite

Usage

1. L2 proxy (recommended)

Start the transparent proxy — sits between your client and ollama:

moe-l2 start --model /models/DeepSeek-V2-Lite.Q4_K_M.gguf --l2-size 4GB

All default ollama tools work through it (curl, open-webui, langchain):

# streaming
curl http://localhost:11435/api/chat -d '{
  "model":"qwen3:4b",
  "messages":[{"role":"user","content":"write a Python script"}],
  "stream":true
}'

# blocking
curl http://localhost:11435/api/chat -d '{
  "model":"qwen3:4b",
  "messages":[{"role":"user","content":"hello"}],
  "stream":false
}'

2. Monitor cache stats

moe-l2 stats --port 11435

Example output:

moe-l2 cache stats
  requests:     47
  hits:         42     (89.4%)
  misses:        5
  slots_used:  32/48  (66.7%)
  memory:     456 MB  (68.3% of 668 MB)

3. Use as a library

from moe_l2 import predict, predict_hybrid, domain_to_expert_ids
from moe_l2.cache import L2Cache

# Predict domain (zero-dependency mode)
domain = predict("print hello world")  # → "codegen"

# Or use the hybrid semantic predictor
domain = predict_hybrid("how do I sort a list?")  # → "codegen"

# Preload experts
cache = L2Cache(model_path="model.gguf", l2_size="4GB")
cache.preload(domain_to_expert_ids[domain])

Architecture

┌────────────────────────────────────────────────────────────┐
│                       HTTP client                          │
│     curl / open-webui / langchain / any OpenAI client      │
└──────────┬─────────────────────────────────────────────────┘
           │ POST /api/chat
           ▼
┌────────────────────────────────────────────────────────────┐
│               moe-l2 Proxy (port 11435)                     │
│                                                            │
│  ┌─────────────────────────────────────────────────────┐   │
│  │  Domain Predictor                                    │   │
│  │  - Keyword mode: zero deps, instant classification   │   │
│  │  - Hybrid mode: +sentence-transformers for context   │   │
│  └─────────────────────┬───────────────────────────────┘   │
│                        │ domain                             │
│  ┌─────────────────────▼───────────────────────────────┐   │
│  │  L2 Cache (mmap'd shared memory)                    │   │
│  │  - LRU eviction policy                              │   │
│  │  - Async preload: next-prediction prefetch          │   │
│  │  - Thread-safe concurrent access                    │   │
│  │  - Zero-copy mmap from SSD → RAM                    │   │
│  └─────────────────────┬───────────────────────────────┘   │
│                        │ forward request                    │
└────────────────────────┼───────────────────────────────────┘
                         ▼
┌────────────────────────────────────────────────────────────┐
│                ollama / llama.cpp (port 11434)              │
│                Hot experts in GPU VRAM                      │
│                Cold experts loaded from RAM/SSD via mmap    │
└────────────────────────────────────────────────────────────┘

CLI reference

Command Description
moe-l2 start --model <path> --l2-size <size> Start proxy + cache
moe-l2 start --model <path> --gpu Start with GPU-accelerated llama-server
moe-l2 stats --port <port> Show live cache stats
moe-l2 download-bins [--release TAG] Download pre-built GPU binaries from GitHub
moe-l2 collect --model <path> Collect MoE routing data → ~/.moe-l2/maps/domain_expert_map.json
moe-l2 stop --port <port> Stop proxy

Options:

  • --model auto: scan /opt/data/models/*.gguf
  • --l2-size 4GB / --l2-size 512MB: target cache size
  • --port 11435 (default)
  • --gpu: enable GPU mode (requires CUDA + NVIDIA GPU)

GPU binaries: Not tracked in git (~530 MB). Fetched at runtime via moe-l2 download-bins. The repo does not track 500MB+ .so files. When you pip install moe-l2, binaries are included. For git-clone users, run moe-l2 download-bins to fetch them from GitHub Release.

Platform requirements

  • Linux x86_64 only — pre-built binaries target Linux AMD64 (CUDA .so + llama-server)
  • macOS, Windows, and ARM Linux are not supported
  • NVMe SSD strongly recommended
  • NVIDIA GPU required for --gpu mode

More data

Metric Standard With moe-l2
Prompt processing 110 t/s 110 t/s
Generation speed 65 t/s ~5-7 t/s
VRAM used 23.3 GB 2.7 GB
Model size / VRAM ratio 0.26× 2.2×

The speed tradeoff is intentional: experts load from system RAM via PCIe. Phase 3 GPU LRU cache is stable — gen speed is CPU-bound (~5-7 t/s), but followup prompt processing reaches ~80-103 t/s from cache hits.

A3 Expert Cache (llama.cpp)

Beyond the proxy layer, moe-l2 ships an A3 (Attention-Aware Expert Cache) patch for llama.cpp that compiles expert LRU caching directly into the CUDA backend — no proxy needed. Run any llama.cpp binary with --cpu-moe --expert-cache <fraction>.

./llama-batched -m DeepSeek-V2-Lite.Q2_K.gguf \
  -p "prompt" -n 128 -ngl 99 \
  --cpu-moe --expert-cache 0.25

Real benchmark (RTX 4090 · DeepSeek-V2-Lite Q2_K 6.4 GB):

Mode VRAM used Speed Savings
OG (---no-mmap, full GPU) 6,635 MB 126.64 t/s baseline
A3 (--expert-cache 0.25) 1,175 MB 8.22 t/s 5,460 MB (5.64×)

A3 caches 25% of the most-recently-used experts on GPU and swaps inactive ones in from CPU RAM on demand. Good for single-user chat: acceptable latency (~7-8 t/s gen) with >80% VRAM savings. For batch / high-throughput scenarios, increase the cache fraction or omit --cpu-moe.

When the cache helps (and when it doesn't)

The expert cache only pays off when experts actually live in CPU RAM (default mmap mode). Verified on RTX 4090 with Mixtral 8x7B:

Mode Expert location Gen speed Cache value
--no-mmap (all weights on GPU) GPU VRAM 3.7 t/s ❌ cache is a no-op layer
--no-mmap + cache GPU VRAM 3.4-3.5 t/s ❌ slower (cache forces the generic expert path, ~3 ms/token fixed overhead regardless of hit rate)
default mmap + cache CPU RAM ~1 t/s (scheduler copy bound) ⚠️ not recommended
default mmap (no cache) CPU RAM 0.9 t/s baseline

Key findings (2026-08-01, per-segment CUDA timing):

  • With --no-mmap, experts are already fully resident in GPU VRAM (copies are device-to-device, 100% on-GPU) — the cache adds nothing but an extra layer.
  • Turning the cache on forces the generic MUL_MAT_ID pipeline (host-side id sorting + two stream syncs ≈ 3.1 ms/token) independent of hit rate — hit rate 27% vs 49% both showed identical ~3.1 ms. The fast expert path only runs with the cache off.
  • Conclusion: use the cache only when experts are CPU-hosted (mmap). The bundled CLI defaults to this configuration (mmap default, no --no-mmap flag).

Run the demo yourself: bash examples/demo_a3_compression.sh (edit paths first).

Related work

TencentYoutuResearch/Palm-Infra / mollm is a C++ engine from Tencent for MoE models with SSD expert offload on Apple Silicon / ARM Linux (16.22 t/s, 122B MoE, 16 GB peak RSS).

Dimension mollm (Tencent) moe-l2
Platform Apple Silicon / ARM Linux Linux x86_64 + GPU (NVIDIA)
Install Build from source (CMake + C++) pip install moe-l2
Model support Qwen-series only Any llama.cpp MoE (DeepSeek, Qwen, Mixtral...)
Backend Custom C++ engine llama.cpp proxy — zero migration
GPU acceleration CPU only (NEON) CUDA + GPU VRAM
Target user Mobile / edge developers Desktop homelab users

Project status

  • ✅ Domain predictor (keyword + optional semantic)
  • ✅ L2 cache (mmap LRU, thread-safe, async preload)
  • ✅ Transparent proxy (HTTP/SSE forwarding)
  • ✅ CLI with auto model detection, GPU mode, and collect (routing data → expert map)
  • ✅ GPU mode verified on RTX 4090 (DS-V2-Lite, ~1.6 GiB VRAM, 95% savings)
  • ✅ GPU LRU expert cache (verified: Qwen3.6 + DS-V2-Lite, 7 levels × 3 types, 0 crashes)
  • ✅ Expert cache boundary verified on Mixtral 8x7B / RTX 4090: under --no-mmap experts are already fully resident in VRAM, so the cache is a no-op layer that adds ~3 ms/token overhead (3.7 → 3.4 t/s); it only pays off when experts are CPU-hosted (mmap). CLI defaults to the mmap configuration.
  • ✅ PyPI package (moe-l2)

License

Apache 2.0. See LICENSE for details.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

moe_l2-0.4.0.tar.gz (77.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

moe_l2-0.4.0-py3-none-any.whl (79.7 kB view details)

Uploaded Python 3

File details

Details for the file moe_l2-0.4.0.tar.gz.

File metadata

  • Download URL: moe_l2-0.4.0.tar.gz
  • Upload date:
  • Size: 77.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.6 {"installer":{"name":"uv","version":"0.11.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for moe_l2-0.4.0.tar.gz
Algorithm Hash digest
SHA256 ca059bdb4e3f3ff470b25ff3bfc1dccebd90b374d66b5dd557ddd2f1f4e3e969
MD5 e3715c80376f69cf0020947270181b98
BLAKE2b-256 a7c70eb6c48654fed07595042560a12d7d2462710564744fa61dd1ab4186820f

See more details on using hashes here.

File details

Details for the file moe_l2-0.4.0-py3-none-any.whl.

File metadata

  • Download URL: moe_l2-0.4.0-py3-none-any.whl
  • Upload date:
  • Size: 79.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.6 {"installer":{"name":"uv","version":"0.11.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for moe_l2-0.4.0-py3-none-any.whl
Algorithm Hash digest
SHA256 d03aaca3b3ddf030ed2acb06719cc38928ed0d58698e51e92916750e59cc827f
MD5 2ba42075106552b5831dcafc970f7454
BLAKE2b-256 07627ac480b692828212cb78b0debd2bba16b4ff1558647bebadec5eaba19c8b

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page