Skip to main content

experts4bit-qlora

CI PyPI

Train and serve fused Mixture-of-Experts models in 4-bit on hardware that cannot hold them in bf16.

transformers v5 stores a MoE's experts as one fused 3-D parameter per layer. bitsandbytes' 4-bit walker only replaces nn.Linear, so it silently skips the experts — the overwhelming majority of the weights (bitsandbytes#1849). This package quantises exactly that fused stack (Experts4bit, the 4-bit face of ExpertsNbit: nf4 / fp4 / int8 / fp8 / bf16 / fp16 storage, with a test-pinned fidelity ordering), pairs it with a streaming loader and per-expert LoRA so you can fine-tune, and serves the result through a paged decode engine that is measured against each model's own attention.

Current position, one page: docs/STATUS.md. Every number, with its evidence and status: docs/claims.json.

Install

pip install experts4bit-qlora           # primitive + adapters (torch + bitsandbytes)
pip install "experts4bit-qlora[train]"  # + the streaming MoE trainer
pip install "experts4bit-qlora[fast]"   # + the fused grouped-GEMM path (grouped-nf4-gemm)

e4b, experts4bit and expertsnbit are aliases of this package. Runs on stock bitsandbytes; every feature has a reference path.

Which door? Start from what does not fit

what ran out call needs
nothing — just train a fused MoE load_moe_4bit_streaming(...) [train]
each step is slow enable_fast_train(model, dgrad=True) [fast]
…and [fast] will not build enable_batched_train(model)
the experts do not fit VRAM load_moe_4bit_streaming(..., offload=True)
the experts do not fit host RAM, serving enable_nvme_residency(...) [fast] + arena
…and they are native MXFP4 enable_mxfp4_nvme_residency(...) [fast] + arena
the experts do not fit host RAM, training enable_nvme_train_residency(...) [fast] + arena + grad ckpt
the dense side does not fit enable_dense_offload(model, "cuda")
serving, want it faster enable_fast(model) [fast]
serving, spare VRAM to trade enable_pipelined_residency(model, hot_sets, k_slots=k) [fast]

Reasoning and caveats for each: docs/CHOOSING.md. Assert the return value of every enable_*: 0 and "silently still on the per-expert loop" look identical from the caller's side.

Quickstart

import torch
from experts4bit_qlora import Experts4bit, ExpertsLoRA, load_moe_4bit_streaming, verify_moe_4bit

# A real fused-MoE checkpoint, quantised on the way to the GPU (never bf16-resident):
model, config = load_moe_4bit_streaming(
    "Qwen/Qwen3-30B-A3B", "cuda", torch.bfloat16, r=8, alpha=16, quant_type="nf4",
)
verify_moe_4bit(model, strict=True)   # raises if any expert stack is still high precision
STEPS=150 R=8 TRAIN_EXPERTS=1 OUT=./out python -m experts4bit_qlora.train      # QLoRA fine-tune
ADAPTER=./out/adapter_best.pt python -m experts4bit_qlora.infer                 # serve it

Do not load these models with stock from_pretrained(..., load_in_4bit=True): it quantises the nn.Linear layers, leaves the experts in bf16, and OOMs.

What is measured

Each row is an entry in docs/claims.json; the last column is its evidence status. measured means the receipt is in this repository; measured-private means the run happened but the receipt lives in a private audit tree and you cannot check it from here.

result status
OLMoE-1B-7B fits a 12 GB card and trains 4.70 GB load; held-out eval 1.4813 → 1.0290 measured
Expert offload trains 30B-class MoEs on 12 GB Qwen3-30B-A3B peaks 7.16 GB, Gemma-4-26B-A4B 8.47 GB measured
Fused training path, two 30B MoEs × five datasets 1.52–1.81× per step at 0.75–0.81× VRAM, loss parity, frozen stack bit-identical over 16.31 GB measured
Arena vs pinned host RAM, at a descending cap 2.56× / 3.80× / 6.40× less host RAM (OLMoE / Gemma-4 / Qwen3-30B) measured
Paged decode vs the model's own attention indistinguishable on Granite (0.00229 nats), gpt-oss (0.00288) and Qwen3 (0.00173) against a chunk-free reference, each below its own floor; Gemma-4 has no reference at this resolution — its own cached forward swings −0.107 … +0.271 nats across windows — and the paged path's one measured cost there is the fp8 cache, 0.046 nats (#359) measured-private
Per-family serving throughput on one rented RTX 5090 class (six families, same protocol) Qwen3-30B 97 → 155 tok/s B=1 and 483 → 944 B=16; OLMoE 248 → 452; Granite 191 → 285; Mixtral 48 → 107; gpt-oss and Gemma-4 NF4 only (124, 71) — the refused arms are the build-out (SERVING-THROUGHPUT.md) measured
Single-stream Qwen3-30B-A3B on an RTX 5090 ≈100 tok/s NF4; 204.6 tok/s with calibrated int4 attention + int4 experts measured-private
Batched (B=16) Qwen3-30B-A3B on an RTX 5090 ≈1,238 tok/s aggregate measured-private
Same box, same prompts, against vLLM (GPTQ-Int4) vLLM ahead 1.47× at B=1, 1.55× at B=16 measured-private
DeepSeek-V4-Flash (284B, 147 GB of experts on disk) loads in ~10 s at 8.74 GiB peak VRAM and generates measured
Informed hot sets vs by-index, identical VRAM +37.1% on DeepSeek-V4-Flash; the gain is a property of the host measured

Three things to read beside that table, because they change what it means:

  • A parity delta is read against a per-model noise floor, never against zero. Two arithmetically equivalent forwards of an MoE disagree, because rounding flips which experts the router picks; on gpt-oss 4.5% of layer-token choices flip and those tokens carry the whole disagreement. "Below the floor" means indistinguishable. docs/METHODOLOGY.md §13.1.
  • 4-bit on a card that already fits the model is a 1.2–2.3× energy penalty, not a saving: NF4 is storage-only and the GEMM runs in bf16 either way. It inverts when memory binds.
  • Ratios travel; absolutes do not. The 5090 class carries ~8.5% inter-box dispersion; the same config on two 4090s moved 8.6% in s/step. Quote the card, or quote a ratio.

What was retired

Claims this project published and then withdrew, each with the measurement that withdrew it, are listed in docs/STATUS.md and kept as retired entries in docs/claims.json so they stay findable. The most recent: the "+0.047 ppl fp8 KV cost" on Qwen3 (below the model's own floor), and the "+0.078 nats gpt-oss sinks/windows defect" (the chunked oracle was the drifting arm, not the serving path).

Scope

The primitives are model-agnostic. The streaming loader and trainer handle SwiGLU fused-MoE families stored per-expert or pre-fused: OLMoE, Qwen3-MoE / Qwen3.5-MoE, Gemma-4 (text tower), GraniteMoe, gpt-oss (MXFP4 experts with per-expert biases and a clamped GLU, dequantised bit-identically), and DeepSeek-V4 (Flash / Pro). Which families load, run and CUDA-graph-capture, with the evidence: docs/ARCHITECTURE_SUPPORT.md. Unsupported architectures fail fast with a clear error.

Known open: Gemma-4-26B-A4B's fp8 K cache wants finer groups on its 512-dim heads, and the family needs a parity instrument that survives its batch-shape variance (#359); the model fails to load on 2 of 5 rented hosts (#344); no shipped tool bakes the training arena from a bf16 checkpoint yet.

Docs

docs/STATUS.md what you get, what was retired, what is open — one page
docs/claims.json every claim with value, hardware, status, evidence
docs/INDEX.md what each of the 42 documents is, and whether it is current
docs/CHOOSING.md which mode, and why
docs/METHODOLOGY.md hosts, protocols, every measurement's provenance
docs/SERVING-PARITY.md paged decode vs each model's own attention
docs/SERVING-THROUGHPUT.md per-family decode throughput under one protocol, with the refusal list
docs/STORAGE-MODES.md the six storage modes and what each promises
docs/RESIDENCY-ENGINES.md residency engines, hot-set selection, host-regime laws
docs/SERVING.md the HTTP shim and Docker deployment
docs/DEEPSEEK-V4.md V4's storage split, epilogue, arena bake
docs/BITSANDBYTES.md relationship to bitsandbytes, prior art

The package family

  • experts4bit-qlora (this repo) owns everything around the expert GEMM: the fused-stack primitives and per-expert LoRA, the streaming loaders, offload, training, the paged serving engine, hot-expert residency.
  • grouped-nf4-gemm owns the GEMM itself: one launch over 4-bit-packed expert stacks with in-register decode and fp32 accumulation, plus the fp8 paged decode attention and the decode glue kernels. [fast] is the seam.

The kernel makes one expert-stack matmul cheap; this package decides which bytes are where.

Provenance

Every number traces to a committed script and a named host, with receipts under bench/ and docs/ — or, where the receipt is private, the register says so. PROVENANCE.md is the OpenTimestamps-anchored record for the v0.2.0 convergence result; anchored documents are never edited in place (see docs/INDEX.md). Falsification work lives under audits/.

License

MIT. experts4bit_qlora/_vendor/experts.py is vendored from bitsandbytes (also MIT) pending upstream merge.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

experts4bit_qlora-0.34.0.tar.gz (772.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

experts4bit_qlora-0.34.0-py3-none-any.whl (414.5 kB view details)

Uploaded Python 3

File details

Details for the file experts4bit_qlora-0.34.0.tar.gz.

File metadata

  • Download URL: experts4bit_qlora-0.34.0.tar.gz
  • Upload date:
  • Size: 772.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for experts4bit_qlora-0.34.0.tar.gz
Algorithm Hash digest
SHA256 be035b6867f849a1b33b5a1fbb80e6b44abf603815a703509e20d467e47cdd11
MD5 161e21c4b9d1d3bf8261359fc3ab6a09
BLAKE2b-256 087aea01470f059bbf89585531639111e8820a12296b492a5336e0383c359594

See more details on using hashes here.

Provenance

The following attestation bundles were made for experts4bit_qlora-0.34.0.tar.gz:

Publisher: release.yml on pjordanandrsn/experts4bit-qlora

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file experts4bit_qlora-0.34.0-py3-none-any.whl.

File metadata

File hashes

Hashes for experts4bit_qlora-0.34.0-py3-none-any.whl
Algorithm Hash digest
SHA256 6da277cc06d609da4baa58177f47a8a492e5e7a3a9c12c71bcb558adb5708aed
MD5 ad85f8e4d768b1c44d9b9e3e2798ed6f
BLAKE2b-256 1f9d8e2aeca49422f92fdf4f8f0b3a0f22c4de5e8fb7e2cd5839c2e429851fee

See more details on using hashes here.

Provenance

The following attestation bundles were made for experts4bit_qlora-0.34.0-py3-none-any.whl:

Publisher: release.yml on pjordanandrsn/experts4bit-qlora

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.35.3

2 files

0.35.2

2 files

0.35.1

2 files

0.35.0

2 files

This release

0.34.0 This release

2 files

0.33.0

2 files

0.32.0

2 files

0.31.2

2 files

0.31.1

2 files

0.31.0

2 files

0.30.0

2 files

0.29.0

2 files

0.28.0

2 files

0.27.0

2 files

0.26.0

2 files

0.25.0

2 files

0.24.2

2 files

0.24.1

2 files

0.24.0

2 files

0.23.0

2 files

0.22.0

2 files

0.21.0

2 files

0.20.0

2 files

0.19.1

2 files

0.19.0

2 files

0.18.0

2 files

0.17.5

2 files

0.17.4

2 files

0.17.3

2 files

0.17.2

2 files

0.17.1

2 files

0.17.0

2 files

0.16.3

2 files

0.16.2

2 files

0.16.1

2 files

0.16.0

2 files

0.15.0

2 files

0.14.0

2 files

0.13.0

2 files

0.12.0

2 files

0.11.1

2 files

0.11.0

2 files

0.10.0

2 files

0.9.1

2 files

0.9.0

2 files

0.8.0

2 files

0.7.1

2 files

0.7.0

2 files

0.6.7

2 files

0.6.6

2 files

0.6.5

2 files

0.6.4

2 files

0.6.3

2 files

0.6.2

2 files

0.6.1

2 files

0.6.0

2 files

0.5.0

2 files

0.4.0

2 files

0.3.0

2 files

0.2.0

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page