Skip to main content

NWC · Neural Weight Compression

Lossless BF16 weights, 31 % smaller, decoded inside the CUDA matvec. Faster than cuBLAS on the uncompressed weights. Bit-identical weights, no quantization.

PyPI downloads ci model license

Quickstart · Results · How it works · Hardware · Roadmap · FAQ · For agents

Qwen3-4B: tokens/s, VRAM, and GPU time per token vs DFloat11

Token generation is memory-bound: every weight is read once per token. NWC stores the weights entropy-coded in VRAM and decodes them in registers, inside the matrix-vector kernel, so the GPU reads 69 % of the bytes and no decompressed weight ever touches memory. On an RTX 4070 that makes Qwen3-4B 22 % faster than native BF16 while using 2.4 GB less VRAM; on a bandwidth-starved NVIDIA A16 it is still faster. Qwen2.5-7B in full BF16 fits a 12 GB card.

News

  • main · tensor-core accumulation layout (format v10, opt-in layout=1): fp8 kernels on the A16 9–13 % faster, at parity on the 4070; v9 checkpoints keep loading (docs/results.md, section 9).
  • main · fp8: weight-only fp8 e4m3 models (per-channel scale) with the fp8 values stored lossless, 0.87 of the fp8 size, same speed as a native fp8 matvec on the RTX 4070 and 1.2–2.1× cuBLAS BF16 (--fp8 in the demo, convert(model, elem="fp8"); docs/results.md, section 8).
  • 2026-09-19 · v0.9.1: python -m nwc.doctor environment check, export_bf16 back to plain checkpoints, Hugging Face repo ids in load_pretrained and the demo, kernels built for Turing (sm_75), CI, and a 0.6B smoke-test checkpoint.
  • 2026-09-16 · v0.9.0: format v9 (prefix code + lookup table) reaches parity on the A16, any matrix shape, PyPI package and the Qwen3-4B checkpoint.

Quickstart

Needs an NVIDIA GPU (Turing or newer, Ampere or newer measured), a driver for CUDA 12.6+ and PyTorch with CUDA.

pip install neural-weight-compression transformers accelerate
python -m nwc.doctor                                          # GPU, driver, library, kernel round trip: all ok?
python -m nwc.demo Parda21/Qwen3-0.6B-NWC --load --graph      # 1 minute, 0.8 GB download: does everything run?
python -m nwc.demo Parda21/Qwen3-4B-NWC --load --graph        # 5.6 GB: the model the numbers above are from
python -m nwc.demo Qwen/Qwen3-4B --native --graph             # the same model uncompressed, for comparison

The demo prints VRAM in use, the generated text and tokens/s, once through HF generate and once as a CUDA graph (no Python overhead). --save DIR writes a compressed checkpoint of any BF16 model you pass it.

from nwc import fuse, convert, save_pretrained, load_pretrained, export_bf16

fuse(model); convert(model)                             # any HF causal LM in BF16: nn.Linear -> NWCLinear, on the GPU
save_pretrained(model, "my-model-NWC", tokenizer=tok)   # compressed checkpoint (safetensors + nwc_config.json)
model = load_pretrained("my-model-NWC")                 # or a HF repo id; no BF16 originals needed
export_bf16("my-model-NWC", "my-model")                 # back to a plain BF16 checkpoint, bit-identical

The converted model is a normal Transformers model: generate, chat templates, StaticCache and CUDA graphs all work. Batch 1 runs through the fused kernel; prefill dequantizes into a temporary buffer and calls cuBLAS. convert(model, elem="fp8") (demo: --fp8) quantizes to weight-only fp8 e4m3 first and stores the fp8 values lossless: Qwen3-4B in 3.61 GB of VRAM at 73.6 tokens/s on the RTX 4070 (native BF16: 8.10 GB, 45 tokens/s), the kernel as fast as a native fp8 matvec. Ready-made: Parda21/Qwen3-4B-NWC-fp8.

Results

Qwen3-4B, greedy decoding as a CUDA graph, same tokens as the reference (64/64). Full tables, the cost model and the negative results are in docs/results.md.

RTX 4070 (452 GB/s) NVIDIA A16 (vGPU 16Q, 10 SMs, 165 GB/s)
VRAM, native BF16 → NWC 8.10 GB → 5.67 GB 8.10 GB → 5.67 GB
tokens/s, native → NWC 45 → 55 (1.22×) 16.8 → 18.2 (1.08×)
GPU time per token, native → NWC 20.4 ms → 18.9 ms 54.1 ms → 52.1 ms
weight kernels vs cuBLAS, 5 layer shapes 1.17–1.55× 1.06–1.18×
logits vs native bit-identical within cuBLAS' own cross-GPU spread (max 0.20, argmax identical)

Against DFloat11 (NeurIPS 2025), which compresses the same exponent bits with Huffman coding but decodes into a buffer before the matmul. Same model, same GPU, GPU time per token by torch.profiler (scripts/compare_df11.py):

RTX 4070 native BF16 DFloat11 NWC
VRAM 8.10 GB 5.73 GB 5.67 GB
GPU time per token 20.4 ms 56.1 ms (2.74× native) 18.9 ms (0.93× native)
weight kernels per token gemv 17.2 ms decode 36.6 + gemv 16.2 ms 17.3 ms, decode fused

Same size, opposite speed: DFloat11's decoded weights pass through memory twice, NWC reads only the compressed bytes. Both land at the entropy of the mantissa; the choice of coder moves the size by 1–2 %.

How it works

BF16 weight → prefix code → fused matvec → output

  • Byte planes. A BF16 weight is split into its mantissa byte (7.97 bits of entropy, stored raw) and its exponent/sign byte (2.6–2.9 bits of entropy, entropy-coded). That is where the 31 % come from; ZipNN and DFloat11 land at the same figure because it is the entropy of the data.
  • Prefix code instead of rANS (format v9). Exponents are ranked by frequency and written as a unary rank code plus a raw sign bit, with an escape for rare values. A 12-bit peek into a 16 KB lookup table in shared memory yields two complete codes per step; the bit window is advanced with funnel shifts and refilled with a single wide multiply-add. Measured on the A16: 11–13 SM clocks per 64 weights, below the 13.6 needed to match cuBLAS.
  • Fused kernel. One persistent block per SM; each lane decodes its own bit stream for a block of 8 rows × 512 columns, keeps its 16 activations in registers and accumulates eight row sums; a transposed shuffle reduction writes partial sums per (row, column block). Any matrix shape works (padding costs 2 bits per filler weight).
  • Dequantization and embedding lookup (tied lm_head) run on the same decoder.

The format is specified in docs/format.md and reproduced by a pure-Python reference decoder in tests/test_format_cpu.py.

Hardware

Whether NWC beats native depends on decoder throughput versus memory bandwidth × 0.69 (see docs/results.md, section 3). Measured rows are ours; projected rows assume Ampere issue rates.

GPU status weight kernels vs cuBLAS Qwen3-4B tokens/s, native → NWC
RTX 4070 ✅ measured 1.17–1.55× 45 → 55
NVIDIA A16 (vGPU 16Q) ✅ measured 1.06–1.18× 16.8 → 18.2
RTX 4080 Super ⏳ next (in house) ~1.4× projected
H100 PCIe 🔍 wanted ~1.1–1.2× projected
H100 SXM 🔍 wanted ~0.8–0.9× projected, tensor-core path planned
A100 🔍 wanted ~0.9× projected
RTX 4090 / 3090 / 30xx 🔍 wanted bandwidth-bound, > 1× expected
T4 / RTX 20xx (Turing) 🔧 builds (sm_75), unmeasured
Blackwell (sm_100/120) 🔧 PTX JIT on Linux, sm_120 built on Windows, unmeasured

Have one of the wanted GPUs? Two commands and the output in a benchmark issue puts your card in this table:

python -m nwc.demo Parda21/Qwen3-4B-NWC --load --graph && python -m nwc.demo Qwen/Qwen3-4B --native --graph

Roadmap

Tracked in milestones and roadmap issues.

  • v0.9 · format v9, A16 parity, shape-independent, PyPI wheels (Windows, Linux), HF checkpoints, CI
  • v0.10 · fp8 weight-only element type (fp8 values lossless, 0.87 of fp8 size, native-fp8 speed on Ada)
  • v0.10 · measure RTX 4080 Super (#1) and H100 (#2); community results into the hardware table (#8)
  • v0.10 · tensor-core accumulation layout (mma.m16n8k16): built and measured (#3); fp8 on the A16 +9–13 %, BF16 at parity, so the H100 question stays open and moves to wider stream refills
  • v0.10 · why does Ada need 17 SM clocks per warp step where Ampere needs 11? (same SASS; docs/results.md section 8)
  • v0.10 · second model family (Llama 3.1 8B, Mistral) with perplexity check (#4); Colab notebook on a T4 (#5)
  • v1.0 · llama.cpp port (GGML tensor type, CUDA mmv kernel, converter), the road to Ollama / LM Studio (#6)
  • v1.0 · paper, draft in docs/paper/ (#7)

FAQ

Can I run the checkpoint in LM Studio, Ollama or llama.cpp? Not yet. An NWC checkpoint is read by nwc.load_pretrained (PyTorch + Transformers). llama.cpp has no NWC tensor type; that port is on the roadmap and is what Ollama and LM Studio would need. vLLM, TGI and SGLang have no loader either.

Do I need NWC to use the model? Only while it is compressed. python -m nwc.export CHECKPOINT OUT gives you the original BF16 checkpoint back, bit for bit, and from there every tool works as usual. There is no lock-in.

Is it really lossless? Yes. Dequantized weights equal the originals (torch.equal, tested on every commit for the format and on the GPU before releases). Only the fp32 summation order of the matvec differs from cuBLAS, the same way cuBLAS differs between two GPUs.

Does it make prefill or batched inference faster? No. Batch 1 (token generation) is the fast path; prefill dequantizes into a temporary BF16 buffer and calls cuBLAS, so it only saves memory.

What about int8 / fp8 models? fp8 e4m3 weight-only is supported (elem="fp8"): the fp8 values are stored lossless at 0.87 of their size, decoded by the same kernel, at native-fp8 speed on the 4070 (0.7× on the A16, where the decoder is the limit). int8 was measured and not built: its symbol alphabet is flat, the prefix code would save only 5 %, an ideal coder 13 %. Numbers in docs/results.md, sections 7 and 8.

Which models work? Any HF causal LM in BF16; the encoder only needs the weights. Tested: Qwen2.5 (3B, 7B), Qwen3 (0.6B, 4B). Small models (0.6B) save memory but are not faster: their matrices are launch-bound.

My GPU is not in the table. Run it and tell us. Turing builds but is unmeasured; Blackwell runs via PTX JIT.

Something fails. python -m nwc.doctor first; paste its output into a bug report.

Install from source

git clone https://github.com/parda21/NWC && cd NWC
python -m nwc.build                 # fatbin sm_75..sm_90 + PTX into nwc/lib (needs nvcc 12.6+; 13.x on Windows)
pip install -e .

Repository builds for the local GPU only: .\build.ps1 (Windows, Visual Studio 2022 Build Tools) or ./build.sh sm_86 (Linux) write build/nwc_ops.dll|so, which the tests and scripts pick up.

python tests/test_format_cpu.py      # no GPU: reference decoder, bit-exact
python scripts/make_wraw.py          # test matrix data/W.raw from a local model
python tests/test_k.py               # GPU: dequantization bit-exact, token path, fp32 path, odd shapes
python tests/test_gather.py          # GPU: embedding lookup
python tests/test_checkpoint.py      # GPU: save -> load -> export round trip, no download
python scripts/kernbench.py --runs 30                              # kernel vs cuBLAS per layer shape
python scripts/graph_decode.py --mode nwc --fusion --tokens 256    # tokens/s as a CUDA graph, vs HF generate
csrc/nwc_ops.cu   the kernel library: encoder, fused matvec, dequantization, gather
nwc/              Python package: nwc_torch (NWCLinear, convert, fuse), checkpoint (save/load/export), demo, doctor, build
tests/ scripts/   correctness tests; benchmarks the numbers above come from
docs/             results.md, format.md, model cards, paper/, announce.md
experiments/      kernel iterations v1–v9, micro-benchmarks, earlier formats: history, not product

For agents

AGENTS.md is the rule book for coding agents (and humans): repository map, which tests need a GPU and which do not, how kernel changes are measured, the conventions, and the pitfalls we already hit. A CLAUDE.md points there. tests/test_format_cpu.py lets an agent without a GPU verify the format bit-exactly.

Contributing, citing, sponsor, license

Contributions: see CONTRIBUTING.md; benchmark results from GPUs we do not have are the most useful thing right now. Cite with the CITATION.cff (GitHub's "Cite this repository" button). This project is sponsored by cloo GmbH. Apache License 2.0, Copyright 2026 Paul Otto, see LICENSE.

About this project. Was it created with the help of AI? Yes. Do I care, as long as it works? No.

Release files for neural-weight-compression 0.10.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distributions (wheels)

Table of built distributions (wheels) for neural-weight-compression 0.10.1
File Interpreter ABI Platform
neural_weight_compression-0.10.1-py3-none-win_amd64.whl Python 3 none Windows x86-64 Details
neural_weight_compression-0.10.1-py3-none-manylinux_2_35_x86_64.whl Python 3 none Linux glibc 2.35+ x86-64 Details

Total release size: 1.6 MB

Release files / neural_weight_compression-0.10.1-py3-none-win_amd64.whl

Download URL neural_weight_compression-0.10.1-py3-none-win_amd64.whl
Size 691.7 kB
Tags Python 3 Windows x86-64
SHA-256 checksum
How to use checksums
268c52e622d43a8099009b735cf13549e6d37cdf5cb4b5bd556e0954c14b9504
BLAKE2b-256 checksum
How to use checksums
aa3e45f89ea83705cbf646cc9baa792211f550b7925f3df8dccd6b7e09a87084
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 19, 2026.

Transparency log

Release files / neural_weight_compression-0.10.1-py3-none-manylinux_2_35_x86_64.whl

Download URL neural_weight_compression-0.10.1-py3-none-manylinux_2_35_x86_64.whl
Size 868.5 kB
Tags Linux glibc 2.35+ x86-64 Python 3
SHA-256 checksum
How to use checksums
f907f0abb620d2d62e1e4804aa5ac4e5c45e1fdc53840e709a24f79c937ecb88
BLAKE2b-256 checksum
How to use checksums
311dcf96c1ad014483e2458dd6890823988eb2668f1575c65b65934015e1616c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 19, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.10.1 This release

2 release files

0.10.0

2 release files

0.9.1

2 release files

0.9.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page