SketchSSM
SketchSSM: Write to the Full State, Read from a Compact Sketch
SketchSSM speeds up decoding in linear-attention and state-space layers (Mamba-2, Gated DeltaNet, KDA). It keeps updating the full recurrent state, but between periodic exact flushes it reads a compact low-rank sketch of the state instead of the full state, which cuts state-read traffic. The sketch size, which sets the traffic reduction, is chosen in the serving configuration; the sketching matrix is computed once per model by offline calibration.
SketchSSM over a window. At a flush step, SketchSSM updates the full state S0 and refreshes the sketch U. At non-flush steps, it multiplies U by query-dependent coefficients ct to reconstruct the output; S0 is not accessed between state updates.
This repository contains:
vllm/: vLLM v0.30.0 with SketchSSM.sketchssm/kernels/: the SketchSSM CUDA decode kernels.sketchssm/calibration/: the offline calibration.benchmarks/: per-layer decode latency in vLLM.
Installation
git clone https://github.com/SNU-ARC/SketchSSM.git
cd SketchSSM/vllm
VLLM_USE_PRECOMPILED=1 python -m pip install -e .
python -m pip install sketchssm # CUDA kernels; without it vLLM uses its Triton kernels
Quick start: serve with vLLM
SketchSSM needs a calibration file (calibration.pt) made by offline
calibration. Calibrations for the models below are on the Hugging Face Hub.
To calibrate a new model yourself, see sketchssm/calibration/.
| Model | Weights | Calibration |
|---|---|---|
| Nemotron Nano 9B v2-BF16 | nvidia/NVIDIA-Nemotron-Nano-9B-v2 | ominn/SketchSSM-Nemotron-Nano-9B-v2-BF16 |
| Nemotron 3 Super-NVFP4 | nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 | ominn/SketchSSM-Nemotron-3-Super-NVFP4 |
| Qwen3.8 Flash-Next-NVFP4 | RadixArk/Qwen3.8-Flash-Next-NVFP4 | ominn/SketchSSM-Qwen3.8-Flash-Next-NVFP4 |
| GLM 5.3 Flash-NVFP4 | RedHatAI/GLM-5.3-Flash-NVFP4 | ominn/SketchSSM-GLM-5.3-Flash-NVFP4 |
| Qwen3.5 9B-BF16 | Qwen/Qwen3.5-9B | ominn/SketchSSM-Qwen3.5-9B-BF16 |
See all calibrations.
Enable SketchSSM with:
--sketchssm: the calibration, a Hub repo id or a local file.--sketchssm-mean-rank: the rank budget, the mean sketch rank per state head (default 10: about 10x less state traffic at standard decode accuracy in the paper).
# Calibration from the Hub
vllm serve nvidia/NVIDIA-Nemotron-Nano-9B-v2 --trust-remote-code \
--sketchssm ominn/SketchSSM-Nemotron-Nano-9B-v2-BF16 --sketchssm-mean-rank 10 \
--mamba-ssm-cache-dtype float32 --no-enable-prefix-caching
# Local calibration file
vllm serve nvidia/NVIDIA-Nemotron-Nano-9B-v2 --trust-remote-code \
--sketchssm outputs/nano/calibration.pt --sketchssm-mean-rank 10 \
--mamba-ssm-cache-dtype float32 --no-enable-prefix-caching
See vllm/README.md for all options.
Benchmarks
See benchmarks/ to measure the per-layer recurrent
decode latency of a model in vLLM, and sketchssm/kernels/
for kernel tuning and microbenchmarks.
Citation
If you use SketchSSM in your research, please cite:
@misc{kwon2026sketchssmwritestateread,
title={SketchSSM: Write to the Full State, Read from a Compact Sketch},
author={Omin Kwon and JoongWon Shin and Minseo Kim and Kurt Keutzer and Sehoon Kim and Jae W. Lee},
year={2026},
eprint={2609.33051},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2609.33051},
}
License
SketchSSM is released under the Apache License 2.0. The
vllm/ directory is a fork of vLLM, also under Apache-2.0. Calibration
files are derived from the base models' weights and are also subject to those
models' licenses.
Metadata
Release files for sketchssm 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| sketchssm-0.1.1.tar.gz | 214.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| sketchssm-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 468.2 kB
Release files / sketchssm-0.1.1.tar.gz
| Download URL | sketchssm-0.1.1.tar.gz |
|---|---|
| Size | 214.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
54c22e9dac8b28ddbfa65b3c051d00bcfde4634b16c3501ca2d8bf814c4c28ff
|
|
BLAKE2b-256 checksum How to use checksums |
24f2fabe5d031b24aab6c4c3610917fdf4c0edadd0a52c31603618ab14346601
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.14
|
Release files / sketchssm-0.1.1-py3-none-any.whl
| Download URL | sketchssm-0.1.1-py3-none-any.whl |
|---|---|
| Size | 254.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
85253aac66e339d4a04760c517d9f53c14b031d4a3fd2d501c73569fa3655af0
|
|
BLAKE2b-256 checksum How to use checksums |
f52ad4515f1bd822edd05c3ef7d355d89db337a0cf83cb56187d1582af152690
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.14
|