Skip to main content

SketchSSM

SketchSSM: Write to the Full State, Read from a Compact Sketch

SketchSSM speeds up decoding in linear-attention and state-space layers (Mamba-2, Gated DeltaNet, KDA). It keeps updating the full recurrent state, but between periodic exact flushes it reads a compact low-rank sketch of the state instead of the full state, which cuts state-read traffic. The sketch size, which sets the traffic reduction, is chosen in the serving configuration; the sketching matrix is computed once per model by offline calibration.

SketchSSM over a window

SketchSSM over a window. At a flush step, SketchSSM updates the full state S0 and refreshes the sketch U. At non-flush steps, it multiplies U by query-dependent coefficients ct to reconstruct the output; S0 is not accessed between state updates.

Paper

This repository contains:

Installation

git clone https://github.com/SNU-ARC/SketchSSM.git
cd SketchSSM/vllm
VLLM_USE_PRECOMPILED=1 python -m pip install -e .
python -m pip install sketchssm   # CUDA kernels; without it vLLM uses its Triton kernels

Quick start: serve with vLLM

SketchSSM needs a calibration file (calibration.pt) made by offline calibration. Calibrations for the models below are on the Hugging Face Hub. To calibrate a new model yourself, see sketchssm/calibration/.

ModelWeightsCalibration
Nemotron Nano 9B v2-BF16nvidia/NVIDIA-Nemotron-Nano-9B-v2ominn/SketchSSM-Nemotron-Nano-9B-v2-BF16
Nemotron 3 Super-NVFP4nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4ominn/SketchSSM-Nemotron-3-Super-NVFP4
Qwen3.8 Flash-Next-NVFP4RadixArk/Qwen3.8-Flash-Next-NVFP4ominn/SketchSSM-Qwen3.8-Flash-Next-NVFP4
GLM 5.3 Flash-NVFP4RedHatAI/GLM-5.3-Flash-NVFP4ominn/SketchSSM-GLM-5.3-Flash-NVFP4
Qwen3.5 9B-BF16Qwen/Qwen3.5-9Bominn/SketchSSM-Qwen3.5-9B-BF16

See all calibrations.

Enable SketchSSM with:

  1. --sketchssm: the calibration, a Hub repo id or a local file.
  2. --sketchssm-mean-rank: the rank budget, the mean sketch rank per state head (default 10: about 10x less state traffic at standard decode accuracy in the paper).
# Calibration from the Hub
vllm serve nvidia/NVIDIA-Nemotron-Nano-9B-v2 --trust-remote-code \
  --sketchssm ominn/SketchSSM-Nemotron-Nano-9B-v2-BF16 --sketchssm-mean-rank 10 \
  --mamba-ssm-cache-dtype float32 --no-enable-prefix-caching

# Local calibration file
vllm serve nvidia/NVIDIA-Nemotron-Nano-9B-v2 --trust-remote-code \
  --sketchssm outputs/nano/calibration.pt --sketchssm-mean-rank 10 \
  --mamba-ssm-cache-dtype float32 --no-enable-prefix-caching

See vllm/README.md for all options.

Benchmarks

See benchmarks/ to measure the per-layer recurrent decode latency of a model in vLLM, and sketchssm/kernels/ for kernel tuning and microbenchmarks.

Citation

If you use SketchSSM in your research, please cite:

@misc{kwon2026sketchssmwritestateread,
  title={SketchSSM: Write to the Full State, Read from a Compact Sketch},
  author={Omin Kwon and JoongWon Shin and Minseo Kim and Kurt Keutzer and Sehoon Kim and Jae W. Lee},
  year={2026},
  eprint={2609.33051},
  archivePrefix={arXiv},
  primaryClass={cs.LG},
  url={https://arxiv.org/abs/2609.33051},
}

License

SketchSSM is released under the Apache License 2.0. The vllm/ directory is a fork of vLLM, also under Apache-2.0. Calibration files are derived from the base models' weights and are also subject to those models' licenses.

Metadata

Release files for sketchssm 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sketchssm 0.1.1
File Size Uploaded
sketchssm-0.1.1.tar.gz 214.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sketchssm 0.1.1
File Interpreter ABI Platform
sketchssm-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 468.2 kB

Release files / sketchssm-0.1.1.tar.gz

Download URL sketchssm-0.1.1.tar.gz
Size 214.0 kB
Tags Source
SHA-256 checksum
How to use checksums
54c22e9dac8b28ddbfa65b3c051d00bcfde4634b16c3501ca2d8bf814c4c28ff
BLAKE2b-256 checksum
How to use checksums
24f2fabe5d031b24aab6c4c3610917fdf4c0edadd0a52c31603618ab14346601
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.14

Release files / sketchssm-0.1.1-py3-none-any.whl

Download URL sketchssm-0.1.1-py3-none-any.whl
Size 254.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
85253aac66e339d4a04760c517d9f53c14b031d4a3fd2d501c73569fa3655af0
BLAKE2b-256 checksum
How to use checksums
f52ad4515f1bd822edd05c3ef7d355d89db337a0cf83cb56187d1582af152690
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.14

Release history Release notifications | RSS feed

0.1.2

2 release files

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page