StreamLLM
StreamLLM optimizes inference memory usage, allowing large language models (such as 14B, 32B, and 70B models) to run on consumer GPUs with as little as 4GB or 8GB VRAM without requiring distributed hardware.
Instead of loading the entire model into GPU memory at once, StreamLLM dynamically streams transformer layers between system RAM/disk and GPU VRAM, using asynchronous prefetching and double buffering so computation overlaps with data movement.
Quickstart
1. Install package
pip install streamllm
2. Discover and install a model
# View available curated models and installation status
streamllm list
# Download and install a model to your local cache
streamllm install smollm-135m
# Or install mid-sized/larger models:
streamllm install qwen-0.5b
streamllm install qwen-14b
3. Inference via Python SDK
Run inference with automatic layer streaming:
from streamllm import AutoModel
# Load any installed preset or HuggingFace repo ID
model = AutoModel.from_pretrained("smollm-135m")
tokens = model.tokenizer("Hello, what is layer streaming?", return_tensors="pt")
output = model.generate(tokens["input_ids"], max_new_tokens=40)
print(model.tokenizer.decode(output.sequences[0]))
You can pass in any preset alias, HuggingFace model repo ID, or local directory path.
Model Management CLI
StreamLLM provides full model management (like Ollama or Docker) for discovering and downloading open-source weights:
# List all available presets, parameters, memory specs, and install status
streamllm list
# List only locally installed models
streamllm list --installed
# Download & install a model (preset alias or HuggingFace repo ID)
streamllm install qwen-0.5b
streamllm install HuggingFaceTB/SmolLM2-135M-Instruct
streamllm pull llama3.2-1b
# Remove an installed model from local cache
streamllm models remove qwen-0.5b
Curated Model Presets
| Preset | Model Repo ID | Parameters | Est. Download | StreamLLM Min VRAM | Standard VRAM |
|---|---|---|---|---|---|
smollm-135m |
HuggingFaceTB/SmolLM2-135M-Instruct |
135M | ~270 MB | ~120 MB | ~600 MB |
smollm-360m |
HuggingFaceTB/SmolLM2-360M-Instruct |
360M | ~720 MB | ~220 MB | ~1.4 GB |
smollm-1.7b |
HuggingFaceTB/SmolLM2-1.7B-Instruct |
1.7B | ~3.4 GB | ~420 MB | ~4.5 GB |
qwen-0.5b |
Qwen/Qwen2.5-0.5B-Instruct |
0.5B | ~980 MB | ~260 MB | ~2.0 GB |
qwen-1.5b |
Qwen/Qwen2.5-1.5B-Instruct |
1.5B | ~3.1 GB | ~400 MB | ~4.8 GB |
llama3.2-1b |
meta-llama/Llama-3.2-1B-Instruct |
1.2B | ~2.4 GB | ~380 MB | ~3.6 GB |
llama3.2-3b |
meta-llama/Llama-3.2-3B-Instruct |
3.2B | ~6.4 GB | ~580 MB | ~8.0 GB |
qwen-7b |
Qwen/Qwen2.5-7B-Instruct |
7B | ~14.5 GB | ~850 MB | ~16.0 GB |
qwen-14b |
Qwen/Qwen2.5-14B-Instruct-AWQ |
14B (AWQ) | ~8.5 GB | ~1200 MB | ~10.0 GB |
llama3-8b |
meta-llama/Meta-Llama-3.1-8B-Instruct |
8B | ~16.0 GB | ~950 MB | ~18.0 GB |
mistral-7b |
mistralai/Mistral-7B-Instruct-v0.3 |
7.3B | ~14.5 GB | ~900 MB | ~16.5 GB |
deepseek-1.5b |
deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B |
1.5B | ~3.1 GB | ~420 MB | ~4.8 GB |
deepseek-7b |
deepseek-ai/DeepSeek-R1-Distill-Qwen-7B |
7B | ~14.5 GB | ~880 MB | ~16.5 GB |
Interactive Chat & One-Shot Run
StreamLLM includes an interactive terminal chat REPL and command-line runner with real-time token streaming:
# 1. Interactive Chat REPL
streamllm chat smollm-135m
streamllm chat qwen-0.5b
streamllm chat qwen-14b
# If you don't provide a model name, StreamLLM automatically selects your installed model!
streamllm chat
# Inside chat, you have built-in slash commands:
# /stats - Check live GPU VRAM & scratchpad allocation
# /system - Update the system instruction prompt
# /clear - Clear conversation history
# /exit - Quit chat
# 2. One-shot command execution
streamllm run smollm-135m "Explain layer streaming in 2 sentences"
streamllm run qwen-0.5b "Summarize quantum computing" --max-tokens 100
# 3. Pipeable raw output for scripts
streamllm run smollm-135m "Generate JSON" --raw
# 4. Check GPU VRAM and live PCIe bandwidth
streamllm hardware
# 5. Run streaming micro-benchmark
streamllm bench --layers 16 --hidden-dim 2048 --seq-len 128
How It Works
During inference, StreamLLM keeps only the currently executing layer on the GPU while asynchronously prefetching upcoming layers over PCIe DMA:
- Static GPU Scratchpad Pool (
GPUScratchpadPool): Pre-allocates two static VRAM buffer slots (Slot A and Slot B). EliminatescudaMallocandcudaFreechurn during generation. - Dual CUDA Streams: Dedicated
compute_streamandtransfer_streamwith lock-free hardware event synchronization (torch.cuda.Event). - Pinned Host Memory: Uses page-locked RAM (
torch.Tensor.pin_memory()) for true non-blocking PCIe transfers. - KV-Cache Manager: Deterministic memory budgeting that ensures context length growth never causes an Out-of-Memory (OOM) crash.
Benchmark Performance
Measured Hardware: NVIDIA GeForce RTX 3050 Laptop GPU (VRAM: 4096 MB, PCIe: 8.65 GB/s)
| Layers | Model Weight Size | Sequential Latency | Prefetch Latency | Speedup | GPU VRAM Scratchpad |
|---|---|---|---|---|---|
| 8 Layers | 512.1 MB | 68.08 ms | 56.70 ms | 1.20x | 128.0 MB |
| 16 Layers | 1024.1 MB | 136.51 ms | 126.11 ms | 1.08x | 128.0 MB |
| 24 Layers | 1536.2 MB | 219.05 ms | 170.54 ms | 1.28x | 128.0 MB |
| 32 Layers | 2048.2 MB | 278.85 ms | 265.03 ms | 1.05x | 128.0 MB |
Key Takeaway: Even as total model weights scale past 2.0 GB, StreamLLM's static scratchpad pool holds physical GPU memory strictly at 128.0 MB while asynchronous prefetching delivers up to 1.28x faster inference.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file streamllm-0.1.2.tar.gz.
File metadata
- Download URL: streamllm-0.1.2.tar.gz
- Upload date:
- Size: 29.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8e8ce551bd3d66df699a7a5700d645e7cda3a61dfa35b9519088b9fdb88385d4
|
|
| MD5 |
b10fff15ad803d0427838da5fb728322
|
|
| BLAKE2b-256 |
4c924a51b939423aadb058d1cde50287ad35ae46db5b07d892c4e3cd167c77e0
|
Provenance
The following attestation bundles were made for streamllm-0.1.2.tar.gz:
Publisher:
publish.yml on prathamc00/streamLLM
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
streamllm-0.1.2.tar.gz -
Subject digest:
8e8ce551bd3d66df699a7a5700d645e7cda3a61dfa35b9519088b9fdb88385d4 - Sigstore transparency entry: 2706671681
- Sigstore integration time:
-
Permalink:
prathamc00/streamLLM@ff8f8d9cd9ebeae6926482b5e25e6a816387962b -
Branch / Tag:
refs/tags/v0.1.2 - Owner: https://github.com/prathamc00
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@ff8f8d9cd9ebeae6926482b5e25e6a816387962b -
Trigger Event:
release
-
Statement type:
File details
Details for the file streamllm-0.1.2-py3-none-any.whl.
File metadata
- Download URL: streamllm-0.1.2-py3-none-any.whl
- Upload date:
- Size: 31.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
dfb96ae0faee355109926c77a2f88612b5d057a935ea1593dfeb66ac912a7ce7
|
|
| MD5 |
043cbcbd2b97289dc8a288f3fe1b20b7
|
|
| BLAKE2b-256 |
9006d9a3db39eacdc1d4c1966fe62562a28524b3bec9220ab86d7e48cd9089b7
|
Provenance
The following attestation bundles were made for streamllm-0.1.2-py3-none-any.whl:
Publisher:
publish.yml on prathamc00/streamLLM
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
streamllm-0.1.2-py3-none-any.whl -
Subject digest:
dfb96ae0faee355109926c77a2f88612b5d057a935ea1593dfeb66ac912a7ce7 - Sigstore transparency entry: 2706671691
- Sigstore integration time:
-
Permalink:
prathamc00/streamLLM@ff8f8d9cd9ebeae6926482b5e25e6a816387962b -
Branch / Tag:
refs/tags/v0.1.2 - Owner: https://github.com/prathamc00
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@ff8f8d9cd9ebeae6926482b5e25e6a816387962b -
Trigger Event:
release
-
Statement type: