Skip to main content

smol-vllm

PyPI version Python 3.10+ License: MIT GitHub

Paged-attention inference engine: KV cache, continuous batching, preemption. Educational, not production.

Install

pip install smol-vllm

Real models (TinyLlama, Qwen2, etc.):

pip install smol-vllm[tinyllama-1.1b]
# or
pip install smol-vllm[qwen2-0.5b]

Quick Start

FakeModel (no extras):

from smol_vllm import LLMEngine

engine = LLMEngine()
for token in engine.generate([1, 2, 3, 4, 5], max_tokens=20):
    print(token, end=" ")

CausalLM (needs [tinyllama-1.1b] or [qwen2-0.5b]):

engine = LLMEngine(use_real_model=True)
tokenizer = engine.model.tokenizer
tokens = tokenizer.encode("Hello!", add_special_tokens=False)
for token in engine.generate(tokens, max_tokens=20):
    print(tokenizer.decode([token]), end="")

Models

Model model_name
TinyLlama 1.1B TinyLlama/TinyLlama-1.1B-Chat-v1.0 (default)
Qwen2 0.5B Qwen/Qwen2-0.5B-Instruct
Phi-2 microsoft/phi-2
Llama 3.2 meta-llama/Llama-3.2-1B-Instruct
Gemma 2 google/gemma-2-2b-it
Mistral mistralai/Mistral-7B-Instruct-v0.3

Gated models (Llama, Gemma, etc.) need a HuggingFace token. Options:

1. Env var (recommended):

export HF_TOKEN=hf_xxxxxxxxxxxx

2. In code:

LLMEngine(use_real_model=True, model_name="meta-llama/Llama-3.2-1B-Instruct", hf_token="hf_xxxx")

Get a token: huggingface.co/settings/tokens. Accept the model's license on its HF page first.

Demo

smol-vllm-demo

What It Teaches

  • PagedAttention — block-based KV cache, ref counting
  • Continuous batching — short jobs fill slots immediately
  • Preemption & swapping — when memory runs low
  • Prefill vs decode — compute-bound → memory-bound

Workflow: run with FakeModel first (zero deps), then switch to CausalLM to compare.

Metrics

Step-level: prefill/decode latency, tok/s, KV util. Summary and CSV logs in logs/.

License

MIT

Metadata

Release files for smol-vllm 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for smol-vllm 0.1.2
File Size Uploaded
smol_vllm-0.1.2.tar.gz 13.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for smol-vllm 0.1.2
File Interpreter ABI Platform
smol_vllm-0.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 29.1 kB

Release files / smol_vllm-0.1.2.tar.gz

Download URL smol_vllm-0.1.2.tar.gz
Size 13.9 kB
Tags Source
SHA-256 checksum
How to use checksums
1f54d9f780c066e923fe202b9e86b794ab2b8233dd5a40fd3ac5eec8f777138d
BLAKE2b-256 checksum
How to use checksums
e8bd90e8214c030ac378e855f4ea3690986ae3539154b1d6cb23d77399187480
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.5

Release files / smol_vllm-0.1.2-py3-none-any.whl

Download URL smol_vllm-0.1.2-py3-none-any.whl
Size 15.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
14c90bfbffe7d8ba09ba70d93a5e103c7ca2793398c7b0c7c464d13075ed0184
BLAKE2b-256 checksum
How to use checksums
a37edab7da883d2bade230cf91fb6c537413bd484e9875d5a4fb46d2f7ccfa90
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.5

Release history Release notifications | RSS feed

This release

0.1.2 This release

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page