Fast LLM
Run any AI model faster on CPU — same output, zero training.
pip install fastllm-turbo
fast-llm calibrate Qwen/Qwen2.5-0.5B
How it works
Most of a neural network's weights don't need full precision. We use a per-layer reconstruction error metric to identify which layers are sensitive to precision loss and which are robust. Sensitive layers stay FP32. Robust layers use INT8 dynamic quantization.
The result: a mixed-precision model that outputs the same tokens as the original, but runs faster because ~60% of computation uses 4× smaller weights.
| Method | Speed | Output matches original? |
|---|---|---|
| INT8 all layers | ~2× | No (quality degrades on small models) |
| Mixed-precision | ~1.6× | Yes |
| FP32 (original) | 1× | — |
Install
pip
pip install fastllm-turbo
Linux / macOS
curl -sSf https://fast-llm.dev/install.sh | sh
Windows (PowerShell)
powershell -c "iex ((New-Object Net.WebClient).DownloadString('https://fast-llm.dev/install.ps1'))"
Usage
# Step 1: Analyze which layers can be INT8
fast-llm calibrate Qwen/Qwen2.5-0.5B
# Step 2: Save optimized model
fast-llm save Qwen/Qwen2.5-0.5B
# Step 3: Run
fast-llm run Qwen/Qwen2.5-0.5B --benchmark
# Interactive chat
fast-llm chat Qwen/Qwen2.5-0.5B
# Desktop GUI
fast-llm gui
Works with any HuggingFace causal LM: Qwen, Llama 3, Phi-3, Mistral, Gemma, DeepSeek.
Requirements
- Python 3.8+
- PyTorch 2.0+ (CPU)
- 1–4 GB RAM per billion parameters
License
MIT
Metadata
Release files for fastllm-turbo 2.0.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| fastllm_turbo-2.0.1.tar.gz | 9.9 kB | Details |
Release files / fastllm_turbo-2.0.1.tar.gz
| Download URL | fastllm_turbo-2.0.1.tar.gz |
|---|---|
| Size | 9.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d1ef0a4b0f08ca499d3cb22a3f718f2d81279a68e227da196674674c149a73da
|
|
BLAKE2b-256 checksum How to use checksums |
1025c07db5ee2d3990205a33a3c0d942ba26d0a493e21f984bf14f36cf2ef645
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.14.6
|