Skip to main content

Fast LLM

Run any AI model faster on CPU — same output, zero training.

pip install fastllm-turbo
fast-llm calibrate Qwen/Qwen2.5-0.5B

How it works

Most of a neural network's weights don't need full precision. We use a per-layer reconstruction error metric to identify which layers are sensitive to precision loss and which are robust. Sensitive layers stay FP32. Robust layers use INT8 dynamic quantization.

The result: a mixed-precision model that outputs the same tokens as the original, but runs faster because ~60% of computation uses 4× smaller weights.

Method Speed Output matches original?
INT8 all layers ~2× No (quality degrades on small models)
Mixed-precision ~1.6× Yes
FP32 (original) 1× —

Install

pip

pip install fastllm-turbo

Linux / macOS

curl -sSf https://fast-llm.dev/install.sh | sh

Windows (PowerShell)

powershell -c "iex ((New-Object Net.WebClient).DownloadString('https://fast-llm.dev/install.ps1'))"

Usage

# Step 1: Analyze which layers can be INT8
fast-llm calibrate Qwen/Qwen2.5-0.5B

# Step 2: Save optimized model
fast-llm save Qwen/Qwen2.5-0.5B

# Step 3: Run
fast-llm run Qwen/Qwen2.5-0.5B --benchmark

# Interactive chat
fast-llm chat Qwen/Qwen2.5-0.5B

# Desktop GUI
fast-llm gui

Works with any HuggingFace causal LM: Qwen, Llama 3, Phi-3, Mistral, Gemma, DeepSeek.

Requirements

  • Python 3.8+
  • PyTorch 2.0+ (CPU)
  • 1–4 GB RAM per billion parameters

License

MIT

Metadata

Release files for fastllm-turbo 2.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for fastllm-turbo 2.0.1
File Size Uploaded
fastllm_turbo-2.0.1.tar.gz 9.9 kB Details

Release files / fastllm_turbo-2.0.1.tar.gz

Download URL fastllm_turbo-2.0.1.tar.gz
Size 9.9 kB
Tags Source
SHA-256 checksum
How to use checksums
d1ef0a4b0f08ca499d3cb22a3f718f2d81279a68e227da196674674c149a73da
BLAKE2b-256 checksum
How to use checksums
1025c07db5ee2d3990205a33a3c0d942ba26d0a493e21f984bf14f36cf2ef645
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.6

Release history Release notifications | RSS feed

This release

2.0.1 This release

1 release file

2.0.0

1 release file

1.0.0

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page