Last released Sep 5, 2026
AirLLM runs 70B large language models on a single 4GB GPU without quantization, distillation or pruning. Kimi K3 2.8T on under 4GB, Qwen3.8-Flash-Next 125B on 6GB, DeepSeek-V3 671B on ~12GB. Train Qwen3.8-Flash-Next under 6GB.