2 projects
pytllm
TLLM runs 70B large language models on a single 4GB GPU without quantization, distillation or pruning. 405B Llama 3.1 on 8GB, DeepSeek-V3 671B on ~12GB, Kimi K3 2.8T on under 4GB, Qwen3.8-27B on 3.3GB.
cognitivess
Official Python SDK for the CognitivessAI API (OpenAI- and Anthropic-compatible).