Last released Jul 27, 2026
Five-value sub-2-bit LLMs: chat with a ~2 GB 8B container at native-runtime speed, load it as a Transformers model, or serve it on an OpenAI-compatible endpoint
Last released Jul 25, 2026
Fermion Research — Neutrino: frontier-grade models below two bits per weight. Full release imminent.
Supported by