A local, MoE-offload inference runtime with an OpenAI- and Anthropic-compatible HTTP API — Run DeepSeek-V4-Flash on your 5090 with 20+ TPS.
Quick start
See docs/install.md for requirements and installation.
ft serve --model ~/models/Qwen3.6-35B-A3B # API server on http://127.0.0.1:1919
ft launch claude # point an agent at it (codex / opencode / openclaw)
ft shell # or chat in the terminal
Documentation
- Install — requirements and setup
- Supported models — model × quantization
- CLI reference —
ftcommands and environment variables
Acknowledgment
FreeToken was deeply inspired by mini-sglang, and learned the design and reused code from the following projects: SGLang, vLLM, FlashInfer, flash-linear-attention, LightLLM and llama.cpp.
License
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
freetoken-0.1.1.tar.gz
(844.7 kB
view details)
File details
Details for the file freetoken-0.1.1.tar.gz.
File metadata
- Download URL: freetoken-0.1.1.tar.gz
- Upload date:
- Size: 844.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.13.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
559f85eea93b69955caa02926a1cb198508b2ab00c662f8e6160ff634204599e
|
|
| MD5 |
ef8d4a216adcfd71734022b545acaa47
|
|
| BLAKE2b-256 |
fcfdf964cb970d512b3dab1bbeec95accd06c40cd1a40714cfca5409aa4fcdc2
|