Skip to main content
FreeToken

Slack Discord

A local, MoE-offload inference runtime with an OpenAI- and Anthropic-compatible HTTP API — Run DeepSeek-V4-Flash on your 5090 with 20+ TPS.

Quick start

See docs/install.md for requirements and installation.

ft serve --model ~/models/Qwen3.6-35B-A3B   # API server on http://127.0.0.1:1919
ft launch claude                            # point an agent at it (codex / opencode / openclaw)
ft shell                                    # or chat in the terminal

Documentation

Acknowledgment

FreeToken was deeply inspired by mini-sglang, and learned the design and reused code from the following projects: SGLang, vLLM, FlashInfer, flash-linear-attention, LightLLM and llama.cpp.

License

Apache License 2.0.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

freetoken-0.1.1.tar.gz (844.7 kB view details)

Uploaded Source

File details

Details for the file freetoken-0.1.1.tar.gz.

File metadata

  • Download URL: freetoken-0.1.1.tar.gz
  • Upload date:
  • Size: 844.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.13

File hashes

Hashes for freetoken-0.1.1.tar.gz
Algorithm Hash digest
SHA256 559f85eea93b69955caa02926a1cb198508b2ab00c662f8e6160ff634204599e
MD5 ef8d4a216adcfd71734022b545acaa47
BLAKE2b-256 fcfdf964cb970d512b3dab1bbeec95accd06c40cd1a40714cfca5409aa4fcdc2

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page