Skip to main content

🧠 Bit-TTT-Engine

PyPI License: MIT Rust

Fast local LLM inference that learns while it runs.

  • 🏎️ 47+ tok/s on RTX 4060 Ti (7B Q4_K_M)
  • 🧠 TTT (Test-Time Training) — adapts during inference (world's first!)
  • 🎨 LoRA — fine-tune with one flag
  • 📦 5 models — Llama-2/3, Gemma-2, Qwen2.5, Mistral
  • 🔌 OpenAI-compatible API — drop-in replacement

🚀 Quick Start

pip install bit-ttt-engine
import cortex_rust

# Load any GGUF model (auto-downloads from HuggingFace!)
model = cortex_rust.load("user/model-GGUF")

# Chat
response = model.chat([
    {"role": "user", "content": "Hello!"}
])
print(response)

# Stream
for token in model.chat_stream([
    {"role": "user", "content": "Tell me a story"}
]):
    print(token, end="", flush=True)

🖥️ CLI

# Interactive chat
bit-ttt chat model.gguf

# Generate text
bit-ttt generate model.gguf -p "Once upon a time"

# OpenAI-compatible API server
bit-ttt serve model.gguf --port 8000

# With LoRA + Q8 KV cache
bit-ttt chat model.gguf --lora adapter.bin --q8-cache

🧠 TTT — Test-Time Training

The model learns while it generates. No other local LLM does this.

model = cortex_rust.load("model.gguf")
model.enable_ttt(True)

# Each conversation makes the model smarter
response = model.chat([{"role": "user", "content": "My name is Alice"}])
# Next time, it remembers context better!

⚡ Performance

Model Speed VRAM
Llama-2 7B (Q4_K_M) 47.8 tok/s ~5 GB
Llama-3 8B (Q4_K_M) 36.8 tok/s ~6 GB
Mistral 7B (Q4_K_M) 40.8 tok/s ~5 GB
Qwen2.5 1.5B (Q4_K_M) 70.4 tok/s ~2 GB

With --q8-cache: 82% VRAM reduction for KV cache.

🔌 OpenAI-Compatible API

bit-ttt serve model.gguf --port 8000
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="none")
response = client.chat.completions.create(
    model="default",
    messages=[{"role": "user", "content": "Hi!"}],
    stream=True,
)

📖 Links

💖 License

MIT License

Metadata

Release files for bit-ttt-engine 0.8.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for bit-ttt-engine 0.8.0
File Size Uploaded
bit_ttt_engine-0.8.0.tar.gz 406.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for bit-ttt-engine 0.8.0
File Interpreter ABI Platform
bit_ttt_engine-0.8.0-cp310-cp310-win_amd64.whl CPython 3.10 CPython 3.10 Windows x86-64 Details

Total release size: 3.4 MB

Release files / bit_ttt_engine-0.8.0.tar.gz

Download URL bit_ttt_engine-0.8.0.tar.gz
Size 406.0 kB
Tags Source
SHA-256 checksum
How to use checksums
8a4e713dd403fff25e88843220a6366135b5ba72317826bd4bf5a05c6661cdb8
BLAKE2b-256 checksum
How to use checksums
2e918bd2b566e42641d2956f05a816f6538803a1e278deef80e97c2e741dfde2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.11.3

Release files / bit_ttt_engine-0.8.0-cp310-cp310-win_amd64.whl

Download URL bit_ttt_engine-0.8.0-cp310-cp310-win_amd64.whl
Size 3.0 MB
Tags CPython 3.10 Windows x86-64
SHA-256 checksum
How to use checksums
7f3e75d71a9e6bdbb3f0f919d76a52813789c98e9d4aff2121ca47a483933dc1
BLAKE2b-256 checksum
How to use checksums
4f83f4657b91ce234d25c6b8805f6eb363cc4a7dddfde89dd5655402c98efc4d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.11.3

Release history Release notifications | RSS feed

This release

0.8.0 This release

2 release files

0.7.0

2 release files

0.6.2

2 release files

0.6.1

1 release file

0.6.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page