Skip to main content

TurboAgent

TurboQuant-powered agentic AI framework for long-context LLMs on consumer hardware.

TurboAgent is a pip-installable Python package that brings Google Research's TurboQuant KV-cache compression to open-source LLMs for local, consumer-hardware agentic AI. It delivers 6x+ memory reduction and up to 8x attention speedup with zero measurable accuracy loss.

Features

  • One-line agent creation with 6x+ KV compression -- 32k-1M+ effective context on a single RTX 4090
  • Hardware-aware auto-tuning -- detects CUDA/ROCm/Metal/CPU and selects optimal configuration
  • Agentic-first primitives -- persistent multi-turn memory, RAG with vector-search, multi-agent swarms
  • Multiple backends -- llama.cpp (consumer GPUs), vLLM (server throughput), PyTorch (research)
  • Zero-calibration, training-free -- just like the paper guarantees

Quick Start

pip install turboagent-ai[llama]
from turboagent import TurboAgent

agent = TurboAgent(
    "meta-llama/Llama-3.1-70B-Instruct",
    kv_mode="turbo3",
    context=131072,
)

response = agent.run("Analyze my 50k-token research doc and suggest experiments...")
print(response)  # KV usage <4 GB total

Installation

# Core + llama.cpp backend (recommended for consumer GPUs)
pip install turboagent-ai[llama]

# With vLLM for server-style throughput
pip install turboagent-ai[vllm]

# With HuggingFace Transformers for research
pip install turboagent-ai[torch]

# With native TurboQuant C++/CUDA kernels (recommended for best performance)
pip install turboagent-ai[native]

# Development
pip install turboagent-ai[dev]

CLI

# Scaffold a new agent project
turboagent init my_agent

# Detect hardware and show optimal configuration
turboagent info

# Run benchmarks
turboagent benchmark --model-size 70

Multi-Agent Swarms

from turboagent.agents.swarm import TurboSwarm, SwarmAgent

swarm = TurboSwarm(
    "meta-llama/Llama-3.1-70B-Instruct",
    agents=[
        SwarmAgent(name="researcher", role="deep research"),
        SwarmAgent(name="critic", role="critical review"),
        SwarmAgent(name="writer", role="clear writing"),
    ],
)

results = swarm.run("Analyze the latest advances in KV cache compression.")

RAG with TurboVectorStore

from turboagent.agents.rag import TurboVectorStore

store = TurboVectorStore(embedding_dim=768)
store.add_documents(texts=chunks, embeddings=embeddings)
results = store.query(query_embedding, top_k=5)

Architecture

turboagent/
├── quant/          # TurboQuantKVCache (PolarQuant + QJL)
├── backends/       # llama.cpp, vLLM, PyTorch engines
├── agents/         # TurboAgent, TurboVectorStore, TurboSwarm
├── hardware/       # Auto-detection and optimal config
├── cli.py          # Project scaffolding and benchmarks
└── utils.py        # Shared helpers

TurboQuant Compression Modes

Mode Bits per Value Compression Best For
turbo3 3.25 bpv 4.9x Maximum context on limited VRAM
turbo4 4.25 bpv 3.8x Higher quality, ample memory

Requirements

  • Python >= 3.10
  • PyTorch >= 2.5.0
  • One of: llama-cpp-python, vLLM, or HuggingFace Transformers

Development

git clone https://github.com/TurboAgentAI/turboagent.git
cd turboagent
pip install -e ".[dev]"
pytest tests/ -v -m "not integration"

Enterprise

The open-source core is free forever under the MIT license.

TurboAgent Enterprise adds commercial extensions for teams and organizations:

  • SSO / SAML authentication
  • Audit logging and compliance exports (SOC-2, GDPR)
  • Air-gapped on-premise licensing
  • SecureMultiAgentSwarm with governance policies and RBAC
  • Multi-node KV cache sharing
  • Priority kernels and dedicated support SLAs
# Enterprise features activate with a license key
# export TURBOAGENT_LICENSE_KEY="TA-ENT-your-key-here"

from turboagent.enterprise.swarm import SecureMultiAgentSwarm
from turboagent.enterprise.audit import AuditLogger

Learn more: turboagent.to/enterprise | Contact: enterprise@turboagent.to

License

MIT — the open-source core is free for commercial and personal use. Commercial extensions are available under a separate license. See Enterprise.

Acknowledgments

Built on community TurboQuant implementations:

Metadata

Release files for turboagent-ai 1.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for turboagent-ai 1.1.0
File Size Uploaded
turboagent_ai-1.1.0.tar.gz 84.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for turboagent-ai 1.1.0
File Interpreter ABI Platform
turboagent_ai-1.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 153.3 kB

Release files / turboagent_ai-1.1.0.tar.gz

Download URL turboagent_ai-1.1.0.tar.gz
Size 84.1 kB
Tags Source
SHA-256 checksum
How to use checksums
f2e15eff01c1b576f233c15ff10e10b29f15bbf7cda64fe7456a6f4089c8be38
BLAKE2b-256 checksum
How to use checksums
f4284e30fdfd7a46da20d735f487c0f46222fe1ecc982f27dd0c71163748ff0e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.12

Release files / turboagent_ai-1.1.0-py3-none-any.whl

Download URL turboagent_ai-1.1.0-py3-none-any.whl
Size 69.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0c771cdac07f35f4d56201462077595ac343996042346b6de348aaf00ef1d7c6
BLAKE2b-256 checksum
How to use checksums
e668031d2899a9b5601d50ebdac5ebb4686b3dc596800ef767fef6cc18a03629
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.12

Release history Release notifications | RSS feed

This release

1.1.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page