Skip to main content

Cottus Runtime

Cottus Runtime Logo

High-performance C++/CUDA LLM inference engine with Python bindings.

PyPI Python C++ CUDA License

Cottus Runtime is a custom inference engine built from scratch for Llama architectures, prioritizing low-latency and strict memory management. It implements its own Transformer execution pipeline, KV cache management, and attention kernels.

Features

  • Core: Custom C++20 Transformer implementation.
  • Memory: PagedAttention with BlockAllocator for efficient KV cache management.
  • Compute: CUDA-accelerated kernels for Attention, RoPE, and GEMM (cuBLAS).
  • Parity: Exact token matching with HuggingFace Transformers (verified).
  • Interface: Clean Python API via PyBind11.

Current Limitations (v0.1)

  • Single‑GPU only
  • Limited model family support (LLaMA‑style)
  • CPU backend not optimized (minor numerical divergence vs CUDA)
  • No quantization support

These constraints are intentional for the initial release.

What to expect in future

Here is what I have planned to add to the project in later iterations. Send a PR if you want to contribute to the project.

  • Multi‑GPU & Distributed Execution – Enable scaling across multiple GPUs and clusters for larger models.
  • Expanded Model Support – Add native support for Mistral, Falcon, and other non‑LLaMA families.
  • Optimized CPU Backend – Introduce a high‑performance CPU path (vectorized kernels, OpenMP) and enable CPU‑only inference.
  • Quantization & INT8 – Provide post‑training quantization pipelines and INT8 kernels for reduced memory and faster inference.
  • FlashAttention‑style Kernels – Integrate memory‑efficient, block‑sparse attention kernels to cut latency and improve throughput.
  • Plugin System – Allow community‑contributed extensions (custom ops, alternative KV‑cache strategies).
  • Better Tooling – CLI utilities for model conversion, benchmarking, and profiling.

Installation

Prerequisites

  • NVIDIA GPU with CUDA 11/12 (Recommended)
  • C++ Compiler (GCC 10+ or Clang 12+)
  • CMake 3.18+
  • Python 3.8+

Install from Source

# Clone repository
git clone https://github.com/cottus-ai/cottus-runtime.git
cd cottus-runtime

# Create virtual environment (Recommended)
python3 -m venv .venv
source .venv/bin/activate

# Install in editable mode
pip install -e .

Quick Start

1. Basic Inference (Tiny Random Model)

python examples/1_basic_inference.py --device cuda

2. CPU Fallback

No GPU? No problem.

python examples/2_cpu_inference.py

3. Real Chat (TinyLlama-1.1B)

Requires ~2.2GB download.

python examples/3_tinyllama_real.py

Usage

The best way to get started is to look at the examples/ directory, which contains complete scripts for various use cases.

Basic Example

from cottus import Engine, EngineConfig
from cottus.model import load_hf_model

# 1. Load Model Weights
weights, _, _, tokenizer, _ = load_hf_model("TinyLlama/TinyLlama-1.1B-Chat-v1.0", device="cuda")

# 2. Config
config = EngineConfig()
config.model_type = "llama"
config.hidden_dim = 2048
config.num_layers = 22
config.num_heads = 32
config.num_kv_heads = 4
config.head_dim = 64
config.intermediate_dim = 5632
config.device = "cuda"

# 3. Helpers
engine = Engine(config, weights)
input_ids = tokenizer.encode("Hello!")

# 4. Generate
output_ids = engine.generate(input_ids, max_new_tokens=20)
print(tokenizer.decode(output_ids))

License

Cottus Runtime is licensed under the Apache License Version 2.0. By contributing to the project, you agree to the license and copyright terms therein and release your contribution under these terms.

Metadata

Release files for cottus 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for cottus 0.1.0
File Size Uploaded
cottus-0.1.0.tar.gz 553.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for cottus 0.1.0
File Interpreter ABI Platform
cottus-0.1.0-cp312-cp312-manylinux_2_34_x86_64.whl CPython 3.12 CPython 3.12 Linux glibc 2.34+ x86-64 Details

Total release size: 924.8 kB

Release files / cottus-0.1.0.tar.gz

Download URL cottus-0.1.0.tar.gz
Size 553.7 kB
Tags Source
SHA-256 checksum
How to use checksums
1ecd785197eca1d809789424e514033bac186c1a6f7e59e766f815bf76c16c8a
BLAKE2b-256 checksum
How to use checksums
b4ffb0d7d6e7c972db41a0fd4eca4c717b54a3e7623c9373ac0aac99440906fe
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jan 4, 2026.

Transparency log

Release files / cottus-0.1.0-cp312-cp312-manylinux_2_34_x86_64.whl

Download URL cottus-0.1.0-cp312-cp312-manylinux_2_34_x86_64.whl
Size 371.1 kB
Tags CPython 3.12 Linux glibc 2.34+ x86-64
SHA-256 checksum
How to use checksums
ead2c92c114a2655549bb423cbea7d94ef304f614a439e9dfa0dbfcca0f3230a
BLAKE2b-256 checksum
How to use checksums
b46602791cfa2e45b5c745de14cd249ee518f3675d1416e61d797d18e16db70d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jan 4, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page