Skip to main content

ThinTensor

thintensor is a unified, high-performance command-line engine for pulling, converting, running, benchmarking, and validating causal language models. It compiles a high-speed Rust-based archive core with an optimized PyTorch/Triton GPU execution runtime.


Fast-Path: Install the shipped release

ThinTensor ships as two packages: the Python/Triton runtime on PyPI and the native archive/conversion core on crates.io. A normal user installs both from the registries; cloning the repository and building the Rust binary is not required.

1. Install prerequisites

Ensure the host has:

  • Python 3.10 or newer
  • Rust and Cargo: install via rustup.rs if missing
  • NVIDIA CUDA Toolkit for GPU execution (ensure nvcc is available)

2. Set up a virtual environment and install PyTorch

Create a fresh python environment and install PyTorch with CUDA support:

python3 -m venv venv
source venv/bin/activate
pip install --upgrade pip

# Install PyTorch with CUDA (matching your system's CUDA version)
pip install torch --index-url https://download.pytorch.org/whl/cu121

3. Install the published Python package and Rust core

Install the Python CLI from PyPI and the native thintensor-core binary from crates.io:

python -m pip install 'thintensor[all]'
cargo install thintensor --locked

cargo install thintensor installs the thintensor-core executable into Cargo's binary directory, normally ~/.cargo/bin. Make sure that directory is on PATH:

command -v thintensor-core
thintensor-core --help

The Python CLI discovers that Cargo-installed binary automatically. For a source checkout or a custom installation, set THINTENSOR_CORE_BIN to the binary path instead.

4. Verify the Installation

Run the doctor command to ensure the GPU runtime and kernel dependencies are fully operational:

thintensor doctor --strict

Developing from source

If you are contributing to ThinTensor rather than using the shipped release, the source build remains available:

git clone https://github.com/random-unknown-username/Thintensor.git
cd Thintensor
python -m pip install -e '.[all]'
cargo build --locked --release --bin thintensor-core

Quickstart: Downloading & Running a Sample Model (Qwen-0.8B)

Follow this fast-path to pull, convert, and execute a lightweight model (Qwen-0.8B):

  1. Download the Hugging Face weights:

    thintensor pull Qwen/Qwen3.5-0.8B
    
  2. Convert the weights into a .thin archive:

    thintensor convert ~/.cache/thintensor/models/Qwen--Qwen3.5-0.8B --out ~/.cache/thintensor/models/Qwen--Qwen3.5-0.8B/Qwen3.5-0.8B.thin
    
  3. Run a prompt through the native GPU runtime:

    thintensor run ~/.cache/thintensor/models/Qwen--Qwen3.5-0.8B/Qwen3.5-0.8B.thin --prompt "Explain quantum computing in one sentence."
    
  4. Verify correctness similarity metrics:

    thintensor validate Qwen3.5-0.8B.thin --hf-model ~/.cache/thintensor/models/Qwen--Qwen3.5-0.8B --profile max-max-perf --suite quick
    

[!NOTE] Base Model vs. Chat Model Behavior: The sample Qwen3.5-0.8B is a raw base model trained only for next-token document completion. It does not engage in interactive conversation.

  • Leading Punctuation: It completes prompts naturally (e.g. Hello -> , I am working with... or What is gravity -> , and how does it affect...).
  • Greedy Decoding Only: To maximize speed and compile highly optimized fused Triton argmax kernels, the runtime is strictly greedy-only (temperature=0, top_p=1, top_k=0).
  • Avoiding Loops: Because there is no stochastic sampling to escape repetition loops, tiny base models (0.8B) may repeat sentences under greedy decoding. To avoid loops and get proper interactive chat responses, always use fine-tuned instruct models (e.g., Qwen/Qwen2.5-3B-Instruct). Bassically right now, we cannot change model parameters like temperature, top_p and top_k, I would try my best to get this fixed in future

Code Architecture & Core Modules

The engine is split into a Rust Archive & Conversion Core and a Python/Triton GPU Execution Runtime. Below is a detailed map of the codebase architecture:

graph TD
    CLI[cli.py: User Commands] --> |Load Archive| Archive[archive.py / archive.rs]
    CLI --> |Deduce Fit/Streaming| Plan[plan.rs: Budget Planning]
    CLI --> |Execute Runtime| Runtime[gpu_runtime.py: ThinGpuCausalLMRuntime]
    Runtime --> |Fused Math| Triton[triton_kernels.py: Fused Kernels]
    Runtime --> |Zero-Copy Views| Archive
    Runtime --> |Causal GQA/MHA| SDPA[PyTorch C++ SDPA Kernel]

1. CLI Entrypoint & Routing

  • CLI Handler: thinruntime/cli.py manages subcommands like run, pull, convert, bench, and validate.
  • Routing Engine: The CLI inspects model metadata via auto_fit.py and routes to the native high-performance runtime for compatible models, falling back to Hugging Face transformers for incompatible architectures.

2. Rust Core (Archive & Conversion)

The Rust modules under src/ handle disk-to-memory layouts and weight packing:

  • Archive Reader/Writer: src/archive.rs and src/manifest.rs define the binary format of .thin packages.
  • Model Converter: src/convert_hf.rs parses Hugging Face safetensors, mapping weights and transforming shapes into contiguous memory layouts.
  • VRAM Budget & Fit Planner: src/plan.rs inspects available VRAM and maps which weight layers must be streamed or pinned to VRAM.

3. High-Performance GPU Runtime

The Python runtime classes coordinate host-device memory mapping and layer execution:

  • ThinGpuCausalLMRuntime: thinruntime/gpu_runtime.py#L2490 is the execution engine.
    • Memory-Mapped Zero-Copy Views: thinruntime/archive.py exposes binary pages as PyTorch tensor views directly from mmap.
    • Forward Causal Decode: forward_token (line 4828+) coordinates prefetch pipelines and sequential layer dispatch.
    • Gated Mixture-of-Experts (MoE): _forward_token_moe (line 6095+) runs MoE routing. When weights exceed VRAM, experts are streamed dynamically using page-pool overlays (line 6260+).
    • Optimized Eager RoPE: _apply_rope (line 5804+) applies rotary positional embeddings. Eager position embeddings are applied in-place to avoid allocations while matching Hugging Face precision perfectly.
    • Scaled Dot-Product Attention: _attention (line 5948+) leverages PyTorch's native C++ scaled_dot_product_attention for fast GQA/MHA execution.

4. Triton Custom Kernels

  • Fused Normalization: thinruntime/triton_kernels.py implements fused add_rms_norm and SwiGLU operations to bypass PyTorch intermediate launch overheads.

Profiles & Configuration

Profiles are intent-based presets defined in thinruntime/profile_presets.py:

  • safe: Preserves full precision (BF16) weights and KV history with the broadest compatibility.
  • balanced: Enables native matvec and GQA/MHA attention kernels without weight compression.
  • max-performance: Opt-in profile targeting INT8 body and tensor-core quantization.
  • max-max-perf: The most aggressive execution preset combining fused projection caches, fast C++ SDPA, and custom in-place memory optimizations.

Verification & Correctness Testing

Logit parity is validated step-by-step against Hugging Face references:

thintensor validate google--gemma-4-E2B.thin \
  --hf-model /path/to/original-gemma-4 \
  --profile max-max-perf \
  --suite quick

Validation execution isolates processes: it runs the Hugging Face trajectory first, caches reference logits, unloads it from VRAM, and then loads the .thin model to calculate the exact minimum cosine similarity across all tokens.

Benchmark results

image

N/A because the models didnt load my 8gb vram gpu without thintensors!

Release files for thintensor 1.0.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for thintensor 1.0.2
File Size Uploaded
thintensor-1.0.2.tar.gz 263.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for thintensor 1.0.2
File Interpreter ABI Platform
thintensor-1.0.2-py3-none-any.whl Python 3 none any Details

Total release size: 537.9 kB

Release files / thintensor-1.0.2.tar.gz

Download URL thintensor-1.0.2.tar.gz
Size 263.3 kB
Tags Source
SHA-256 checksum
How to use checksums
a3352ac733e3b303978adf9ca75802134b85f2c40a7de5b1cefdb4877d0df9bf
BLAKE2b-256 checksum
How to use checksums
ad7d16fd9232bf5e8198b3b0f8b9efe3c6714323588fbab5092437966caaa435
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.6

Release files / thintensor-1.0.2-py3-none-any.whl

Download URL thintensor-1.0.2-py3-none-any.whl
Size 274.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
99765a935e4c48e719b5f49aab0cd94e514c98a990bdc5f9a07ca9fc1f6a01fb
BLAKE2b-256 checksum
How to use checksums
fbd1ce50a588c4c1802ff2cd45d07abb3ed5deab7ef8b6939dfb37a06f9abc89
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.6

Release history Release notifications | RSS feed

This release

1.0.2 This release

2 release files

1.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page