Hardware-aware local LLM deployment: probe, resolve, run.
Project description
Rigma
Hardware-aware local LLM deployment for consumer machines: rigma up probes your GPU/RAM,
picks the community-verified best model + quant + flag combo for your exact hardware,
downloads a pinned llama.cpp build and the model, and serves an OpenAI-compatible endpoint —
no knob-mashing required.
Unlike generic runners, Rigma applies the tuning that actually matters per machine:
MoE expert offload (--n-cpu-moe) sized to your RAM, architecture-aware KV-cache policies
(e.g. q8_0 K-cache floor on DeltaNet-family models), backend selection per GPU generation
(e.g. Vulkan over ROCm on RDNA4), flash attention, and session persistence
(--slot-save-path) on by default. Every decision is auditable: rigma plan --explain
shows the arithmetic and sources.
Quickstart (pre-alpha)
pip install rigma
rigma up # probes your machine, downloads the best model, opens the chat UI
That's it — a browser tab opens with a chat connected to your tuned local model, and any
OpenAI-compatible tool can use http://127.0.0.1:11500/v1.
Commands
| Command | What it does |
|---|---|
rigma up |
Start everything; opens the chat UI in your browser |
rigma chat |
Chat with the running model in the terminal |
rigma status |
What's running, where |
rigma stop |
Stop the model server and UI |
rigma models |
What fits your machine |
rigma plan --explain |
What up would run, with the math |
rigma doctor |
What Rigma detects on this machine |
rigma update |
Pull the latest community combo registry |
rigma up flags: --use-case coding · --model SLUG · --port 11500 · --no-browser ·
--turbo (fast download, may hog your bandwidth) · --yes · --dry-run
Status
Pre-alpha (M2). Combos come in two grades: verified (benchmarked on real hardware, evidence attached) and provisional (research-seeded fit math — run one and PR your numbers to rigma-registry). Verified so far:
| Hardware | Model | Backend | Result |
|---|---|---|---|
| RX 9070 XT 16GB + 16GB RAM (Windows) | Qwen3.6-35B-A3B UD-Q3_K_XL, ctx 32K, n_cpu_moe 10 | Vulkan (llama.cpp b9867) | verified 2026-07-06: 57.1 t/s gen, 689 t/s prefill @ 4K prompt |
Design: docs/superpowers/specs/2026-07-03-rigma-design.md. License: Apache-2.0.
RAG integration (via Raggity, AGPL-3.0, separate process) lands in M4.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file rigma-0.3.0.tar.gz.
File metadata
- Download URL: rigma-0.3.0.tar.gz
- Upload date:
- Size: 74.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
96d402922e474b841bccf271d311507889e8424a2bb8ee7659e8427ec4a2ba49
|
|
| MD5 |
48453263b081dacf3eee7a83e7741149
|
|
| BLAKE2b-256 |
5045748328d065aa5e70574a67d5b57dc9e7508af6fc6ba6870f12849ccb461d
|
Provenance
The following attestation bundles were made for rigma-0.3.0.tar.gz:
Publisher:
publish.yml on IxMxAMAR/rigma
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
rigma-0.3.0.tar.gz -
Subject digest:
96d402922e474b841bccf271d311507889e8424a2bb8ee7659e8427ec4a2ba49 - Sigstore transparency entry: 2116012764
- Sigstore integration time:
-
Permalink:
IxMxAMAR/rigma@fd6de1776ce77fbab291b1d553f30c044a8d12d2 -
Branch / Tag:
refs/tags/v0.3.0 - Owner: https://github.com/IxMxAMAR
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@fd6de1776ce77fbab291b1d553f30c044a8d12d2 -
Trigger Event:
release
-
Statement type:
File details
Details for the file rigma-0.3.0-py3-none-any.whl.
File metadata
- Download URL: rigma-0.3.0-py3-none-any.whl
- Upload date:
- Size: 26.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
fbe2702f1cfd2a8cc7fddf2df3cdb6d85921e0648447def3c5015346de5e45f8
|
|
| MD5 |
151edb3da8bc862a0dd68c55e5ab90ca
|
|
| BLAKE2b-256 |
cf1109d972f5379d1bcda70711c2497691db4c07d67380106ffb83db0dd1ec37
|
Provenance
The following attestation bundles were made for rigma-0.3.0-py3-none-any.whl:
Publisher:
publish.yml on IxMxAMAR/rigma
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
rigma-0.3.0-py3-none-any.whl -
Subject digest:
fbe2702f1cfd2a8cc7fddf2df3cdb6d85921e0648447def3c5015346de5e45f8 - Sigstore transparency entry: 2116012842
- Sigstore integration time:
-
Permalink:
IxMxAMAR/rigma@fd6de1776ce77fbab291b1d553f30c044a8d12d2 -
Branch / Tag:
refs/tags/v0.3.0 - Owner: https://github.com/IxMxAMAR
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@fd6de1776ce77fbab291b1d553f30c044a8d12d2 -
Trigger Event:
release
-
Statement type: