LLMRig
Evidence-driven control plane for local AI.
Detect runtimes. Resolve artifacts. Decide from evidence. Run locally. Verify with measurement.
LLMRig answers a deceptively hard local-AI question: what can this machine actually run, through which runtime, with what evidence?
It keeps hardware, model identity, artifact format, quantization, runtime capability, local availability, execution state, context, and measured performance separate so missing evidence never silently becomes a confident claim.
unknown != false
compatible != local
local != executable
executable != measurable
measurable != measured
measured performance != model quality
What v0.8 adds
LLMRig v0.8 turns the v0.7 Autopilot foundation into a runtime-aware local-AI control plane foundation:
- a common runtime-adapter registry for oMLX, Ollama, MLX-LM, and llama.cpp
llmrig runtimesfor read-only runtime readiness and capability inspection- Hugging Face-first artifact resolution with conservative format, quantization, context, and provenance rules
- first-class oMLX inventory/provenance and explicit
solve --verifyexecution - authenticated local oMLX support through
OMLX_API_KEYwithout exposing secrets - cross-runtime
race-v2measurement that leaves unavailable metrics unknown - no arbitrary filesystem scans and no surprise model/runtime installation
The longer-term Autopilot direction is Detect → Decide → Configure → Run → Verify. v0.8 deliberately keeps configuration and acquisition actions read-only/manual; mutating plan/apply actions are the next roadmap stage.
Quick start
Install the isolated CLI:
pipx install llmrig
llmrig --version
Inspect the local runtime surface:
llmrig runtimes
Then analyze a curated model or an exact Hugging Face repository:
llmrig solve qwen3:0.6b
llmrig solve mlx-community/Qwen3.5-27B-4bit
Nothing is downloaded or executed by default.
Runtime intelligence
llmrig runtimes
llmrig runtimes --json
LLMRig reports runtime installation/readiness separately from model locality and execution support. The v0.8 registry currently covers:
| Runtime | Detection | LLMRig execution / measurement | Primary artifact evidence |
|---|---|---|---|
| oMLX | Yes | Yes, for provenance-backed local models | MLX |
| Ollama | Yes | Yes | Ollama-managed artifacts |
| MLX-LM | Yes | Yes, for explicit local artifacts | MLX |
| llama.cpp | Yes | Yes, for explicit local artifacts | GGUF |
Runtime detection is read-only. LLMRig does not install these runtimes or start services while probing them.
Solve first
llmrig solve MODEL
llmrig solve MODEL --json
llmrig solve MODEL --context 32768
solve constructs independent evidence dimensions for each candidate:
discovery
compatibility
runtime availability
local availability
execution
measurement capability
measurement
recommendation
A candidate can therefore be compatible but not local, local but not executable, or measurable but not measured. An inconclusive result is a valid result.
LLMRig does not calculate a universal score or infer model quality from throughput. Planning recommendations only appear when the available evidence supports them.
Explicit local artifacts
LLMRig never scans arbitrary directories for native model files. Supply a locator explicitly when you want a native local artifact considered:
llmrig solve MODEL --local-artifact llama.cpp=/path/to/model.gguf
llmrig solve MODEL --local-artifact mlx-lm=/path/to/model-directory
A user-supplied path establishes a local association, not independent content attestation. LLMRig checks bounded structure, keeps the private locator out of public solve output, and leaves unknown identity/quantization facts unknown.
Hugging Face resolution
Exact owner/repository identifiers are resolved through read-only Hugging Face
metadata. LLMRig may read the repository's small config.json when necessary for
structured format, quantization, or context evidence; it does not download model
weights during resolution.
Important evidence rules include:
.ggufestablishes GGUF packaging; filename quantization is accepted only when exactly one recognized token is present- generic
.safetensorsdoes not by itself establish MLX/oMLX compatibility - an
mlxtag/path hint alone is not proof of MLX packaging - MLX packaging requires stronger structured evidence, such as explicit
library_name=mlxor an MLX hint backed by the MLX-LM quantization contract - conflicting context, quantization, base-model provenance, or incomplete shard groupings fail closed instead of being guessed through
Discovery metadata is evidence, not installation trust.
oMLX in v0.8
LLMRig can observe an API-visible oMLX model and associate it with an exact Hugging Face repository only when provenance is defensible:
- oMLX directly reports the exact source repository; or
- the completed oMLX Hugging Face download registry reports the exact repository and the API-visible local model mapping is unique and unambiguous.
Display names alone are not accepted as provenance.
Authenticated local endpoints are supported through:
export OMLX_API_KEY="..."
llmrig runtimes
The API key, admin session cookie, filesystem paths, and private execution locator are never serialized into public solve output.
Verify explicitly
Default solve is read-only. Measurement requires explicit intent:
llmrig solve MODEL \
--local-artifact mlx-lm=/path/to/model-directory \
--verify
solve --verify requires at least two comparable, already-local, executable,
measurable candidates and reuses the deterministic race-v2 workload.
For oMLX, LLMRig prefers server-reported prompt/generation timing and throughput.
If the server reports only total_time, LLMRig records latency only and leaves
throughput unmeasured. It does not synthesize tokens/second from incomparable data.
Balanced decisions remain inconclusive when too few comparable measured dimensions remain or multiple Pareto tradeoffs survive.
Solve exit codes:
0— analysis completed, including completed-but-inconclusive verification1— operational solve or attempted verification failure2— invalid/unresolvable input or unavailable requested verification
Minimal Python SDK
import llmrig
result = llmrig.solve("mlx-community/Qwen3.5-27B-4bit")
print(result.plan.recommendation_status)
Stable entry point:
llmrig.solve(
model,
*,
context=None,
local_artifacts=(),
verify=False,
) -> llmrig.SolveResult
The SDK prints nothing and never exits the process. Invalid/unresolvable input raises
SolveInputError; operational failures raise SolveEngineError. SolveResult and
SolveCandidate are deliberate public contracts. _llmrig is private implementation.
Measurement commands
| Command | Purpose | Decision semantics |
|---|---|---|
solve --verify |
Verify a solve candidate set | balanced Pareto or inconclusive |
race |
Compare local configurations | metric-specific generation/prompt/latency results |
choose |
Explain one measured objective | generation, prompt, latency, or balanced |
optimize |
Expose measured tradeoffs | unranked noise-aware Pareto frontier |
bench |
Full Ollama benchmark | throughput, residency when available, smoke tests |
Example native race:
llmrig race MODEL \
--local-artifact llama.cpp=/path/to/model.gguf \
--local-artifact mlx-lm=/path/to/model-directory
Results within the configured 5% race threshold are inconclusive. Performance measurement does not establish model quality.
Installation
pipx is recommended:
pipx install llmrig
pipx upgrade llmrig
Or use a virtual environment:
python3 -m venv .venv
source .venv/bin/activate # macOS/Linux
python -m pip install llmrig
Windows PowerShell:
python -m venv .venv
.venv\Scripts\Activate.ps1
python -m pip install llmrig
LLMRig supports Python 3.9+ on macOS, Linux, and Windows and has no third-party
Python runtime dependencies. A source checkout can also run python3 llmrig.py.
Architecture direction
LLMRig
┌───────────┐
│ RigGraph │
└─────┬─────┘
│
┌──────────┴──────────┐
│ Autopilot Engine │
└──────────┬──────────┘
│
┌──────────┬───────┼───────┬──────────┐
↓ ↓ ↓ ↓ ↓
oMLX Ollama MLX-LM llama.cpp future
v0.8 establishes the adapter/evidence foundation. v0.9 is planned to add explicit
plan/apply actions for acquisition, configuration, and launch. See
ROADMAP_V1.md.
Project docs
CHANGELOG.md— canonical release historyROADMAP_V1.md— product directionRELEASING.md— GitHub Release and PyPI processSECURITY.md— security policy and trust boundariesCONTRIBUTING.md— contribution guide
License
MIT
Release files for llmrig 0.8.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llmrig-0.8.1.tar.gz | 142.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llmrig-0.8.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 246.5 kB
Release files / llmrig-0.8.1.tar.gz
| Download URL | llmrig-0.8.1.tar.gz |
|---|---|
| Size | 142.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c56e9d394e319cfa5d8cf6eb931676178a98f92c6d284be9bf7e068019c4f15a
|
|
BLAKE2b-256 checksum How to use checksums |
b76a348e6f37b71d4caab8e05340578ec79c1921df89913733334734eb709ca9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 24, 2026.
Transparency logRelease files / llmrig-0.8.1-py3-none-any.whl
| Download URL | llmrig-0.8.1-py3-none-any.whl |
|---|---|
| Size | 103.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
51dbdf8d3a7141e791ce1d03f115703794e6d2c588088a8e99b0c0daaef7ffec
|
|
BLAKE2b-256 checksum How to use checksums |
138441b3d570d625c13297d623e3dfd0e19ebf4da2f857cf6480f2af5763e9b1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 24, 2026.
Transparency log