inferris
A trace-native, memory-safe inference engine for on-device AI — built so you can see what a model is doing while it runs, in real time, and act on it.
Why
On-device inference runtimes are black boxes: weights go in, tokens come out, and everything in between is invisible — which is exactly where the interesting failures live, and where guardrails will have to live. inferris is built engine-first around the opposite premise: the runtime itself should be observable while it runs. Every instrument that was used to build and debug this engine ships in the engine. Real-time inference insight is the product today; a guardrail surface grows on top of it as the engine grows.
Install
pip install inferris
or run it without installing anything:
uvx inferris chat models/Qwen3.8-27B-Q4_K_M.gguf
Quickstart
# grab a model (Qwen3.8 family GGUF)
uvx --from huggingface_hub hf download unsloth/Qwen3.8-27B-GGUF \
--include "*Q4_K_M*.gguf" --local-dir models/
# one-shot generation
inferris generate models/Qwen3.8-27B-Q4_K_M.gguf "The capital of France is"
# interactive chat (REPL, chat template read from GGUF metadata)
inferris chat models/Qwen3.8-27B-Q4_K_M.gguf
# NVIDIA machines: opt in to the fast path (auto-selection is on the short list)
INFERRIS_BACKEND=cuda INFERRIS_MTP=1 inferris chat models/Qwen3.8-27B-Q4_K_M.gguf
Sampling defaults resolve CLI flag → GGUF header (general.sampling.*) →
model-card fallback. --temp 0 is exact greedy; --seed makes sampling
reproducible. Every run reports a # perf: prefill/decode split on stderr.
What it is today — deliberately narrow
- One model family: Qwen3.8 (hybrid Gated DeltaNet + gated attention), with its Qwen3.5/3.6 siblings as verification rungs. GDN — the recurrent half of the hybrid — has first-class support and is where the throughput work went.
- One optimization target: decode throughput on a single NVIDIA RTX 5090, via hand-written CUDA kernels and the family's built-in MTP speculative decoding (byte-identical to non-speculative output, enforced by a gate on every commit).
- A memory-safe core: the engine is pure Rust from GGUF parsing to
sampling;
unsafeis confined to the CUDA kernel boundary.
That narrowness is a design position, not an early-days apology: depth of verification and observability on one family first, breadth second.
Performance
Decode throughput, Qwen3.8-27B Q4_K_M, single RTX 5090 (32 GB, sm_120), identical GGUF bytes, greedy, hot cache, median of three runs:
| engine | plain decode | speculative decode (MTP) |
|---|---|---|
| llama.cpp (CUDA, build b10217) | 78.3 tok/s | 117.8 tok/s |
| inferris | 76.5 tok/s | 113.5 tok/s |
Honest ledger: prefill is not yet optimized (decode-first engine; prefill is the next campaign), and the remaining decode gap vs llama.cpp is tracked openly. Numbers are as of this release and will move.
Trace-native
Observability is the product surface, not a debug flag:
- Loud attestations on stderr for every mode-class decision: backend, weights mode, tied vs untied logits head, MTP layer detection — "probably took the fast path" is never a thing.
INFERRIS_TRACE— per-layer norm probes (|mix| / |ffn| / |x|) to locate where a forward pass goes wrong.INFERRIS_LENS— a logit lens over chosen token ids at every layer.- Speculative acceptance counters (
# spec: accepted a/b) on every run.
Verification
- Same-bytes differential testing against llama.cpp on the identical GGUF file — judge backend pinned and recorded, margin-aware verdicts (near-ties at low bit-width abstain instead of failing).
- A three-judge protocol (bf16 reference / same-bytes peer / engine) that separates quantization noise from real bugs.
- 90 unit and property tests, bit-exact anchors on the quant codecs, and the speculative-decoding byte-identity gate.
Model & platform support
| model | status |
|---|---|
| Qwen3.8-27B (Q4_K_M) | CUDA fast path, benchmarked above |
| Qwen3.5-0.8B (Q8_0 / Q4_K_M) | verification rung — all gates green |
Quantizations: Q4_K, Q5_K, Q6_K, Q8_0, F16, F32. Text path only (Qwen3.8's vision tower ships as a separate mmproj file and is not loaded).
| platform | wheel | path |
|---|---|---|
| Linux x86_64 + NVIDIA GPU | ✅ | CUDA fast path (hand-written kernels, tuned on sm_120) |
| macOS arm64 | ✅ | CPU reference path (correctness-grade) |
Roadmap
- More models and more hardware — with a deliberate bias toward the software/hardware architectures designed to put more intelligence on small hardware: linear/hybrid-attention families like GDN, and the on-device runtimes and silicon they are meant for.
- The guardrail layer: constrained decoding and intervention hooks, built on the trace surface that already ships.
- Progressive open-sourcing across the 0.x series. PyPI wheels are the supported distribution today.
Python API
from inferris import Engine
eng = Engine("models/Qwen3.5-0.8B-Q8_0.gguf")
text = eng.generate("The capital of France is", max_tokens=16) # greedy by default
ids = eng.encode("hi 🦀")
raw = eng.decode(ids) # returns bytes — UTF-8 only re-establishes after concatenation
The API defaults to greedy (reproducibility-first); the CLI defaults to the model card's recommended sampling. Type stubs ship in the wheel.
License
Dual-licensed under MIT OR Apache-2.0.
Release files for inferris 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| inferris-0.1.0-cp39-abi3-manylinux_2_34_x86_64.whl | CPython 3.9 | abi3 | Linux glibc 2.34+ x86-64 | Details |
Release files / inferris-0.1.0-cp39-abi3-manylinux_2_34_x86_64.whl
| Download URL | inferris-0.1.0-cp39-abi3-manylinux_2_34_x86_64.whl |
|---|---|
| Size | 931.0 kB |
| Tags | CPython 3.9 Linux glibc 2.34+ x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
2b336a9673c42d39efc4bfe3d45cd345fc8c8f4d98266069ad5d0db08bdea656
|
|
BLAKE2b-256 checksum How to use checksums |
98f50508ebb78ce0c5304de221dab0d1025ae773d9177c0b3cdfcd70383dca78
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
maturin/1.14.1
|