Skip to main content

inferris

A trace-native, memory-safe inference engine for on-device AI — built so you can see what a model is doing while it runs, in real time, and act on it.

Why

On-device inference runtimes are black boxes: weights go in, tokens come out, and everything in between is invisible — which is exactly where the interesting failures live, and where guardrails will have to live. inferris is built engine-first around the opposite premise: the runtime itself should be observable while it runs. Every instrument that was used to build and debug this engine ships in the engine. Real-time inference insight is the product today; a guardrail surface grows on top of it as the engine grows.

Install

pip install inferris

or run it without installing anything:

uvx inferris chat models/Qwen3.8-27B-Q4_K_M.gguf

Quickstart

# grab a model (Qwen3.8 family GGUF)
uvx --from huggingface_hub hf download unsloth/Qwen3.8-27B-GGUF \
  --include "*Q4_K_M*.gguf" --local-dir models/

# one-shot generation
inferris generate models/Qwen3.8-27B-Q4_K_M.gguf "The capital of France is"

# interactive chat (REPL, chat template read from GGUF metadata)
inferris chat models/Qwen3.8-27B-Q4_K_M.gguf

# NVIDIA machines: opt in to the fast path (auto-selection is on the short list)
INFERRIS_BACKEND=cuda INFERRIS_MTP=1 inferris chat models/Qwen3.8-27B-Q4_K_M.gguf

Sampling defaults resolve CLI flag → GGUF header (general.sampling.*) → model-card fallback. --temp 0 is exact greedy; --seed makes sampling reproducible. Every run reports a # perf: prefill/decode split on stderr.

What it is today — deliberately narrow

  • One model family: Qwen3.8 (hybrid Gated DeltaNet + gated attention), with its Qwen3.5/3.6 siblings as verification rungs. GDN — the recurrent half of the hybrid — has first-class support and is where the throughput work went.
  • One optimization target: decode throughput on a single NVIDIA RTX 5090, via hand-written CUDA kernels and the family's built-in MTP speculative decoding (byte-identical to non-speculative output, enforced by a gate on every commit).
  • A memory-safe core: the engine is pure Rust from GGUF parsing to sampling; unsafe is confined to the CUDA kernel boundary.

That narrowness is a design position, not an early-days apology: depth of verification and observability on one family first, breadth second.

Performance

Decode throughput, Qwen3.8-27B Q4_K_M, single RTX 5090 (32 GB, sm_120), identical GGUF bytes, greedy, hot cache, median of three runs:

engine plain decode speculative decode (MTP)
llama.cpp (CUDA, build b10217) 78.3 tok/s 117.8 tok/s
inferris 76.5 tok/s 113.5 tok/s

Honest ledger: prefill is not yet optimized (decode-first engine; prefill is the next campaign), and the remaining decode gap vs llama.cpp is tracked openly. Numbers are as of this release and will move.

Trace-native

Observability is the product surface, not a debug flag:

  • Loud attestations on stderr for every mode-class decision: backend, weights mode, tied vs untied logits head, MTP layer detection — "probably took the fast path" is never a thing.
  • INFERRIS_TRACE — per-layer norm probes (|mix| / |ffn| / |x|) to locate where a forward pass goes wrong.
  • INFERRIS_LENS — a logit lens over chosen token ids at every layer.
  • Speculative acceptance counters (# spec: accepted a/b) on every run.

Verification

  • Same-bytes differential testing against llama.cpp on the identical GGUF file — judge backend pinned and recorded, margin-aware verdicts (near-ties at low bit-width abstain instead of failing).
  • A three-judge protocol (bf16 reference / same-bytes peer / engine) that separates quantization noise from real bugs.
  • 90 unit and property tests, bit-exact anchors on the quant codecs, and the speculative-decoding byte-identity gate.

Model & platform support

model status
Qwen3.8-27B (Q4_K_M) CUDA fast path, benchmarked above
Qwen3.5-0.8B (Q8_0 / Q4_K_M) verification rung — all gates green

Quantizations: Q4_K, Q5_K, Q6_K, Q8_0, F16, F32. Text path only (Qwen3.8's vision tower ships as a separate mmproj file and is not loaded).

platform wheel path
Linux x86_64 + NVIDIA GPU ✅ CUDA fast path (hand-written kernels, tuned on sm_120)
macOS arm64 ✅ CPU reference path (correctness-grade)

Roadmap

  • More models and more hardware — with a deliberate bias toward the software/hardware architectures designed to put more intelligence on small hardware: linear/hybrid-attention families like GDN, and the on-device runtimes and silicon they are meant for.
  • The guardrail layer: constrained decoding and intervention hooks, built on the trace surface that already ships.
  • Progressive open-sourcing across the 0.x series. PyPI wheels are the supported distribution today.

Python API

from inferris import Engine

eng = Engine("models/Qwen3.5-0.8B-Q8_0.gguf")
text = eng.generate("The capital of France is", max_tokens=16)  # greedy by default
ids = eng.encode("hi 🦀")
raw = eng.decode(ids)  # returns bytes — UTF-8 only re-establishes after concatenation

The API defaults to greedy (reproducibility-first); the CLI defaults to the model card's recommended sampling. Type stubs ship in the wheel.

License

Dual-licensed under MIT OR Apache-2.0.

Release files for inferris 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distribution (wheel)

Table of built distributions (wheels) for inferris 0.1.0
File Interpreter ABI Platform
inferris-0.1.0-cp39-abi3-manylinux_2_34_x86_64.whl CPython 3.9 abi3 Linux glibc 2.34+ x86-64 Details

Release files / inferris-0.1.0-cp39-abi3-manylinux_2_34_x86_64.whl

Download URL inferris-0.1.0-cp39-abi3-manylinux_2_34_x86_64.whl
Size 931.0 kB
Tags CPython 3.9 Linux glibc 2.34+ x86-64 abi3
SHA-256 checksum
How to use checksums
2b336a9673c42d39efc4bfe3d45cd345fc8c8f4d98266069ad5d0db08bdea656
BLAKE2b-256 checksum
How to use checksums
98f50508ebb78ce0c5304de221dab0d1025ae773d9177c0b3cdfcd70383dca78
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.14.1

Release history Release notifications | RSS feed

This release

0.1.0 This release

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page