Skip to main content

Zenthrix

Hardware-Adaptive Edge Neural Graph Compiler

PyPI License CI

Zenthrix is an edge-native model compiler frontend for compiling open-weight neural networks (LLMs, SLMs, and vision models) into zero-copy, memory-optimized binaries tailored for consumer edge silicon — Apple Silicon, Qualcomm Snapdragon NPU, and Arm Cortex/Ethos.

This repository is the public developer entry point: the PyPI package, CLI, and model-ingestion layer. The proprietary compilation engine itself lives in a separate private repository and is distributed as a precompiled binary. The v0.1.0 frontend validates inputs and exposes the integration boundary; compilation and inference require that separately provisioned engine.


Table of Contents


Key Features

  • Direct Ingestion — Native loaders for ONNX, PyTorch Export (AOTInductor), and GGUF architectures, with no intermediate format conversion required.
  • Unified Memory Tiling — Schedules compute passes against unified memory architectures, reducing peak active RAM allocation by up to 40%.
  • Zero-Copy Runtime — Emits standalone, relocatable .zx binaries that execute locally without a heavy Python runtime dependency.
  • Privacy-First Compilation — Models compile entirely on-device; weights and computational graphs never leave the local environment.

Installation

Install the precompiled command-line client and runtime via pip:

pip install zenthrix

You can also install and run it with uv:

uv tool install zenthrix
zenthrix --version

For a one-off invocation without installing the command globally:

uvx zenthrix --version

System Requirements

Platform Minimum Version
macOS 14.0+ (Apple Silicon M1/M2/M3/M4)
Linux Ubuntu 22.04+ (aarch64 / x86_64)
Android NDK r25+ (for targeting Snapdragon platforms)

Quickstart

1. Compile a Model

Compilation requires the separately distributed native engine. Without it, the CLI reports an actionable error rather than producing an invalid .zx file.

Compile an ONNX or GGUF model targeting local hardware execution:

zenthrix compile \
  --model meta-llama/Llama-3.2-1B-Instruct \
  --format onnx \
  --target auto \
  --quantization int4 \
  --output ./llama-3.2-1b.zx

2. Inspect Graph Optimizations

Analyze operator fusions and projected memory footprints prior to compilation:

zenthrix inspect ./llama-3.2-1b.zx --memory-profile

3. Run Inference via CLI

Verify compiled throughput directly in your terminal:

zenthrix run \
  --model ./llama-3.2-1b.zx \
  --prompt "Explain quantum decoherence in two sentences." \
  --max-tokens 128

4. Python API Usage

import zenthrix

# Load and initialize the compiled runtime
engine = zenthrix.Engine(model_path="./llama-3.2-1b.zx")

# Execute a deterministic inference pass
output = engine.generate(
    prompt="Synthesize the primary risks of high inference latency.",
    temperature=0.2,
    max_tokens=256,
)

print(output.text)
print(f"Time to First Token (TTFT): {output.ttft_ms} ms")
print(f"Throughput: {output.tokens_per_second} tokens/sec")

Supported Target Architectures

Silicon Target Optimization Backend Compute Units
Apple Silicon (M-Series / A-Series) Metal MSL & AMX Matrix Intrinsics GPU / Neural Engine
Qualcomm Snapdragon (8 Gen 2/3/4) Hexagon HTP Architecture (C++) NPU / HVX
Arm Neoverse / Cortex Arm NEON / SVE2 Assembly CPU Vector Extensions

Repository Layout

zenthrix/
├── .github/
│   ├── workflows/
│   │   ├── ci.yml
│   └── release.yml
├── python/
│   └── zenthrix/
│       ├── __init__.py
│       ├── cli.py
│       ├── config.py
│       ├── engine.py
│       ├── exceptions.py
│       ├── validation.py
│       ├── adapters/
│       │   ├── __init__.py
│       │   ├── gguf_loader.py
│       │   ├── onnx_loader.py
│       │   └── pytorch_loader.py
├── tests/
│   ├── test_cli.py
│   ├── test_engine.py
│   └── test_validation.py
├── .gitignore
├── CONTRIBUTING.md
├── LICENSE
└── pyproject.toml

Contributing

We welcome community contributions to adapters, loaders, and frontend parsers. All contributions require signing our Contributor License Agreement (CLA) during the pull request process.

See CONTRIBUTING.md for local environment setup instructions.

License

The Zenthrix CLI and client adapters are distributed under the Apache License 2.0. The underlying compilation engine dynamic binary is subject to the WithBrian Technologies Commercial EULA embedded in binary distributions.

Metadata

Release files for zenthrix 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for zenthrix 0.1.1
File Size Uploaded
zenthrix-0.1.1.tar.gz 49.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for zenthrix 0.1.1
File Interpreter ABI Platform
zenthrix-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 66.1 kB

Release files / zenthrix-0.1.1.tar.gz

Download URL zenthrix-0.1.1.tar.gz
Size 49.3 kB
Tags Source
SHA-256 checksum
How to use checksums
ae8662366b4e2db90c52fc05f90a81aeaee67f29e7552466563b00106f2aca64
BLAKE2b-256 checksum
How to use checksums
523a328baf9b1455bd4c40926587e2275f39d5d19b094e182e9750478eeaae25
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 18, 2026.

Transparency log

Release files / zenthrix-0.1.1-py3-none-any.whl

Download URL zenthrix-0.1.1-py3-none-any.whl
Size 16.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
412d3d8f5623da948c2318d80cb58274bacec58dae4f7ed66e79e66034d17ef4
BLAKE2b-256 checksum
How to use checksums
af4dc3939547e55981fecb17a013f751619f8aa3d72f70781d45087cd4d19b52
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 18, 2026.

Transparency log

Release history Release notifications | RSS feed

0.2.0

2 release files

0.1.3

2 release files

0.1.2

2 release files

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page