Zenthrix
Hardware-Adaptive Edge Neural Graph Compiler
Zenthrix is an edge-native model compiler frontend for compiling open-weight neural networks (LLMs, SLMs, and vision models) into zero-copy, memory-optimized binaries tailored for consumer edge silicon — Apple Silicon, Qualcomm Snapdragon NPU, and Arm Cortex/Ethos.
Table of Contents
Key Features
- Direct Ingestion — Native loaders for ONNX, PyTorch Export (AOTInductor), and GGUF architectures, with no intermediate format conversion required.
- Unified Memory Tiling — Schedules compute passes against unified memory architectures, reducing peak active RAM allocation by up to 40%.
- Zero-Copy Runtime — Emits standalone, relocatable
.zxbinaries that execute locally without a heavy Python runtime dependency. - Privacy-First Compilation — Models compile entirely on-device; weights and computational graphs never leave the local environment.
Installation
Install the precompiled command-line client and runtime via pip:
pip install zenthrix
You can also install and run it with uv:
uv tool install zenthrix
zenthrix --version
For a one-off invocation without installing the command globally:
uvx zenthrix --version
System Requirements
| Platform | Minimum Version |
|---|---|
| macOS | 14.0+ (Apple Silicon M1/M2/M3/M4) |
| Linux | Ubuntu 22.04+ (aarch64 / x86_64) |
| Android | NDK r25+ (for targeting Snapdragon platforms) |
Quickstart
1. Compile a Model
Compilation requires the separately distributed native engine. Without it, the
CLI reports an actionable error rather than producing an invalid .zx file.
Compile an ONNX or GGUF model targeting local hardware execution:
zenthrix compile \
--model meta-llama/Llama-3.2-1B-Instruct \
--format onnx \
--target auto \
--quantization int4 \
--output ./llama-3.2-1b.zx
2. Inspect Graph Optimizations
Analyze operator fusions and projected memory footprints prior to compilation:
zenthrix inspect ./llama-3.2-1b.zx --memory-profile
3. Run Inference via CLI
Verify compiled throughput directly in your terminal:
zenthrix run \
--model ./llama-3.2-1b.zx \
--prompt "Explain quantum decoherence in two sentences." \
--max-tokens 128
4. Python API Usage
import zenthrix
# Load and initialize the compiled runtime
engine = zenthrix.Engine(model_path="./llama-3.2-1b.zx")
# Execute a deterministic inference pass
output = engine.generate(
prompt="Synthesize the primary risks of high inference latency.",
temperature=0.2,
max_tokens=256,
)
print(output.text)
print(f"Time to First Token (TTFT): {output.ttft_ms} ms")
print(f"Throughput: {output.tokens_per_second} tokens/sec")
Supported Target Architectures
| Silicon Target | Optimization Backend | Compute Units |
|---|---|---|
| Apple Silicon (M-Series / A-Series) | Metal MSL & AMX Matrix Intrinsics | GPU / Neural Engine |
| Qualcomm Snapdragon (8 Gen 2/3/4) | Hexagon HTP Architecture (C++) | NPU / HVX |
| Arm Neoverse / Cortex | Arm NEON / SVE2 Assembly | CPU Vector Extensions |
Contributing
We welcome community contributions to adapters, loaders, and frontend parsers. All contributions require signing our Contributor License Agreement (CLA) during the pull request process.
See CONTRIBUTING.md for local environment setup instructions.
License
The Zenthrix CLI and client adapters are distributed under the Apache License 2.0. The underlying compilation engine dynamic binary is subject to the WithBrian Technologies Commercial EULA embedded in binary distributions.
Metadata
Release files for zenthrix 0.1.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| zenthrix-0.1.2.tar.gz | 49.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| zenthrix-0.1.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 66.5 kB
Release files / zenthrix-0.1.2.tar.gz
| Download URL | zenthrix-0.1.2.tar.gz |
|---|---|
| Size | 49.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
32c9f9b295e3f81fd8400a86fd588b6b81a91c3115328499e23f8e04c15199ab
|
|
BLAKE2b-256 checksum How to use checksums |
48f36bd7bde9db919e4f19ca16f3c9f3a244a8cf9966a54eeb079b62681c9a7e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 18, 2026.
Transparency logRelease files / zenthrix-0.1.2-py3-none-any.whl
| Download URL | zenthrix-0.1.2-py3-none-any.whl |
|---|---|
| Size | 16.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
489dc1494abfe3400758e6cad79d69b9be047854f4276342287b2a487199d435
|
|
BLAKE2b-256 checksum How to use checksums |
517f43012abeb2e8d66269010f4bcfa54b14987a665772dc571ad5a0e7a7214e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 18, 2026.
Transparency log