Termux-BitNet (v2.0.1 Sovereign Ternary)
Production-Grade Universal 1.58-Bit (i2_s) On-Device LLM Inference Engine with ARM64 NEON DotProd SIMD & Native Vulkan GPU Acceleration.
🏛️ Executive Engineering Disclosure: v2.0.1 Generational Breakthroughs
termux-bitnet v2.0.1 extends the sovereign ternary architecture with hallucination-prevention prompt wrapping, hardware-level GPU watchdog defense, and dynamic LM Head memory slicing, delivering fully calibrated conversational generation across diverse 1.58-bit models.
📚 Official Academic & Upstream Credibility
- Foundation Technical Whitepaper: AOSF-TR-2026-BITNET-TERNARY-02: Root Cause Analysis of Ternary Numerical Collapse and Dynamic Activation Engine
- Upstream Contributions to Microsoft:
- microsoft/BitNet #551: Resolved ARM QK=128 stride memory layout desynchronization and word salad.
- microsoft/BitNet #624: Completed 1x4_32W parallel NEON
sdothardware acceleration kernel and automated Android Termux tooling.
- Upstream Contributions to GGML:
- ggml-org/whisper.cpp #4089: Mobile heterogeneous GPU-Encoder / CPU-Decoder split-mode architecture.
📑 Supported Pretrained Models & Optimal Settings
| Model Identifier | Parameter Count | Disk Footprint | Activation Kernel | Recommended Prompt Template | Recommended Acceleration Flags |
|---|---|---|---|---|---|
| Microsoft BitNet 2B-4T | 2.0B | 1.13 GB | Squared ReLU (--act-fn relu2) |
--prompt-template raw |
-ngl 30 (100% GPU Offload) |
| TII Falcon-E 1B Instruct | 1.0B | 635 MB | SwiGLU (--act-fn swiglu) |
--prompt-template falcon |
-ngl 24 (100% GPU Offload) |
| TII Falcon3 7B Instruct | 7.45B | 3.05 GB | SwiGLU (--act-fn swiglu) |
--prompt-template chatml |
--chunk-layers 4 --vocab-slice 32768 |
| BitNet Embedding 270M | 268M | 367 MB | Linear / RMSNorm | --prompt-template raw |
CPU Zero-Copy Mmap (30+ tok/s) |
🛠️ Comprehensive CLI & SDK Parameter Reference Manual
The inference runtime provides user-directed controls to eliminate hallucinations and tailor execution to mobile hardware:
| Option Flag | Type / Range | Default | Purpose & Architectural Behavior |
|---|---|---|---|
--prompt-template |
chatml | falcon | raw | none |
none |
Wraps user input in model-compliant dialogue tokens to eliminate drift and word salad. |
--act-fn |
auto | relu2 | swiglu |
auto |
Enforces mathematical activation kernel routing (relu(x)^2 vs SiLU(x) * up). |
--chunk-layers |
Integer (0 = disabled, e.g. 4) | 0 |
Submits GPU command buffers in chunks of N layers, bypassing the 2.5s Mali watchdog fence. |
--vocab-slice |
Integer (0 = disabled, e.g. 32768) | 0 |
Slices FP16 LM Head projection rows, slashing VRAM consumption by 576MB on 7B models. |
-ngl, --n-gpu-layers |
Integer (0 to total layers) | 0 |
Number of transformer layers permanently offloaded into Vulkan GPU VRAM. |
--prompt-prefix |
String | "" |
Custom prefix prepended to user prompt before tokenization. |
--prompt-suffix |
String | "" |
Custom suffix appended to user prompt before tokenization. |
--eos-token-id |
Integer | Model default | Explicit override for End-of-Sequence token ID to guarantee generation termination. |
⚡ Empirical Real-Device Hardware Fleet Scorecard (Ground Truth)
All benchmarks were empirically measured on genuine Samsung Galaxy hardware under unrooted Android Termux Bionic libc environments with steady thermal equilibrium and verified semantic outputs:
| Target Model | Test Device & Hardware | Offload Mode | Token Generation Speed | Verified Semantic Output String | Status |
|---|---|---|---|---|---|
| BitNet-2B | Galaxy S25 (Snapdragon 8 Elite) | GPU (30/30) | 19.37 tok/s | "Paris. Cathy has a lot more money..." |
PASS |
| BitNet-2B | Galaxy S20 (Turnip Adreno 650) | GPU (30/30) | 7.71 tok/s | "Paris. Cathy has a lot more money..." |
PASS |
| BitNet-2B | Galaxy A35 (Exynos 1380 Mali-G68) | GPU (30/30) | 4.22 tok/s | "Paris. Cathy has a lot more money..." |
PASS |
| BitNet-2B | Galaxy A53 (Exynos 1280 Mali-G68) | GPU (30/30) | 3.26 tok/s | "Paris. Cathy has a lot more money..." |
PASS |
| Falcon-E-1B | Galaxy S25 (Snapdragon 8 Elite) | GPU (24/24) | 34.35 tok/s | "a complex and often subject to numerous..." |
PASS |
| Falcon-E-1B | Galaxy S20 (Turnip Adreno 650) | GPU (24/24) | 10.76 tok/s | "a complex and often subject to numerous..." |
PASS |
| Falcon-E-1B | Galaxy A35 (Exynos 1380 Mali-G68) | GPU (24/24) | 5.81 tok/s | "a complex and often subject to numerous..." |
PASS |
| Falcon-E-1B | Galaxy A53 (Exynos 1280 Mali-G68) | GPU (24/24) | 4.46 tok/s | "a complex and often subject to numerous..." |
PASS |
| Falcon3-7B | Galaxy S25 (Snapdragon 8 Elite) | GPU (28/28) | 8.30 tok/s | "the only power that can be considered..." |
PASS |
| Falcon3-7B | Galaxy S20 (Turnip Adreno 650) | GPU (-ngl 8) | 2.62 tok/s | "a large city in the north-west region..." |
PASS |
| Falcon3-7B | Galaxy A53 (Exynos 1280, 6GB RAM) | GPU Chunked (4) | 1.74 tok/s | "the only power that can be considered..." |
PASS |
| Falcon3-7B | Galaxy A35 (Exynos 1380, 6GB RAM) | GPU Chunked (4) | 0.78 tok/s | "the only power that can be considered..." |
PASS |
📦 Installation
1. Python SDK (PyPI)
# Inside Android Termux (prerequisites: clang cmake python openblas)
pkg update && pkg install -y clang cmake python openblas
# Install Termux-BitNet and AMEVA-Runtime
pip install --upgrade termux-bitnet ameva-runtime
2. Node.js CLI (npm)
npm install -g termux-bitnet @ameva/runtime
3. One-Touch Native Installer
curl -sSL https://raw.githubusercontent.com/uno-km/termux-bitnet/main/install.sh | bash
🚀 Quickstart Recipes
1. Python Programmatic Inference
from termux_bitnet import BitNetEngine, BitNetConfig
# Configure for real-time Falcon-E-1B or BitNet-2B with template wrapping
config = BitNetConfig(
model_path="~/.cache/termux-bitnet/models/falcon-e-1b-instruct-i2_s.gguf",
device="gpu", # "auto", "cpu", or "gpu"
n_gpu_layers=24, # 100% GPU offload
chat_template="falcon", # Formats with <|im_start|>user\n...
chunk_layers=4, # Prevent Mali watchdog timeouts
vocab_slice=32768, # Slashing LM Head VRAM
n_threads=4,
temperature=0.7,
top_p=0.95
)
with BitNetEngine(config) as engine:
print("[Prompt]: What is the capital of France?")
print("[Response]: ", end="", flush=True)
for token in engine.generate_stream("What is the capital of France?"):
print(token, end="", flush=True)
print()
metrics = engine.get_last_metrics()
print(f"Speed: {metrics.tokens_per_second:.2f} tok/s")
2. Standalone CLI Usage
# 1. Run 1B real-time conversational model with Falcon wrapping
termux-bitnet run -m models/falcon-e-1b-instruct-i2_s.gguf \
-p "What is the capital of France?" --prompt-template falcon -ngl 24 -t 4
# 2. Run Microsoft BitNet 2B with Squared ReLU
termux-bitnet run -m models/bitnet-2b-ggml-model-i2_s.gguf \
-p "What is the capital of France?" --act-fn relu2 --prompt-template raw -ngl 30 -t 4
# 3. Run 7.45B model on 6GB RAM device with GPU Chunking and Vocab Slicing
termux-bitnet run -m models/falcon3-7b-instruct-1.58bit-i2_s.gguf \
-p "Explain quantum computing simply" --prompt-template chatml \
--chunk-layers 4 --vocab-slice 32768 -ngl 12 -t 4
🌐 Ecosystem Links & Documentation
- Official Documentation Portal: https://uno-km.vercel.app/lib/bitnet/
- Foundation Research Lab: https://uno-km.vercel.app/labs/
- License: Apache-2.0
- Governance: AMEVA Open-Source Foundation (AOSF) & @uno-km
Metadata
Release files for termux-bitnet 2.0.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| termux_bitnet-2.0.1.tar.gz | 124.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| termux_bitnet-2.0.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 168.4 kB
Release files / termux_bitnet-2.0.1.tar.gz
| Download URL | termux_bitnet-2.0.1.tar.gz |
|---|---|
| Size | 124.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8d99080587e1bfc63afcc25ee1a9da7d69cec93a8dd7fe77d689c4edda897a0d
|
|
BLAKE2b-256 checksum How to use checksums |
bf7c465ec069466e6fa800f53eddbe09e63396389efebfcac798af21c21ffebe
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.0
|
Release files / termux_bitnet-2.0.1-py3-none-any.whl
| Download URL | termux_bitnet-2.0.1-py3-none-any.whl |
|---|---|
| Size | 44.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
9290dc110625db2b122a5c78b78e4d534f18d284246d1dfc01ba6507c68e0386
|
|
BLAKE2b-256 checksum How to use checksums |
4ec7622191031682794db6ac97a12f0d830e3d7326ce04e2fb9035ff063e51d7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.0
|