Termux-BitNet
Production 1.58-bit (i2_s) BitNet On-Device Inference SDK with Native Vulkan Compute GPU Acceleration for Android Termux & ARM64.
Empirical Real-Device Benchmarks (Microsoft BitNet b1.58 2B-4T)
| Target Device | SoC / Microarchitecture | Compute Backend | Token Generation | Prompt Eval Time | Speedup | Verification Status |
|---|---|---|---|---|---|---|
| Samsung Galaxy S25 | Snapdragon 8 Elite (Adreno 830) | AMEVA Vulkan GPU Full Pipeline | 17.558 tok/s | 205.9 ms | 12.58x | Verified (Ground Truth) |
| Samsung Galaxy S25 | Snapdragon 8 Elite (Oryon CPU) | Native CPU (4 Threads) | 1.396 tok/s | 2,041.0 ms | 1.00x | Baseline |
| Samsung Galaxy A35 5G | Exynos 1380 (Mali-G68) | AMEVA Vulkan GPU Full Pipeline | 3.471 tok/s | 1,552.8 ms | 5.94x | Verified (Ground Truth) |
| Samsung Galaxy A35 5G | Exynos 1380 (Cortex-A78 CPU) | Native CPU (4 Threads) | 0.584 tok/s | 8,775.0 ms | 1.00x | Baseline |
| Samsung Galaxy S20 | Snapdragon 865 (Kryo 585 CPU) | Native CPU (4 Threads) | 2.990 tok/s | 334.4 ms | Reference | Historical Baseline |
1. Overview & Architecture
termux-bitnet is an optimized on-device inference engine and dual SDK (Python & Node.js) engineered for running 1.58-bit quantized Large Language Models (BitNet b1.58) natively on Android Termux, ARM64 mobile processors, and edge devices.
[Python Application / CLI] [Node.js / TypeScript App]
│ │
▼ (BitNetEngine / ctypes) ▼ (BitNetEngine / FFI)
[termux-bitnet Python SDK] [termux-bitnet npm Thin Gateway]
│ │
└──────────────────┬──────────────────┘
│
┌────────────┴────────────┐
▼ ▼
[Native C++17 BitNet Core] [AMEVA-Runtime Vulkan Engine]
│ (ARM NEON DotProd) │ (SPIR-V Compute Shaders)
▼ ▼
[ARM64 CPU: 4 Big Cores] [Mobile GPU: Adreno 830 / Mali-G68]
2. Installation & Verification
2.1 Python Package with Vulkan GPU Acceleration (PyPI)
# Inside Android Termux (prerequisites: clang cmake python openblas)
pkg update && pkg install -y clang cmake python openblas
# Install Termux-BitNet and AMEVA-Runtime for Vulkan GPU acceleration
pip install termux-bitnet ameva-runtime
2.2 Node.js / TypeScript Thin Gateway (npm)
# Global CLI installation (Provides 'termux-bitnet-js' command)
npm install -g termux-bitnet @ameva/runtime
2.3 Automated Installer (Precompiled Binary or Fast Native Build)
curl -sL https://raw.githubusercontent.com/uno-km/termux-bitnet/main/install.sh | bash
3. CLI Usage
3.1 Python CLI (termux-bitnet)
# 1. Hardware Diagnostic (ARM NEON, DotProd SIMD & Vulkan GPU Verification)
termux-bitnet info
# 2. Download Model with HTTP Range Resume Support
termux-bitnet download bitnet-2b
# 3. Run On-Device Inference with Native Vulkan GPU Acceleration
termux-bitnet run -m ~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf \
-p "Explain quantum computing in one sentence." \
--device gpu -t 4 -c 2048 -n 64 --temp 0.7 --top-p 0.95
3.2 Node.js CLI (termux-bitnet-js)
# 1. Hardware Diagnostic
termux-bitnet-js info
# 2. Run Inference via Node.js Gateway with GPU backend
termux-bitnet-js run -m ~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf \
-p "Explain quantum computing in one sentence." -d gpu -t 4 -n 64
4. Programmatic API
4.1 Python SDK (with Vulkan GPU Mode)
from termux_bitnet import BitNetEngine, BitNetConfig
# 1. Configure Engine Parameters with Vulkan GPU Acceleration
config = BitNetConfig(
model_path="~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf",
device="gpu", # "gpu" (Vulkan via AMEVA-Runtime), "cpu", or "auto"
n_threads=4,
temperature=0.7,
top_p=0.95,
top_k=40,
min_p=0.05,
repeat_penalty=1.15,
)
# 2. Stream Generation with Dynamic Parameters
with BitNetEngine(config) as engine:
print("[Prompt]: Write a Python palindrome check function")
print("[Response]: ", end="", flush=True)
for token in engine.generate_stream("Write a Python palindrome check function:"):
print(token, end="", flush=True)
print()
metrics = engine.get_last_metrics()
print(f"Speed: {metrics.tokens_per_second:.2f} tok/s, Prompt tokens: {metrics.prompt_tokens}")
4.2 Node.js & TypeScript SDK
const { createEngine } = require('termux-bitnet');
async function main() {
const engine = createEngine({
modelPath: '~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf',
device: 'gpu', // 'gpu' (Vulkan), 'cpu', or 'auto'
threads: 4,
temperature: 0.7,
topP: 0.95,
});
console.log('[Prompt]: Explain quantum computing in one sentence');
console.log('[Response]: ');
await engine.generateStream(
'Explain quantum computing in one sentence',
64,
(token) => {
process.stdout.write(token);
}
);
console.log('\n');
}
main();
4.3 Direct AMEVA-Runtime Adapter API
from ameva_runtime.adapters.bitnet import BitnetAdapter
adapter = BitnetAdapter()
adapter.load_model("~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf")
for token in adapter.generate("Explain quantum computing in one sentence:"):
print(token, end="", flush=True)
print()
5. Configuration Parameter Matrix (BitNetConfig)
| CLI Flag | Python (BitNetConfig) |
Node.js (BitNetOptions) |
Default | Description |
|---|---|---|---|---|
-m, --model |
model_path |
modelPath |
"" |
Path to GGUF model binary |
-p, --prompt |
prompt |
prompt |
"" |
Input prompt text |
-d, --device |
device |
device |
"auto" |
Compute backend (auto, gpu, cpu) |
-t, --threads |
n_threads |
threads |
cores |
Number of CPU worker threads |
-c, --ctx-size |
n_ctx |
contextSize |
2048 |
KV Cache context window size |
-b, --batch-size |
n_batch |
batchSize |
512 |
Prompt evaluation batch size |
-n, --n-predict |
n_predict |
maxTokens |
128 |
Maximum tokens to generate |
--temp |
temperature |
temperature |
0.7 |
Softmax temperature (0.0 = Greedy) |
--top-p |
top_p |
topP |
0.95 |
Nucleus Top-P sampling cutoff |
--top-k |
top_k |
topK |
40 |
Top-K sampling cutoff |
--min-p |
min_p |
minP |
0.05 |
Min-P relative probability cutoff |
--repeat-penalty |
repeat_penalty |
repeatPenalty |
1.15 |
Repetition penalty coefficient |
-s, --seed |
seed |
seed |
0 |
Random seed (0 = non-deterministic) |
--system-prompt |
system_prompt |
systemPrompt |
"" |
System prompt prefix |
-r, --stop |
stop_tokens |
stopTokens |
"" |
Stop sequence tokens |
6. Official Documentation & Specifications
- Official Documentation Site: https://uno-km.vercel.app/lib/bitnet/
- AI Agent Context Feed: llms.txt
- Full Technical Specification: llms-full.txt
7. License & Foundation
Released under the Apache License 2.0.
Engineered under the AMEVA Open-Source Foundation (AOSF) & uno-km ecosystem.
Metadata
Release files for termux-bitnet 1.4.5
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| termux_bitnet-1.4.5.tar.gz | 73.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| termux_bitnet-1.4.5-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 117.8 kB
Release files / termux_bitnet-1.4.5.tar.gz
| Download URL | termux_bitnet-1.4.5.tar.gz |
|---|---|
| Size | 73.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f5548f7f3b13bb24dedfbd0754e254a4f3a1052264c2169129271a9cf922e753
|
|
BLAKE2b-256 checksum How to use checksums |
923e4b0bd2d153dd88fad8da3915a99061b4b5d7039fa88db1a619dbcb4ce0a4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.16
|
Release files / termux_bitnet-1.4.5-py3-none-any.whl
| Download URL | termux_bitnet-1.4.5-py3-none-any.whl |
|---|---|
| Size | 44.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
8209b8980afbdf435d31f34280b98ad2f1d75a77a529eba5d20b28c0ccb0ae2e
|
|
BLAKE2b-256 checksum How to use checksums |
a0654ce1aaf39091f4f1437568414799db26729a4df91bbc4f85d48c822a5c17
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.16
|