Skip to main content

Termux-BitNet

Production 1.58-bit (i2_s) BitNet On-Device Inference SDK with Native Vulkan Compute GPU Acceleration for Android Termux & ARM64.

PyPI Version npm Version GitHub Release Documentation Portal License Vulkan GPU Compute Core C++ Platform



Empirical Real-Device Benchmarks (Microsoft BitNet b1.58 2B-4T)

Target Device SoC / Microarchitecture Compute Backend Token Generation Prompt Eval Time Speedup Verification Status
Samsung Galaxy S25 Snapdragon 8 Elite (Adreno 830) AMEVA Vulkan GPU Full Pipeline 17.558 tok/s 205.9 ms 12.58x Verified (Ground Truth)
Samsung Galaxy S25 Snapdragon 8 Elite (Oryon CPU) Native CPU (4 Threads) 1.396 tok/s 2,041.0 ms 1.00x Baseline
Samsung Galaxy A35 5G Exynos 1380 (Mali-G68) AMEVA Vulkan GPU Full Pipeline 3.471 tok/s 1,552.8 ms 5.94x Verified (Ground Truth)
Samsung Galaxy A35 5G Exynos 1380 (Cortex-A78 CPU) Native CPU (4 Threads) 0.584 tok/s 8,775.0 ms 1.00x Baseline
Samsung Galaxy S20 Snapdragon 865 (Kryo 585 CPU) Native CPU (4 Threads) 2.990 tok/s 334.4 ms Reference Historical Baseline


1. Overview & Architecture

termux-bitnet is an optimized on-device inference engine and dual SDK (Python & Node.js) engineered for running 1.58-bit quantized Large Language Models (BitNet b1.58) natively on Android Termux, ARM64 mobile processors, and edge devices.

[Python Application / CLI]        [Node.js / TypeScript App]
       │                                     │
       ▼ (BitNetEngine / ctypes)             ▼ (BitNetEngine / FFI)
[termux-bitnet Python SDK]        [termux-bitnet npm Thin Gateway]
       │                                     │
       └──────────────────┬──────────────────┘
                          │
             ┌────────────┴────────────┐
             ▼                         ▼
[Native C++17 BitNet Core]    [AMEVA-Runtime Vulkan Engine]
       │ (ARM NEON DotProd)              │ (SPIR-V Compute Shaders)
       ▼                                 ▼
[ARM64 CPU: 4 Big Cores]      [Mobile GPU: Adreno 830 / Mali-G68]

2. Installation & Verification

2.1 Python Package with Vulkan GPU Acceleration (PyPI)

# Inside Android Termux (prerequisites: clang cmake python openblas)
pkg update && pkg install -y clang cmake python openblas

# Install Termux-BitNet and AMEVA-Runtime for Vulkan GPU acceleration
pip install termux-bitnet ameva-runtime

2.2 Node.js / TypeScript Thin Gateway (npm)

# Global CLI installation (Provides 'termux-bitnet-js' command)
npm install -g termux-bitnet @ameva/runtime

2.3 Automated Installer (Precompiled Binary or Fast Native Build)

curl -sL https://raw.githubusercontent.com/uno-km/termux-bitnet/main/install.sh | bash

3. CLI Usage

3.1 Python CLI (termux-bitnet)

# 1. Hardware Diagnostic (ARM NEON, DotProd SIMD & Vulkan GPU Verification)
termux-bitnet info

# 2. Download Model with HTTP Range Resume Support
termux-bitnet download bitnet-2b

# 3. Run On-Device Inference with Native Vulkan GPU Acceleration
termux-bitnet run -m ~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf \
  -p "Explain quantum computing in one sentence." \
  --device gpu -t 4 -c 2048 -n 64 --temp 0.7 --top-p 0.95

3.2 Node.js CLI (termux-bitnet-js)

# 1. Hardware Diagnostic
termux-bitnet-js info

# 2. Run Inference via Node.js Gateway with GPU backend
termux-bitnet-js run -m ~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf \
  -p "Explain quantum computing in one sentence." -d gpu -t 4 -n 64

4. Programmatic API

4.1 Python SDK (with Vulkan GPU Mode)

from termux_bitnet import BitNetEngine, BitNetConfig

# 1. Configure Engine Parameters with Vulkan GPU Acceleration
config = BitNetConfig(
    model_path="~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf",
    device="gpu",          # "gpu" (Vulkan via AMEVA-Runtime), "cpu", or "auto"
    n_threads=4,
    temperature=0.7,
    top_p=0.95,
    top_k=40,
    min_p=0.05,
    repeat_penalty=1.15,
)

# 2. Stream Generation with Dynamic Parameters
with BitNetEngine(config) as engine:
    print("[Prompt]: Write a Python palindrome check function")
    print("[Response]: ", end="", flush=True)
    for token in engine.generate_stream("Write a Python palindrome check function:"):
        print(token, end="", flush=True)
    print()
    metrics = engine.get_last_metrics()
    print(f"Speed: {metrics.tokens_per_second:.2f} tok/s, Prompt tokens: {metrics.prompt_tokens}")

4.2 Node.js & TypeScript SDK

const { createEngine } = require('termux-bitnet');

async function main() {
  const engine = createEngine({
    modelPath: '~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf',
    device: 'gpu',         // 'gpu' (Vulkan), 'cpu', or 'auto'
    threads: 4,
    temperature: 0.7,
    topP: 0.95,
  });

  console.log('[Prompt]: Explain quantum computing in one sentence');
  console.log('[Response]: ');

  await engine.generateStream(
    'Explain quantum computing in one sentence',
    64,
    (token) => {
      process.stdout.write(token);
    }
  );
  console.log('\n');
}

main();

4.3 Direct AMEVA-Runtime Adapter API

from ameva_runtime.adapters.bitnet import BitnetAdapter

adapter = BitnetAdapter()
adapter.load_model("~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf")

for token in adapter.generate("Explain quantum computing in one sentence:"):
    print(token, end="", flush=True)
print()

5. Configuration Parameter Matrix (BitNetConfig)

CLI Flag Python (BitNetConfig) Node.js (BitNetOptions) Default Description
-m, --model model_path modelPath "" Path to GGUF model binary
-p, --prompt prompt prompt "" Input prompt text
-d, --device device device "auto" Compute backend (auto, gpu, cpu)
-t, --threads n_threads threads cores Number of CPU worker threads
-c, --ctx-size n_ctx contextSize 2048 KV Cache context window size
-b, --batch-size n_batch batchSize 512 Prompt evaluation batch size
-n, --n-predict n_predict maxTokens 128 Maximum tokens to generate
--temp temperature temperature 0.7 Softmax temperature (0.0 = Greedy)
--top-p top_p topP 0.95 Nucleus Top-P sampling cutoff
--top-k top_k topK 40 Top-K sampling cutoff
--min-p min_p minP 0.05 Min-P relative probability cutoff
--repeat-penalty repeat_penalty repeatPenalty 1.15 Repetition penalty coefficient
-s, --seed seed seed 0 Random seed (0 = non-deterministic)
--system-prompt system_prompt systemPrompt "" System prompt prefix
-r, --stop stop_tokens stopTokens "" Stop sequence tokens

6. Official Documentation & Specifications


7. License & Foundation

Released under the Apache License 2.0.
Engineered under the AMEVA Open-Source Foundation (AOSF) & uno-km ecosystem.

Metadata

Release files for termux-bitnet 1.4.6

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for termux-bitnet 1.4.6
File Size Uploaded
termux_bitnet-1.4.6.tar.gz 118.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for termux-bitnet 1.4.6
File Interpreter ABI Platform
termux_bitnet-1.4.6-py3-none-any.whl Python 3 none any Details

Total release size: 163.1 kB

Release files / termux_bitnet-1.4.6.tar.gz

Download URL termux_bitnet-1.4.6.tar.gz
Size 118.7 kB
Tags Source
SHA-256 checksum
How to use checksums
4fee5b0d3eb2e9b34ad24b3e9d1bb03ad208ba3a56dc26894f7ab5ecaeab5b8e
BLAKE2b-256 checksum
How to use checksums
8f282731e68d61509c17beba9db8f251cf4f797fa2ebaa6bbba24c02dcfa582b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.0

Release files / termux_bitnet-1.4.6-py3-none-any.whl

Download URL termux_bitnet-1.4.6-py3-none-any.whl
Size 44.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7396a8e20c3942cc9509f6615ed49c7a6d9564dfb8bee1acc121c9dcbb84de29
BLAKE2b-256 checksum
How to use checksums
200d6b4eb1c694ea827133c1747a4076696b037fba95afd92a278bcf7684d4da
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.0

Release history Release notifications | RSS feed

2.1.0

2 release files

2.0.1

2 release files

2.0.0

2 release files

This release

1.4.6 This release

2 release files

1.4.5

2 release files

1.4.4

2 release files

1.4.3

2 release files

1.4.2

2 release files

1.4.1

2 release files

1.4.0

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.1.5

2 release files

1.1.3

1 release file

1.1.2

1 release file

1.1.1

1 release file

1.1.0

2 release files

1.0.16

2 release files

1.0.15

2 release files

1.0.14

2 release files

1.0.13

2 release files

1.0.12

2 release files

1.0.11

2 release files

1.0.10

2 release files

1.0.9

2 release files

1.0.8

2 release files

1.0.7

2 release files

1.0.5

2 release files

1.0.4

2 release files

1.0.3

2 release files

1.0.1

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page