Skip to main content

Termux-BitNet

Production 1.58-bit (i2_s) BitNet On-Device Inference SDK with Native Vulkan Compute GPU Acceleration for Android Termux & ARM64.

PyPI Version npm Version GitHub Release Documentation Portal License Vulkan GPU Compute Core C++ Platform



Empirical Real-Device Benchmarks (Microsoft BitNet b1.58 2B-4T)

Target Device SoC / Microarchitecture Compute Backend Token Generation Prompt Eval Time Speedup Verification Status
Samsung Galaxy S25 Snapdragon 8 Elite (Adreno 830) AMEVA Vulkan GPU Full Pipeline 17.558 tok/s 205.9 ms 12.58x Verified (Ground Truth)
Samsung Galaxy S25 Snapdragon 8 Elite (Oryon CPU) Native CPU (4 Threads) 1.396 tok/s 2,041.0 ms 1.00x Baseline
Samsung Galaxy A35 5G Exynos 1380 (Mali-G68) AMEVA Vulkan GPU Full Pipeline 3.471 tok/s 1,552.8 ms 5.94x Verified (Ground Truth)
Samsung Galaxy A35 5G Exynos 1380 (Cortex-A78 CPU) Native CPU (4 Threads) 0.584 tok/s 8,775.0 ms 1.00x Baseline
Samsung Galaxy S20 Snapdragon 865 (Kryo 585 CPU) Native CPU (4 Threads) 2.990 tok/s 334.4 ms Reference Historical Baseline


1. Overview & Architecture

termux-bitnet is an optimized on-device inference engine and dual SDK (Python & Node.js) engineered for running 1.58-bit quantized Large Language Models (BitNet b1.58) natively on Android Termux, ARM64 mobile processors, and edge devices.

[Python Application / CLI]        [Node.js / TypeScript App]
       │                                     │
       ▼ (BitNetEngine / ctypes)             ▼ (BitNetEngine / FFI)
[termux-bitnet Python SDK]        [termux-bitnet npm Thin Gateway]
       │                                     │
       └──────────────────┬──────────────────┘
                          │
             ┌────────────┴────────────┐
             ▼                         ▼
[Native C++17 BitNet Core]    [AMEVA-Runtime Vulkan Engine]
       │ (ARM NEON DotProd)              │ (SPIR-V Compute Shaders)
       ▼                                 ▼
[ARM64 CPU: 4 Big Cores]      [Mobile GPU: Adreno 830 / Mali-G68]

2. Installation & Verification

2.1 Python Package with Vulkan GPU Acceleration (PyPI)

# Inside Android Termux (prerequisites: clang cmake python openblas)
pkg update && pkg install -y clang cmake python openblas

# Install Termux-BitNet and AMEVA-Runtime for Vulkan GPU acceleration
pip install termux-bitnet ameva-runtime

2.2 Node.js / TypeScript Thin Gateway (npm)

# Global CLI installation (Provides 'termux-bitnet-js' command)
npm install -g termux-bitnet @ameva/runtime

2.3 Automated Installer (Precompiled Binary or Fast Native Build)

curl -sL https://raw.githubusercontent.com/uno-km/termux-bitnet/main/install.sh | bash

3. CLI Usage

3.1 Python CLI (termux-bitnet)

# 1. Hardware Diagnostic (ARM NEON, DotProd SIMD & Vulkan GPU Verification)
termux-bitnet info

# 2. Download Model with HTTP Range Resume Support
termux-bitnet download bitnet-2b

# 3. Run On-Device Inference with Native Vulkan GPU Acceleration
termux-bitnet run -m ~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf \
  -p "Explain quantum computing in one sentence." \
  --device gpu -t 4 -c 2048 -n 64 --temp 0.7 --top-p 0.95

3.2 Node.js CLI (termux-bitnet-js)

# 1. Hardware Diagnostic
termux-bitnet-js info

# 2. Run Inference via Node.js Gateway with GPU backend
termux-bitnet-js run -m ~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf \
  -p "Explain quantum computing in one sentence." -d gpu -t 4 -n 64

4. Programmatic API

4.1 Python SDK (with Vulkan GPU Mode)

from termux_bitnet import BitNetEngine, BitNetConfig

# 1. Configure Engine Parameters with Vulkan GPU Acceleration
config = BitNetConfig(
    model_path="~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf",
    device="gpu",          # "gpu" (Vulkan via AMEVA-Runtime), "cpu", or "auto"
    n_threads=4,
    temperature=0.7,
    top_p=0.95,
    top_k=40,
    min_p=0.05,
    repeat_penalty=1.15,
)

# 2. Stream Generation with Dynamic Parameters
with BitNetEngine(config) as engine:
    print("[Prompt]: Write a Python palindrome check function")
    print("[Response]: ", end="", flush=True)
    for token in engine.generate_stream("Write a Python palindrome check function:"):
        print(token, end="", flush=True)
    print()
    metrics = engine.get_last_metrics()
    print(f"Speed: {metrics.tokens_per_second:.2f} tok/s, Prompt tokens: {metrics.prompt_tokens}")

4.2 Node.js & TypeScript SDK

const { createEngine } = require('termux-bitnet');

async function main() {
  const engine = createEngine({
    modelPath: '~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf',
    device: 'gpu',         // 'gpu' (Vulkan), 'cpu', or 'auto'
    threads: 4,
    temperature: 0.7,
    topP: 0.95,
  });

  console.log('[Prompt]: Explain quantum computing in one sentence');
  console.log('[Response]: ');

  await engine.generateStream(
    'Explain quantum computing in one sentence',
    64,
    (token) => {
      process.stdout.write(token);
    }
  );
  console.log('\n');
}

main();

4.3 Direct AMEVA-Runtime Adapter API

from ameva_runtime.adapters.bitnet import BitnetAdapter

adapter = BitnetAdapter()
adapter.load_model("~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf")

for token in adapter.generate("Explain quantum computing in one sentence:"):
    print(token, end="", flush=True)
print()

5. Configuration Parameter Matrix (BitNetConfig)

CLI Flag Python (BitNetConfig) Node.js (BitNetOptions) Default Description
-m, --model model_path modelPath "" Path to GGUF model binary
-p, --prompt prompt prompt "" Input prompt text
-d, --device device device "auto" Compute backend (auto, gpu, cpu)
-t, --threads n_threads threads cores Number of CPU worker threads
-c, --ctx-size n_ctx contextSize 2048 KV Cache context window size
-b, --batch-size n_batch batchSize 512 Prompt evaluation batch size
-n, --n-predict n_predict maxTokens 128 Maximum tokens to generate
--temp temperature temperature 0.7 Softmax temperature (0.0 = Greedy)
--top-p top_p topP 0.95 Nucleus Top-P sampling cutoff
--top-k top_k topK 40 Top-K sampling cutoff
--min-p min_p minP 0.05 Min-P relative probability cutoff
--repeat-penalty repeat_penalty repeatPenalty 1.15 Repetition penalty coefficient
-s, --seed seed seed 0 Random seed (0 = non-deterministic)
--system-prompt system_prompt systemPrompt "" System prompt prefix
-r, --stop stop_tokens stopTokens "" Stop sequence tokens

6. Official Documentation & Specifications


7. License & Foundation

Released under the Apache License 2.0.
Engineered under the AMEVA Open-Source Foundation (AOSF) & uno-km ecosystem.

Metadata

Release files for termux-bitnet 1.4.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for termux-bitnet 1.4.2
File Size Uploaded
termux_bitnet-1.4.2.tar.gz 73.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for termux-bitnet 1.4.2
File Interpreter ABI Platform
termux_bitnet-1.4.2-py3-none-any.whl Python 3 none any Details

Total release size: 117.4 kB

Release files / termux_bitnet-1.4.2.tar.gz

Download URL termux_bitnet-1.4.2.tar.gz
Size 73.5 kB
Tags Source
SHA-256 checksum
How to use checksums
053e5fd5e94a66fa0130220f586d52da78e29cb15db8f385bcbcd940d00edd37
BLAKE2b-256 checksum
How to use checksums
3fc13510b0616cf1bebc83dd12ce248c9d93621cbd08823ebc3d0a821dde6b68
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release files / termux_bitnet-1.4.2-py3-none-any.whl

Download URL termux_bitnet-1.4.2-py3-none-any.whl
Size 43.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f0339c0f2f0e07efeae21e49b053463ce765891c3a815ad5caf0a92febb92103
BLAKE2b-256 checksum
How to use checksums
55129fa5850d397d02e423e061a96579e12e31fe4d7df6ad2604395097a49bb1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release history Release notifications | RSS feed

2.1.0

2 release files

2.0.1

2 release files

2.0.0

2 release files

1.4.6

2 release files

1.4.5

2 release files

1.4.4

2 release files

1.4.3

2 release files

This release

1.4.2 This release

2 release files

1.4.1

2 release files

1.4.0

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.1.5

2 release files

1.1.3

1 release file

1.1.2

1 release file

1.1.1

1 release file

1.1.0

2 release files

1.0.16

2 release files

1.0.15

2 release files

1.0.14

2 release files

1.0.13

2 release files

1.0.12

2 release files

1.0.11

2 release files

1.0.10

2 release files

1.0.9

2 release files

1.0.8

2 release files

1.0.7

2 release files

1.0.5

2 release files

1.0.4

2 release files

1.0.3

2 release files

1.0.1

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page