Skip to main content

Termux-BitNet

Production 1.58-bit (i2_s) BitNet On-Device Inference SDK with Native Vulkan Compute GPU Acceleration for Android Termux & ARM64.

PyPI Version npm Version GitHub Release Documentation Portal License Vulkan GPU Compute Core C++ Platform



Empirical Real-Device Benchmarks (Microsoft BitNet b1.58 2B-4T)

Target Device SoC / Microarchitecture Compute Backend Token Generation Prompt Eval Time Speedup Verification Status
Samsung Galaxy S25 Snapdragon 8 Elite (Adreno 830) AMEVA Vulkan GPU Full Pipeline 17.558 tok/s 205.9 ms 12.58x Verified (Ground Truth)
Samsung Galaxy S25 Snapdragon 8 Elite (Oryon CPU) Native CPU (4 Threads) 1.396 tok/s 2,041.0 ms 1.00x Baseline
Samsung Galaxy A35 5G Exynos 1380 (Mali-G68) AMEVA Vulkan GPU Full Pipeline 3.471 tok/s 1,552.8 ms 5.94x Verified (Ground Truth)
Samsung Galaxy A35 5G Exynos 1380 (Cortex-A78 CPU) Native CPU (4 Threads) 0.584 tok/s 8,775.0 ms 1.00x Baseline
Samsung Galaxy S20 Snapdragon 865 (Kryo 585 CPU) Native CPU (4 Threads) 2.990 tok/s 334.4 ms Reference Historical Baseline


1. Overview & Architecture

termux-bitnet is an optimized on-device inference engine and dual SDK (Python & Node.js) engineered for running 1.58-bit quantized Large Language Models (BitNet b1.58) natively on Android Termux, ARM64 mobile processors, and edge devices.

[Python Application / CLI]        [Node.js / TypeScript App]
       │                                     │
       ▼ (BitNetEngine / ctypes)             ▼ (BitNetEngine / FFI)
[termux-bitnet Python SDK]        [termux-bitnet npm Thin Gateway]
       │                                     │
       └──────────────────┬──────────────────┘
                          │
             ┌────────────┴────────────┐
             ▼                         ▼
[Native C++17 BitNet Core]    [AMEVA-Runtime Vulkan Engine]
       │ (ARM NEON DotProd)              │ (SPIR-V Compute Shaders)
       ▼                                 ▼
[ARM64 CPU: 4 Big Cores]      [Mobile GPU: Adreno 830 / Mali-G68]

2. Installation & Verification

2.1 Python Package with Vulkan GPU Acceleration (PyPI)

# Inside Android Termux (prerequisites: clang cmake python openblas)
pkg update && pkg install -y clang cmake python openblas

# Install Termux-BitNet and AMEVA-Runtime for Vulkan GPU acceleration
pip install termux-bitnet ameva-runtime

2.2 Node.js / TypeScript Thin Gateway (npm)

# Global CLI installation (Provides 'termux-bitnet-js' command)
npm install -g termux-bitnet @ameva/runtime

2.3 Automated Installer (Precompiled Binary or Fast Native Build)

curl -sL https://raw.githubusercontent.com/uno-km/termux-bitnet/main/install.sh | bash

3. CLI Usage

3.1 Python CLI (termux-bitnet)

# 1. Hardware Diagnostic (ARM NEON, DotProd SIMD & Vulkan GPU Verification)
termux-bitnet info

# 2. Download Model with HTTP Range Resume Support
termux-bitnet download bitnet-2b

# 3. Run On-Device Inference with Native Vulkan GPU Acceleration
termux-bitnet run -m ~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf \
  -p "Explain quantum computing in one sentence." \
  --device gpu -t 4 -c 2048 -n 64 --temp 0.7 --top-p 0.95

3.2 Node.js CLI (termux-bitnet-js)

# 1. Hardware Diagnostic
termux-bitnet-js info

# 2. Run Inference via Node.js Gateway with GPU backend
termux-bitnet-js run -m ~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf \
  -p "Explain quantum computing in one sentence." -d gpu -t 4 -n 64

4. Programmatic API

4.1 Python SDK (with Vulkan GPU Mode)

from termux_bitnet import BitNetEngine, BitNetConfig

# 1. Configure Engine Parameters with Vulkan GPU Acceleration
config = BitNetConfig(
    model_path="~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf",
    device="gpu",          # "gpu" (Vulkan via AMEVA-Runtime), "cpu", or "auto"
    n_threads=4,
    temperature=0.7,
    top_p=0.95,
    top_k=40,
    min_p=0.05,
    repeat_penalty=1.15,
)

# 2. Stream Generation with Dynamic Parameters
with BitNetEngine(config) as engine:
    print("[Prompt]: Write a Python palindrome check function")
    print("[Response]: ", end="", flush=True)
    for token in engine.generate_stream("Write a Python palindrome check function:"):
        print(token, end="", flush=True)
    print()
    metrics = engine.get_last_metrics()
    print(f"Speed: {metrics.tokens_per_second:.2f} tok/s, Prompt tokens: {metrics.prompt_tokens}")

4.2 Node.js & TypeScript SDK

const { createEngine } = require('termux-bitnet');

async function main() {
  const engine = createEngine({
    modelPath: '~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf',
    device: 'gpu',         // 'gpu' (Vulkan), 'cpu', or 'auto'
    threads: 4,
    temperature: 0.7,
    topP: 0.95,
  });

  console.log('[Prompt]: Explain quantum computing in one sentence');
  console.log('[Response]: ');

  await engine.generateStream(
    'Explain quantum computing in one sentence',
    64,
    (token) => {
      process.stdout.write(token);
    }
  );
  console.log('\n');
}

main();

4.3 Direct AMEVA-Runtime Adapter API

from ameva_runtime.adapters.bitnet import BitnetAdapter

adapter = BitnetAdapter()
adapter.load_model("~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf")

for token in adapter.generate("Explain quantum computing in one sentence:"):
    print(token, end="", flush=True)
print()

5. Configuration Parameter Matrix (BitNetConfig)

CLI Flag Python (BitNetConfig) Node.js (BitNetOptions) Default Description
-m, --model model_path modelPath "" Path to GGUF model binary
-p, --prompt prompt prompt "" Input prompt text
-d, --device device device "auto" Compute backend (auto, gpu, cpu)
-t, --threads n_threads threads cores Number of CPU worker threads
-c, --ctx-size n_ctx contextSize 2048 KV Cache context window size
-b, --batch-size n_batch batchSize 512 Prompt evaluation batch size
-n, --n-predict n_predict maxTokens 128 Maximum tokens to generate
--temp temperature temperature 0.7 Softmax temperature (0.0 = Greedy)
--top-p top_p topP 0.95 Nucleus Top-P sampling cutoff
--top-k top_k topK 40 Top-K sampling cutoff
--min-p min_p minP 0.05 Min-P relative probability cutoff
--repeat-penalty repeat_penalty repeatPenalty 1.15 Repetition penalty coefficient
-s, --seed seed seed 0 Random seed (0 = non-deterministic)
--system-prompt system_prompt systemPrompt "" System prompt prefix
-r, --stop stop_tokens stopTokens "" Stop sequence tokens

6. Official Documentation & Specifications


7. License & Foundation

Released under the Apache License 2.0.
Engineered under the AMEVA Open-Source Foundation (AOSF) & uno-km ecosystem.

Metadata

Release files for termux-bitnet 1.4.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for termux-bitnet 1.4.4
File Size Uploaded
termux_bitnet-1.4.4.tar.gz 73.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for termux-bitnet 1.4.4
File Interpreter ABI Platform
termux_bitnet-1.4.4-py3-none-any.whl Python 3 none any Details

Total release size: 117.4 kB

Release files / termux_bitnet-1.4.4.tar.gz

Download URL termux_bitnet-1.4.4.tar.gz
Size 73.5 kB
Tags Source
SHA-256 checksum
How to use checksums
ce8811f5691bba09337ac8c36d7205aefc5c469dba1bb418b0ff0da76d4e91c0
BLAKE2b-256 checksum
How to use checksums
114bf10284c10705e616a5323b02594b75205034548963a25141630fe7eb2833
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release files / termux_bitnet-1.4.4-py3-none-any.whl

Download URL termux_bitnet-1.4.4-py3-none-any.whl
Size 43.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
eccb00f9aca6125e2df23bc43f8d8e1067ac519b52e5c7c4f40ea299a6a03be2
BLAKE2b-256 checksum
How to use checksums
098dd3f5386abb228aaede049e6e3d1a0133c3bd4c504b3c9010705633ecd395
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release history Release notifications | RSS feed

2.1.0

2 release files

2.0.1

2 release files

2.0.0

2 release files

1.4.6

2 release files

1.4.5

2 release files

This release

1.4.4 This release

2 release files

1.4.3

2 release files

1.4.2

2 release files

1.4.1

2 release files

1.4.0

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.1.5

2 release files

1.1.3

1 release file

1.1.2

1 release file

1.1.1

1 release file

1.1.0

2 release files

1.0.16

2 release files

1.0.15

2 release files

1.0.14

2 release files

1.0.13

2 release files

1.0.12

2 release files

1.0.11

2 release files

1.0.10

2 release files

1.0.9

2 release files

1.0.8

2 release files

1.0.7

2 release files

1.0.5

2 release files

1.0.4

2 release files

1.0.3

2 release files

1.0.1

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page