Skip to main content

Termux-BitNet

Production 1.58-bit (i2_s) BitNet On-Device Inference SDK & Dual Engine for Android Termux & ARM64.

PyPI Version npm Version GitHub Release Documentation Portal License Core C++ Platform


[!WARNING]

Public Engineering Disclosure & Upstream Contributions: Permanent Elimination of Interim Heuristics

In earlier development iterations prior to v1.3.0, when confronted with upstream ARM dequantization mismatches and missing NEON kernel paths in the upstream repository, an interim heuristic fallback was temporarily utilized to produce candidate outputs under mobile device constraints.

This technical limitation has been completely resolved, validated on real hardware, and permanently eliminated. Through architectural reverse engineering and direct upstream contributions to Microsoft's official microsoft/BitNet ecosystem:

  1. Upstream PR #551 (microsoft/BitNet#551): Identified and isolated the original ARM i2_s tensor corruption ("word salad") and layout divergence on mobile architectures.
  2. Upstream PR #624 (microsoft/BitNet#624): Fully implemented the missing __ARM_NEON 4-row parallel kernel (1x4_32W) utilizing ARMv8.2-A sdot hardware dot-product acceleration, automated Android Termux environment detection, and validated genuine on-device inference on Samsung Galaxy devices (Snapdragon 8 Elite and Exynos 1380).
  3. Mathematical Resolution of Dequantization Centering: Solved the 32-way interleaved dequantization center mismatch (correcting unsigned raw ${0, 1, 2}$ mapping vs. centered ${-1, 0, 1}$ dot product with activation summation), eradicating repetitive token degeneration defects.
  4. Strict Zero-Mock & Fail-Fast Engineering Standard: Operating strictly with 100% genuine on-device C++ inference. If a native binary or kernel cannot execute, the engine strictly fails fast with explicit error codes and remediation steps rather than emitting deceptive mock responses.

Verified Real-Device On-Device Benchmarks (BitNet b1.58 2B-4T i2_s natively inside Android Termux):

  • Samsung Galaxy S25 (Snapdragon 8 Elite / Oryon CPU): 5.12 tokens/sec (~195.3 ms/tok prompt eval)
  • Samsung Galaxy A35 5G (Samsung Exynos 1380): 1.57 tokens/sec (~636.9 ms/tok prompt eval)

1. Overview & Architecture

termux-bitnet is an optimized on-device inference engine and dual SDK (Python & Node.js) engineered for running 1.58-bit quantized Large Language Models (BitNet b1.58) natively on Android Termux, ARM64 mobile processors, and edge devices.

The underlying computation engine executes 1.58-bit ternary quantized weights {-1, 0, +1} directly via hand-vectorized ARM64 NEON SIMD and DotProd vector instructions (vdotq_s32), replacing floating-point matrix multiplications with integer additions and subtractions under a sub-250MB RAM footprint.

[Python Application / CLI]        [Node.js / TypeScript App]
       │                                     │
       ▼ (BitNetEngine / ctypes)             ▼ (BitNetEngine / FFI)
[termux-bitnet Python SDK]        [termux-bitnet npm Thin Gateway]
       │                                     │
       └──────────────────┬──────────────────┘
                          │
                          ▼ (Strict C ABI: libtermux_bitnet.so)
         [Native C++17 BitNet Core Engine]
                          │
                          ▼
    [ARM64 NEON + DotProd (vdotq_s32) Vector Kernels]

2. Verified BitNet GGUF Model Registry

termux-bitnet provides deterministic model downloading and caching from verified Hugging Face repositories with HTTP Range resume capability:

Model Alias Hugging Face Repository & File Parameters Quantization File Size Target Device
bitnet-2b microsoft/bitnet-b1.58-2B-4T-gguf 2.4B i2_s 1.13 GB Flagship Phones (Galaxy S20+, S24, S25, Pixel)
bitnet-large RichardErkhov/1bitLLM_-_bitnet_b1_58-large-gguf 0.7B Q4_0 404 MB Entry-level / Low-RAM ARM64 Devices
bitnet-3b Green-Sky/bitnet_b1_58-3B-GGUF 3.3B q1_3 730 MB High-Capacity Mobile Workstations
bitnet-3b-q4 RichardErkhov/1bitLLM_-_bitnet_b1_58-3B-gguf 3.3B Q4_0 1.83 GB High-Precision Quantized Model

3. Installation & Verification

3.1 Automated Installer (Precompiled Binary or Fast Native Build)

The recommended installation method uses install.sh, which automatically downloads verified ARM64 prebuilt binaries from GitHub Releases or compiles the native C++ core on the device:

curl -sL https://raw.githubusercontent.com/uno-km/termux-bitnet/main/install.sh | bash

3.2 Python Package (PyPI)

# Inside Android Termux (prerequisites: clang cmake python openblas)
pkg update && pkg install -y clang cmake python openblas
pip install termux-bitnet

3.3 Node.js / TypeScript Thin Gateway (npm)

# Global CLI installation (Provides 'termux-bitnet-js' command)
npm install -g termux-bitnet

3.4 Hardened Manual Compilation Workflow

To compile the native C++ engine manually with verified hardware acceleration on Samsung Exynos (Cortex-A78/A55) or Qualcomm Snapdragon (Oryon/Kryo):

# 1. Install prerequisites in Termux
pkg install -y clang cmake openblas libandroid-execinfo

# 2. Configure CMake with explicit NEON + DotProd and Clang toolchain
cmake -B build   -DGGML_NEON=ON   -DGGML_ARM_DOTPROD=ON   -DCMAKE_C_COMPILER=clang   -DCMAKE_CXX_COMPILER=clang++   -DCMAKE_BUILD_TYPE=Release

# 3. Build native standalone CLI and shared library
cmake --build build --target termux-bitnet-cli -j$(nproc 2>/dev/null || echo 4)

4. CLI Usage

4.1 Python CLI (termux-bitnet)

# 1. Hardware Diagnostic (ARM NEON & DotProd SIMD Verification)
termux-bitnet info

# 2. List Available Verified Models
termux-bitnet models

# 3. Download Model with Range Resume Support
termux-bitnet download bitnet-2b

# 4. Run On-Device Inference
termux-bitnet run -m ~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf   -p "Explain quantum computing in one sentence."   -t 4 -c 2048 -n 64 --temp 0.7 --top-p 0.95

4.2 Node.js CLI (termux-bitnet-js)

# 1. Hardware Diagnostic
termux-bitnet-js info

# 2. Model Registry List
termux-bitnet-js models

# 3. Run Inference via Node.js Gateway
termux-bitnet-js run -m ~/.cache/termux-bitnet/models/bitnet-2b-ggml-model-i2_s.gguf   -p "Explain quantum computing in one sentence." -t 4 -n 64

5. Programmatic API

5.1 Python SDK

from termux_bitnet import BitNetEngine, BitNetConfig

# 1. Configure Engine Parameters
config = BitNetConfig(
    model_path="models/bitnet-2b.gguf",
    n_threads=4,
    device="auto",
    temperature=0.7,
    top_p=0.95,
    top_k=40,
    min_p=0.05,
    repeat_penalty=1.15,
)

# 2. Stream Generation with Dynamic Parameters
with BitNetEngine(config) as engine:
    print("[Prompt]: Write a Python palindrome check function")
    print("[Response]: ", end="", flush=True)
    for token in engine.generate_stream("Write a Python palindrome check function:"):
        print(token, end="", flush=True)
    print()
    metrics = engine.get_last_metrics()
    print(f"Speed: {metrics.tokens_per_second:.2f} tok/s, Prompt tokens: {metrics.prompt_tokens}")

5.2 Node.js & TypeScript SDK

const { createEngine } = require('termux-bitnet');

async function main() {
  const engine = createEngine({
    modelPath: 'models/bitnet-2b.gguf',
    threads: 4,
    device: 'auto',
    temperature: 0.7,
    topP: 0.95,
  });

  console.log('[Prompt]: Explain quantum computing in one sentence');
  console.log('[Response]: ');

  await engine.generateStream(
    'Explain quantum computing in one sentence',
    64,
    (token) => {
      process.stdout.write(token);
    }
  );
  console.log('\n');
}

main();

6. Configuration Parameter Matrix (BitNetConfig)

CLI Flag Python (BitNetConfig) Node.js (BitNetOptions) Default Description
-m, --model model_path modelPath "" Path to GGUF model binary
-p, --prompt prompt prompt "" Input prompt text
-t, --threads n_threads threads cores Number of CPU worker threads
-d, --device device device "auto" Compute backend (auto, gpu, cpu)
-c, --ctx-size n_ctx contextSize 2048 KV Cache context window size
-b, --batch-size n_batch batchSize 512 Prompt evaluation batch size
-n, --n-predict n_predict maxTokens 128 Maximum tokens to generate
--temp temperature temperature 0.7 Softmax temperature (0.0 = Greedy)
--top-p top_p topP 0.95 Nucleus Top-P sampling cutoff
--top-k top_k topK 40 Top-K sampling cutoff
--min-p min_p minP 0.05 Min-P relative probability cutoff
--repeat-penalty repeat_penalty repeatPenalty 1.15 Repetition penalty coefficient
-s, --seed seed seed 0 Random seed (0 = non-deterministic)
--system-prompt system_prompt systemPrompt "" System prompt prefix
-r, --stop stop_tokens stopTokens "" Stop sequence tokens

7. Official Documentation & Specifications


8. License & Foundation

Released under the Apache License 2.0.
Engineered under the AMEVA Open-Source Foundation (AOSF) & uno-km ecosystem.

Metadata

Release files for termux-bitnet 1.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for termux-bitnet 1.4.0
File Size Uploaded
termux_bitnet-1.4.0.tar.gz 70.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for termux-bitnet 1.4.0
File Interpreter ABI Platform
termux_bitnet-1.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 110.9 kB

Release files / termux_bitnet-1.4.0.tar.gz

Download URL termux_bitnet-1.4.0.tar.gz
Size 70.4 kB
Tags Source
SHA-256 checksum
How to use checksums
73342a03e964990c85f82384dc8b2ad7f70d8a5593cd2de7e0b059f0679120c5
BLAKE2b-256 checksum
How to use checksums
08c70b47a7f0b5a0d29fd15b6d9d7414f6ccc93309bdfbdb1d5287a1e0130281
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release files / termux_bitnet-1.4.0-py3-none-any.whl

Download URL termux_bitnet-1.4.0-py3-none-any.whl
Size 40.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1c68101f5b339da27ba3e1544fa8486e10d741fa61442f6fb52fe3a518be2a65
BLAKE2b-256 checksum
How to use checksums
c06ed3c896fd7efaf18c0382a8efaff4eb4a1223894df03bc564d755642e7980
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release history Release notifications | RSS feed

2.1.0

2 release files

2.0.1

2 release files

2.0.0

2 release files

1.4.6

2 release files

1.4.5

2 release files

1.4.4

2 release files

1.4.3

2 release files

1.4.2

2 release files

1.4.1

2 release files

This release

1.4.0 This release

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.1.5

2 release files

1.1.3

1 release file

1.1.2

1 release file

1.1.1

1 release file

1.1.0

2 release files

1.0.16

2 release files

1.0.15

2 release files

1.0.14

2 release files

1.0.13

2 release files

1.0.12

2 release files

1.0.11

2 release files

1.0.10

2 release files

1.0.9

2 release files

1.0.8

2 release files

1.0.7

2 release files

1.0.5

2 release files

1.0.4

2 release files

1.0.3

2 release files

1.0.1

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page