Skip to main content

Termux-BitNet (v2.0.1 Sovereign Ternary)

Production-Grade Universal 1.58-Bit (i2_s) On-Device LLM Inference Engine with ARM64 NEON DotProd SIMD & Native Vulkan GPU Acceleration.

PyPI Version npm Version GitHub Release Documentation Portal License Core Acceleration Supported Models Tested Hardware


🏛️ Executive Engineering Disclosure: v2.0.1 Generational Breakthroughs

termux-bitnet v2.0.1 extends the sovereign ternary architecture with hallucination-prevention prompt wrapping, hardware-level GPU watchdog defense, and dynamic LM Head memory slicing, delivering fully calibrated conversational generation across diverse 1.58-bit models.

📚 Official Academic & Upstream Credibility


📑 Supported Pretrained Models & Optimal Settings

Model Identifier Parameter Count Disk Footprint Activation Kernel Recommended Prompt Template Recommended Acceleration Flags
Microsoft BitNet 2B-4T 2.0B 1.13 GB Squared ReLU (--act-fn relu2) --prompt-template raw -ngl 30 (100% GPU Offload)
TII Falcon-E 1B Instruct 1.0B 635 MB SwiGLU (--act-fn swiglu) --prompt-template falcon -ngl 24 (100% GPU Offload)
TII Falcon3 7B Instruct 7.45B 3.05 GB SwiGLU (--act-fn swiglu) --prompt-template chatml --chunk-layers 4 --vocab-slice 32768
BitNet Embedding 270M 268M 367 MB Linear / RMSNorm --prompt-template raw CPU Zero-Copy Mmap (30+ tok/s)

🛠️ Comprehensive CLI & SDK Parameter Reference Manual

The inference runtime provides user-directed controls to eliminate hallucinations and tailor execution to mobile hardware:

Option Flag Type / Range Default Purpose & Architectural Behavior
--prompt-template chatml | falcon | raw | none none Wraps user input in model-compliant dialogue tokens to eliminate drift and word salad.
--act-fn auto | relu2 | swiglu auto Enforces mathematical activation kernel routing (relu(x)^2 vs SiLU(x) * up).
--chunk-layers Integer (0 = disabled, e.g. 4) 0 Submits GPU command buffers in chunks of N layers, bypassing the 2.5s Mali watchdog fence.
--vocab-slice Integer (0 = disabled, e.g. 32768) 0 Slices FP16 LM Head projection rows, slashing VRAM consumption by 576MB on 7B models.
-ngl, --n-gpu-layers Integer (0 to total layers) 0 Number of transformer layers permanently offloaded into Vulkan GPU VRAM.
--prompt-prefix String "" Custom prefix prepended to user prompt before tokenization.
--prompt-suffix String "" Custom suffix appended to user prompt before tokenization.
--eos-token-id Integer Model default Explicit override for End-of-Sequence token ID to guarantee generation termination.

⚡ Empirical Real-Device Hardware Fleet Scorecard (Ground Truth)

All benchmarks were empirically measured on genuine Samsung Galaxy hardware under unrooted Android Termux Bionic libc environments with steady thermal equilibrium and verified semantic outputs:

Target Model Test Device & Hardware Offload Mode Token Generation Speed Verified Semantic Output String Status
BitNet-2B Galaxy S25 (Snapdragon 8 Elite) GPU (30/30) 19.37 tok/s "Paris. Cathy has a lot more money..." PASS
BitNet-2B Galaxy S20 (Turnip Adreno 650) GPU (30/30) 7.71 tok/s "Paris. Cathy has a lot more money..." PASS
BitNet-2B Galaxy A35 (Exynos 1380 Mali-G68) GPU (30/30) 4.22 tok/s "Paris. Cathy has a lot more money..." PASS
BitNet-2B Galaxy A53 (Exynos 1280 Mali-G68) GPU (30/30) 3.26 tok/s "Paris. Cathy has a lot more money..." PASS
Falcon-E-1B Galaxy S25 (Snapdragon 8 Elite) GPU (24/24) 34.35 tok/s "a complex and often subject to numerous..." PASS
Falcon-E-1B Galaxy S20 (Turnip Adreno 650) GPU (24/24) 10.76 tok/s "a complex and often subject to numerous..." PASS
Falcon-E-1B Galaxy A35 (Exynos 1380 Mali-G68) GPU (24/24) 5.81 tok/s "a complex and often subject to numerous..." PASS
Falcon-E-1B Galaxy A53 (Exynos 1280 Mali-G68) GPU (24/24) 4.46 tok/s "a complex and often subject to numerous..." PASS
Falcon3-7B Galaxy S25 (Snapdragon 8 Elite) GPU (28/28) 8.30 tok/s "the only power that can be considered..." PASS
Falcon3-7B Galaxy S20 (Turnip Adreno 650) GPU (-ngl 8) 2.62 tok/s "a large city in the north-west region..." PASS
Falcon3-7B Galaxy A53 (Exynos 1280, 6GB RAM) GPU Chunked (4) 1.74 tok/s "the only power that can be considered..." PASS
Falcon3-7B Galaxy A35 (Exynos 1380, 6GB RAM) GPU Chunked (4) 0.78 tok/s "the only power that can be considered..." PASS

📦 Installation

1. Python SDK (PyPI)

# Inside Android Termux (prerequisites: clang cmake python openblas)
pkg update && pkg install -y clang cmake python openblas

# Install Termux-BitNet and AMEVA-Runtime
pip install --upgrade termux-bitnet ameva-runtime

2. Node.js CLI (npm)

npm install -g termux-bitnet @ameva/runtime

3. One-Touch Native Installer

curl -sSL https://raw.githubusercontent.com/uno-km/termux-bitnet/main/install.sh | bash

🚀 Quickstart Recipes

1. Python Programmatic Inference

from termux_bitnet import BitNetEngine, BitNetConfig

# Configure for real-time Falcon-E-1B or BitNet-2B with template wrapping
config = BitNetConfig(
    model_path="~/.cache/termux-bitnet/models/falcon-e-1b-instruct-i2_s.gguf",
    device="gpu",           # "auto", "cpu", or "gpu"
    n_gpu_layers=24,        # 100% GPU offload
    chat_template="falcon", # Formats with <|im_start|>user\n...
    chunk_layers=4,         # Prevent Mali watchdog timeouts
    vocab_slice=32768,      # Slashing LM Head VRAM
    n_threads=4,
    temperature=0.7,
    top_p=0.95
)

with BitNetEngine(config) as engine:
    print("[Prompt]: What is the capital of France?")
    print("[Response]: ", end="", flush=True)
    for token in engine.generate_stream("What is the capital of France?"):
        print(token, end="", flush=True)
    print()
    metrics = engine.get_last_metrics()
    print(f"Speed: {metrics.tokens_per_second:.2f} tok/s")

2. Standalone CLI Usage

# 1. Run 1B real-time conversational model with Falcon wrapping
termux-bitnet run -m models/falcon-e-1b-instruct-i2_s.gguf \
  -p "What is the capital of France?" --prompt-template falcon -ngl 24 -t 4

# 2. Run Microsoft BitNet 2B with Squared ReLU
termux-bitnet run -m models/bitnet-2b-ggml-model-i2_s.gguf \
  -p "What is the capital of France?" --act-fn relu2 --prompt-template raw -ngl 30 -t 4

# 3. Run 7.45B model on 6GB RAM device with GPU Chunking and Vocab Slicing
termux-bitnet run -m models/falcon3-7b-instruct-1.58bit-i2_s.gguf \
  -p "Explain quantum computing simply" --prompt-template chatml \
  --chunk-layers 4 --vocab-slice 32768 -ngl 12 -t 4

Metadata

Release files for termux-bitnet 2.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for termux-bitnet 2.0.1
File Size Uploaded
termux_bitnet-2.0.1.tar.gz 124.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for termux-bitnet 2.0.1
File Interpreter ABI Platform
termux_bitnet-2.0.1-py3-none-any.whl Python 3 none any Details

Total release size: 168.4 kB

Release files / termux_bitnet-2.0.1.tar.gz

Download URL termux_bitnet-2.0.1.tar.gz
Size 124.1 kB
Tags Source
SHA-256 checksum
How to use checksums
8d99080587e1bfc63afcc25ee1a9da7d69cec93a8dd7fe77d689c4edda897a0d
BLAKE2b-256 checksum
How to use checksums
bf7c465ec069466e6fa800f53eddbe09e63396389efebfcac798af21c21ffebe
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.0

Release files / termux_bitnet-2.0.1-py3-none-any.whl

Download URL termux_bitnet-2.0.1-py3-none-any.whl
Size 44.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9290dc110625db2b122a5c78b78e4d534f18d284246d1dfc01ba6507c68e0386
BLAKE2b-256 checksum
How to use checksums
4ec7622191031682794db6ac97a12f0d830e3d7326ce04e2fb9035ff063e51d7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.0

Release history Release notifications | RSS feed

2.1.0

2 release files

This release

2.0.1 This release

2 release files

2.0.0

2 release files

1.4.6

2 release files

1.4.5

2 release files

1.4.4

2 release files

1.4.3

2 release files

1.4.2

2 release files

1.4.1

2 release files

1.4.0

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.1.5

2 release files

1.1.3

1 release file

1.1.2

1 release file

1.1.1

1 release file

1.1.0

2 release files

1.0.16

2 release files

1.0.15

2 release files

1.0.14

2 release files

1.0.13

2 release files

1.0.12

2 release files

1.0.11

2 release files

1.0.10

2 release files

1.0.9

2 release files

1.0.8

2 release files

1.0.7

2 release files

1.0.5

2 release files

1.0.4

2 release files

1.0.3

2 release files

1.0.1

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page