Skip to main content

Termux-BitNet (v2.1.0 Sovereign Ternary)

Production-Grade Universal 1.58-Bit (i2_s) On-Device LLM Inference Engine with ARM64 NEON DotProd SIMD & Native Vulkan GPU Acceleration.

PyPI Version npm Version GitHub Release Documentation Portal License Core Acceleration Supported Models Tested Hardware


🏛️ Executive Engineering Disclosure: v2.0.1 Generational Breakthroughs

termux-bitnet v2.0.1 extends the sovereign ternary architecture with hallucination-prevention prompt wrapping, hardware-level GPU watchdog defense, and dynamic LM Head memory slicing, delivering fully calibrated conversational generation across diverse 1.58-bit models.

📚 Official Academic & Upstream Credibility


📑 Supported Pretrained Models & Optimal Settings

Model Identifier Parameter Count Disk Footprint Activation Kernel Recommended Prompt Template Recommended Acceleration Flags
Microsoft BitNet 2B-4T 2.0B 1.13 GB Squared ReLU (--act-fn relu2) --prompt-template raw -ngl 30 (100% GPU Offload)
TII Falcon-E 1B Instruct 1.0B 635 MB SwiGLU (--act-fn swiglu) --prompt-template falcon -ngl 24 (100% GPU Offload)
TII Falcon3 7B Instruct 7.45B 3.05 GB SwiGLU (--act-fn swiglu) --prompt-template chatml --chunk-layers 4 --vocab-slice 32768
BitNet Embedding 270M 268M 367 MB Linear / RMSNorm --prompt-template raw CPU Zero-Copy Mmap (30+ tok/s)

🛠️ Comprehensive CLI & SDK Parameter Reference Manual

The inference runtime provides user-directed controls to eliminate hallucinations and tailor execution to mobile hardware:

Option Flag Type / Range Default Purpose & Architectural Behavior
--prompt-template chatml | falcon | raw | none none Wraps user input in model-compliant dialogue tokens to eliminate drift and word salad.
--act-fn auto | relu2 | swiglu auto Enforces mathematical activation kernel routing (relu(x)^2 vs SiLU(x) * up).
--chunk-layers Integer (0 = disabled, e.g. 4) 0 Submits GPU command buffers in chunks of N layers, bypassing the 2.5s Mali watchdog fence.
--vocab-slice Integer (0 = disabled, e.g. 32768) 0 Slices FP16 LM Head projection rows, slashing VRAM consumption by 576MB on 7B models.
-ngl, --n-gpu-layers Integer (0 to total layers) 0 Number of transformer layers permanently offloaded into Vulkan GPU VRAM.
--prompt-prefix String "" Custom prefix prepended to user prompt before tokenization.
--prompt-suffix String "" Custom suffix appended to user prompt before tokenization.
--eos-token-id Integer Model default Explicit override for End-of-Sequence token ID to guarantee generation termination.

⚡ Empirical Real-Device Hardware Fleet Scorecard (Ground Truth)

All benchmarks were empirically measured on genuine Samsung Galaxy hardware under unrooted Android Termux Bionic libc environments with steady thermal equilibrium and verified per-device semantic outputs:

Target Model Test Device & Hardware Offload Mode Token Generation Speed Verified Semantic Output String Status
BitNet-2B Galaxy S25 (Snapdragon 8 Elite) GPU (30/30) 19.37 tok/s "Paris. The Eiffel Tower, located in the heart of Paris," PASS
BitNet-2B Galaxy S20 (Turnip Adreno 650) GPU (30/30) 7.71 tok/s "Paris.\nIt is located in the region of Île-de-France," PASS
BitNet-2B Galaxy A35 (Exynos 1380 Mali-G68) GPU (30/30) 4.22 tok/s " Paris. London has a very long name, but its people are friendly and" PASS
BitNet-2B Galaxy A53 (Exynos 1280 Mali-G68) GPU (30/30) 3.26 tok/s " Paris. Cathy has a lot more money than David does. The new" PASS
Falcon-E-1B Galaxy S25 (Snapdragon 8 Elite) GPU (24/24) 34.35 tok/s "situated in Paris, a city known for its rich history and cul" PASS
Falcon-E-1B Galaxy S20 (Turnip Adreno 650) GPU (24/24) 10.76 tok/s "located in the Southern Italy, bordered by the Danube (D" PASS
Falcon-E-1B Galaxy A35 (Exynos 1380 Mali-G68) GPU (24/24) 5.81 tok/s " a city in the heart of art and culture, where history and innovation meet" PASS
Falcon-E-1B Galaxy A53 (Exynos 1280 Mali-G68) GPU (24/24) 4.46 tok/s " a complex and often subject to numerous laws and regulations, with different types of" PASS
Falcon3-7B Galaxy S25 (Snapdragon 8 Elite) GPU (28/28) 8.30 tok/s "a sovereign state situated mainly in Western Europe..." PASS
Falcon3-7B Galaxy S20 (Turnip Adreno 650) GPU (-ngl 8) 2.62 tok/s "a large city in the north-west region..." PASS
Falcon3-7B Galaxy A53 (Exynos 1280, 6GB RAM) GPU Chunked (4) 1.74 tok/s " the only power that can be considered to exist for" PASS
Falcon3-7B Galaxy A35 (Exynos 1380, 6GB RAM) GPU Chunked (4) 0.78 tok/s " the largest city in Europe by population, and its" PASS

📦 Installation

1. Python SDK (PyPI)

# Inside Android Termux (prerequisites: clang cmake python openblas)
pkg update && pkg install -y clang cmake python openblas

# Install Termux-BitNet and AMEVA-Runtime
pip install --upgrade termux-bitnet ameva-runtime

2. Node.js CLI (npm)

npm install -g termux-bitnet @ameva/runtime

3. One-Touch Native Installer

curl -sSL https://raw.githubusercontent.com/uno-km/termux-bitnet/main/install.sh | bash

🚀 Quickstart Recipes

1. Python Programmatic Inference

from termux_bitnet import BitNetEngine, BitNetConfig

# Configure for real-time Falcon-E-1B or BitNet-2B with template wrapping
config = BitNetConfig(
    model_path="~/.cache/termux-bitnet/models/falcon-e-1b-instruct-i2_s.gguf",
    device="gpu",           # "auto", "cpu", or "gpu"
    n_gpu_layers=24,        # 100% GPU offload
    chat_template="falcon", # Formats with <|im_start|>user\n...
    chunk_layers=4,         # Prevent Mali watchdog timeouts
    vocab_slice=32768,      # Slashing LM Head VRAM
    n_threads=4,
    temperature=0.7,
    top_p=0.95
)

with BitNetEngine(config) as engine:
    print("[Prompt]: What is the capital of France?")
    print("[Response]: ", end="", flush=True)
    for token in engine.generate_stream("What is the capital of France?"):
        print(token, end="", flush=True)
    print()
    metrics = engine.get_last_metrics()
    print(f"Speed: {metrics.tokens_per_second:.2f} tok/s")

2. Standalone CLI Usage

# 1. Run 1B real-time conversational model with Falcon wrapping
termux-bitnet run -m models/falcon-e-1b-instruct-i2_s.gguf \
  -p "What is the capital of France?" --prompt-template falcon -ngl 24 -t 4

# 2. Run Microsoft BitNet 2B with Squared ReLU
termux-bitnet run -m models/bitnet-2b-ggml-model-i2_s.gguf \
  -p "What is the capital of France?" --act-fn relu2 --prompt-template raw -ngl 30 -t 4

# 3. Run 7.45B model on 6GB RAM device with GPU Chunking and Vocab Slicing
termux-bitnet run -m models/falcon3-7b-instruct-1.58bit-i2_s.gguf \
  -p "Explain quantum computing simply" --prompt-template chatml \
  --chunk-layers 4 --vocab-slice 32768 -ngl 12 -t 4

Distributed Clustering & Memory Pooling (AMEVA-Cluster)

Termux-BitNet integrates with AMEVA-Cluster (pip install ameva-cluster) for distributed 1.58-bit ternary tensor sharding and cross-device memory pooling.

1. Install Cluster Runtime

pip install ameva-cluster
# or Node.js:
npm install @ameva/cluster

2. Launch Worker Node on Remote Phone

# On remote worker device (e.g. Galaxy A53):
ameva-cluster worker --port 50052

3. Run Distributed 1.58-Bit Inference

# Master node sharding Falcon3-7B or BitNet-2B across remote phones:
termux-bitnet run -m models/bitnet_b1_58-3B.gguf \
  --rpc 192.0.2.10:50052,192.0.2.11:50052 \
  -p "Explain ternary quantization benefits."
from termux_bitnet import BitNetEngine, BitNetConfig

config = BitNetConfig(
    model_path="models/bitnet_b1_58-3B.gguf",
    cluster_rpc_servers="192.0.2.10:50052,192.0.2.11:50052"
)
with BitNetEngine(config) as engine:
    for token in engine.generate_stream("Distributed clustering active across ternary nodes."):
        print(token, end="", flush=True)

Metadata

Release files for termux-bitnet 2.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for termux-bitnet 2.1.0
File Size Uploaded
termux_bitnet-2.1.0.tar.gz 132.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for termux-bitnet 2.1.0
File Interpreter ABI Platform
termux_bitnet-2.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 182.6 kB

Release files / termux_bitnet-2.1.0.tar.gz

Download URL termux_bitnet-2.1.0.tar.gz
Size 132.7 kB
Tags Source
SHA-256 checksum
How to use checksums
948357ba42c0b9721f9106dceedc5ad372bbef555512aad1373e2c2f8ff02b10
BLAKE2b-256 checksum
How to use checksums
8f07d142adc655eada2184473bb6c73203377144a2321d43b8637d5b6bef5c4a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.0

Release files / termux_bitnet-2.1.0-py3-none-any.whl

Download URL termux_bitnet-2.1.0-py3-none-any.whl
Size 49.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c15c2e1d8992531f02c5083f72f434c859a604e5bc5bc28517a6ceaefd25a139
BLAKE2b-256 checksum
How to use checksums
f4066cb8947c099f1beb5370d5f564bf016ce30728103a491ce13ec6b92b77b2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.0

Release history Release notifications | RSS feed

This release

2.1.0 This release

2 release files

2.0.1

2 release files

2.0.0

2 release files

1.4.6

2 release files

1.4.5

2 release files

1.4.4

2 release files

1.4.3

2 release files

1.4.2

2 release files

1.4.1

2 release files

1.4.0

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.1.5

2 release files

1.1.3

1 release file

1.1.2

1 release file

1.1.1

1 release file

1.1.0

2 release files

1.0.16

2 release files

1.0.15

2 release files

1.0.14

2 release files

1.0.13

2 release files

1.0.12

2 release files

1.0.11

2 release files

1.0.10

2 release files

1.0.9

2 release files

1.0.8

2 release files

1.0.7

2 release files

1.0.5

2 release files

1.0.4

2 release files

1.0.3

2 release files

1.0.1

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page