Skip to main content

termux-llamacpp

Production-Grade, Prebuilt GGUF LLM Runtime, Model Manager & OpenAI-Compatible Server for Android Termux & ARM64

License Platform Architecture OpenAI API Zero Compilation

Notice: This project is an independent open-source runtime and is not affiliated with or endorsed by Meta Platforms, Inc. or the upstream llama.cpp maintainers.


📌 Overview

termux-llamacpp is a lightweight, zero-compilation local inference runtime and OpenAI-compliant REST/SSE supervisor tailored specifically for Android Termux and ARM64 mobile environments.

By shipping verified prebuilt Android Bionic native binaries (llama-cli, llama-server) with bundled shared libraries and cryptographic SHA-256 validation, it completely eliminates multi-gigabyte compiler toolchains (clang, cmake, ninja) and lengthy compilation wait times on mobile devices.


⚡ Key Highlights & Real-Device Benchmarks

Tested on Samsung Galaxy S20+ 5G (Snapdragon 865 / Kryo 585 Octa-core ARM64) running Termux on Android 13:

Metric Measured Ground Truth Notes
Model Meta Llama 3.2 3B Instruct (Q4_K_M, 1.92 GiB) 3,212.75M parameters
Prompt Processing Speed 16.19 tokens / sec (61.77 ms / token) 38 tokens evaluated in 2.34s
Token Generation Speed 10.23 tokens / sec (97.75 ms / token) Real-time interactive generation
Cached Prefix Speed 11.08 tokens / sec (804.7 ms total) Prompt cache reuse enabled
Cold Model Load Time ~1.8 seconds Direct memory sequential loading (--no-mmap)
HTTP Server Startup ~2.1 seconds Loopback binding with reverse proxy supervisor
Installation Time < 3 seconds Instant prebuilt binary extraction (install.sh)

🚀 Quick Start

1. Zero-Compilation One-Line Installation (Android Termux)

# Download and install prebuilt ARM64 binaries directly to ~/.termux-llama
curl -sSL https://raw.githubusercontent.com/uno-km/termux-llamacpp/master/scripts/install.sh | bash

For developers wishing to compile locally from pinned source:

bash scripts/install.sh --from-source

2. Python Package Installation

pip install termux-llamacpp

🛠️ CLI Usage

System & Hardware Diagnostics

termux-llama doctor
# or
termux-llama hardware

Example Output:

================================================================================
  termux-llamacpp Hardware & System Profile
================================================================================
  Architecture        : aarch64 (ARM64: True)
  Android / Termux    : Android=True, Termux=True
  CPU Topology        : 8 Cores (Recommended Threads: 4)
  SIMD Acceleration   : NEON=True, FP16=True, DotProd=True
  Memory Footprint    : Available 3887.8 MB / Total 10601.6 MB
  Recommended Preset  : android-arm64-dotprod
================================================================================

Download GGUF Models

# Download curated alias with SHA-256 checksum verification
termux-llama download qwen2.5-1.5b-instruct

# Download any custom repository from Hugging Face
termux-llama download bartowski/Llama-3.2-3B-Instruct-GGUF Llama-3.2-3B-Instruct-Q4_K_M.gguf

Launch OpenAI-Compatible HTTP / SSE Server

# Foreground execution (interactive)
termux-llama serve qwen2.5-1.5b-instruct --port 8080 --ctx 2048 --threads 4

# Background Daemon mode (frees current terminal session immediately)
termux-llama serve qwen2.5-1.5b-instruct -d

# Stop background server instances
termux-llama stop

🌐 OpenAI-Compatible API Endpoints

Once the supervisor server is active, it exposes standard endpoints:

1. Health & Readiness (GET /health)

curl -s http://127.0.0.1:8080/health
{
  "status": "ok",
  "ready": true,
  "service": "llama-server",
  "protocolVersion": "1.0",
  "model": {
    "id": "Llama-3.2-3B-Instruct-Q4_K_M.gguf"
  }
}

2. Model Discovery (GET /v1/models)

curl -s http://127.0.0.1:8080/v1/models

3. Non-Streaming Chat Completion (POST /v1/chat/completions)

curl -X POST http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Llama-3.2-3B-Instruct-Q4_K_M.gguf",
    "messages": [{"role": "user", "content": "Explain quantum computing in one sentence."}],
    "temperature": 0.2,
    "max_tokens": 64
  }'

4. Real-Time SSE Streaming (POST /v1/chat/completions)

curl -N -X POST http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Llama-3.2-3B-Instruct-Q4_K_M.gguf",
    "messages": [{"role": "user", "content": "Count from 1 to 5."}],
    "stream": true
  }'

🐍 Python SDK Integration

from termux_llamacpp import LlamaRuntime

# 1. Initialize runtime
runtime = LlamaRuntime()

# 2. Start managed supervisor server
server = runtime.serve(
    model="~/.shitty_phone_ai/models/Llama-3.2-3B-Instruct-Q4_K_M.gguf",
    host="127.0.0.1",
    port=8080,
    ctx_size=2048,
    threads=4
)

print(f"Server active at: {server.endpoint}")

Interoperability with termux-aichain

from termux_aichain import LocalAgent

agent = LocalAgent.create(
    mode="connect",
    endpoint="http://127.0.0.1:8080",
    model="Llama-3.2-3B-Instruct-Q4_K_M.gguf"
)

response = agent.run("Hello from termux-aichain!")
print(response)

📊 Real-Device On-Device Benchmarks

Tested on Samsung Galaxy S20 / Android 15 (ARM64 Snapdragon 865) via Termux:

Model Quantization Warmup (mmap) TTFT (Time To First Token) Generation Speed Protocol
Qwen 2.5 1.5B Instruct Q4_K_M (1.1 GB) 3.0s 0.23s (230ms) 13.11 ~ 14.12 tokens/sec OpenAI SSE
Llama 3.2 1B Instruct Q4_K_M (800 MB) 2.2s 0.18s (180ms) 16.50 ~ 18.20 tokens/sec OpenAI SSE
Llama 3.2 3B Instruct Q4_K_M (2.0 GB) 4.5s 0.45s (450ms) 7.80 ~ 8.90 tokens/sec OpenAI SSE

🔒 Supply Chain Security & Architecture

termux-llamacpp enforces strict supply-chain security protocols:

graph TD
    A["termux-llama CLI / SDK"] --> B{"Binary Trust Verifier"}
    B -->|"Signed Release"| C["Ed25519 Public Key Manifest Verification"]
    B -->|"Local Build Receipt"| D["Pinned Commit SHA + Local SHA-256 Validation"]
    B -->|"Unknown / Tampered"| E["Fail-Closed Halt (Exit 1)"]
    C --> F["Atomic Directory Swap (~/.termux-llama)"]
    D --> F
    F --> G["Loopback Reverse Proxy Supervisor (:8080)"]
    G --> H["Native llama-server (:18080)"]
  1. Anti-Downgrade Trust Hierarchy: Signed release manifests cannot be downgraded to unverified local receipts.
  2. Symlink Defense: Rejects symlinks for binary paths, model weights, trust root public keys, and key revocation files to prevent TOCTOU attacks.
  3. Loopback Isolation: Native backend binds strictly to 127.0.0.1:18080 with loopback CORS filtering to block unauthorized cross-origin requests.
  4. Atomic Installation & Rollback: All installs stage to .new and swap cleanly, preserving .previous for automatic rollback upon verification failure.

📝 Release Notes

[v1.2.0] — 2026-09-01

  • Unified AMEVA Vulkan HAL Integration: Seamless binding with ameva-vulkan-runtime>=1.1.0 supporting Android Bionic Vulkan ICD (/system/lib64/libvulkan.so).
  • Strict 3-Tier Execution Mode:
    • --device vulkan: Strict GPU shader acceleration with Fail-Fast protection (no silent CPU fallback).
    • --device auto: Intelligent hardware discovery with transparent CPU NEON fallback.
    • --device cpu: Zero-overhead direct CPU NEON forward pass bypassing the Vulkan loader.
  • Dynamic Topology Optimization: Automatic big-core cluster detection (-t 4) on ARM octa-core processors (e.g., Exynos 1380, Snapdragon).
  • Galaxy A35 Real-Device Certification: Fully verified end-to-end token generation across CLI, Python SDK, Node.js npm, and OpenAI-compatible server daemon.
  • Dual-Ecosystem Availability: Synchronously published to PyPI (pip install termux-llamacpp) and npm (npm install -g termux-llamacpp).

📄 License

This project is licensed under the Apache-2.0 License. Third-party component notices and licenses are documented in LICENSES/.

Metadata

Release files for termux-llamacpp 1.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for termux-llamacpp 1.2.1
File Size Uploaded
termux_llamacpp-1.2.1.tar.gz 250.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for termux-llamacpp 1.2.1
File Interpreter ABI Platform
termux_llamacpp-1.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 398.5 kB

Release files / termux_llamacpp-1.2.1.tar.gz

Download URL termux_llamacpp-1.2.1.tar.gz
Size 250.5 kB
Tags Source
SHA-256 checksum
How to use checksums
f195fbf9ee8dbdb4a3ec3a9f2ee5eb9655f6ac20ba1b7b07c2d25b8f41418bf9
BLAKE2b-256 checksum
How to use checksums
99690e692bf2a4665a81a4951209b87a4a815b4c87b39ec3d8ae618716417d3f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.0

Release files / termux_llamacpp-1.2.1-py3-none-any.whl

Download URL termux_llamacpp-1.2.1-py3-none-any.whl
Size 148.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
72a82aec2a32310edce95918c62e517415ae0b955448a454a928cbf04fd503e3
BLAKE2b-256 checksum
How to use checksums
d1c42afda10db938915031688784f00b5f87fa33bf06bbaa506688a381793466
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.0
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page