Termux-LlamaCpp
디바이스 리소스를 활용한 Android Termux ARM64 전용 GGUF LLM 런타임, 모델 매니저 및 OpenAI 호환 REST/SSE 서버
Production-Grade GGUF LLM Runtime Utilizing Device Resources, Model Manager & OpenAI Server for Android ARM64
Architecture & Overview
검증된 Android Bionic ARM64 네이티브 바이너리와 필수 공유 라이브러리, 디바이스 리소스 최적화 연동을 통해 제로 컴파일 로컬 실행과 안정적인 OpenAI 호환 REST/SSE 스트리밍 서버를 제공합니다.
Ships verified Android ARM64 native binaries with bundled shared libraries and device resource optimization, enabling instant zero-compilation local inference and a robust OpenAI-compatible REST/SSE server.
Installation & Quickstart
Python (PyPI)
pip install termux-llamacpp
from termux_llamacpp import LlamaRuntime, RuntimeConfig
# 1. Initialize Runtime with Vulkan GPU Acceleration
config = RuntimeConfig(
model_path="models/Llama-3.2-1B-Instruct-Q4_K_M.gguf",
device="auto",
threads=4,
context_size=2048
)
runtime = LlamaRuntime(config)
# 2. Synchronous or Streaming Text Generation
output = runtime.generate("한국의 사계절 중 가을의 매력에 대해 설명해줘.", max_tokens=256)
print(output.text)
print(f"Speed: {output.metrics.eval_tokens_per_sec:.2f} t/s (Prompt: {output.metrics.prompt_tokens_per_sec:.2f} t/s)")
Node.js / TypeScript (npm)
npm install termux-llamacpp
import { LlamaRuntime } from "termux-llamacpp";
// 1. Initialize ESM Runtime
const runtime = new LlamaRuntime({
modelPath: "models/Llama-3.2-1B-Instruct-Q4_K_M.gguf",
device: "auto",
threads: 4
});
// 2. Generate Completion
const result = await runtime.generate({
prompt: "Explain the architectural philosophy of edge computing.",
maxTokens: 256
});
console.log("Response:", result.text);
console.log(`Generation Speed: ${result.metrics.evalTokensPerSec} t/s`);
Distributed Clustering & Memory Pooling (AMEVA-Cluster)
Termux-LlamaCpp natively integrates with AMEVA-Cluster (pip install ameva-cluster) for distributed RAM pooling across heterogeneous mobile fleets.
1. Install Cluster Runtime
pip install ameva-cluster
# or Node.js:
npm install @ameva/cluster
2. Launch Worker Node on Remote Phone
# On worker device (e.g. Galaxy A53):
ameva-cluster worker --port 50052
3. Distributed Inference via Master Node
# Master CLI with automatic memory guardband balancing:
termux-llama run -m Qwen2.5-7B-Instruct-Q4_K_M.gguf \
--rpc 192.0.2.10:50052,192.0.2.11:50052 \
-ts auto \
-p "Explain quantum computing in simple terms."
from termux_llamacpp import LlamaRuntime, RuntimeConfig
# Python SDK Distributed Memory Pooling
config = RuntimeConfig(
model_path="models/Qwen2.5-7B-Instruct-Q4_K_M.gguf",
cluster_rpc_servers="192.0.2.10:50052,192.0.2.11:50052",
tensor_split="auto"
)
runtime = LlamaRuntime(config)
res = runtime.generate("Distributed clustering active across multiple mobile devices.")
print(res.text)
--rpc / cluster_rpc_servers: Comma-separated RPC worker addresses.-ts auto: Dynamic tensor-split automatically balanced by hardware VRAM/RAM capacity.- Protocol Guard: Cryptographically authenticated via
ClusterGuardProxywith HMAC-SHA256 handshake.
- Official Architecture & API Reference
- Ecosystem Metrics & Registry Stats
- AMEVA Open-Source Foundation Portal
License
Licensed under the Apache-2.0 License. Copyright (c) 2026 Eunho Kim (@uno-km).
Metadata
Release files for termux-llamacpp 1.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| termux_llamacpp-1.4.0.tar.gz | 301.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| termux_llamacpp-1.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 500.3 kB
Release files / termux_llamacpp-1.4.0.tar.gz
| Download URL | termux_llamacpp-1.4.0.tar.gz |
|---|---|
| Size | 301.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e5012e269fbf5b4e20524c9b1e6134c450966ce3ba427e82f3d788fb754b704b
|
|
BLAKE2b-256 checksum How to use checksums |
092f97fc49701f8ff64cb41849744e4421ce4c281dff17d794593faa0d8c015f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.0
|
Release files / termux_llamacpp-1.4.0-py3-none-any.whl
| Download URL | termux_llamacpp-1.4.0-py3-none-any.whl |
|---|---|
| Size | 198.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
beaea6ef878a6871ce85bd62809b06a0426deebf0bd48074adeffcc5d521509f
|
|
BLAKE2b-256 checksum How to use checksums |
cbf64518c8b97844d70c81c026e4101f0916456cc1c32fcce6c3e6d6de522ae5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.0
|