Skip to main content

Termux-LlamaCpp

PyPI Python npm npm downloads License

디바이스 리소스를 활용한 Android Termux ARM64 전용 GGUF LLM 런타임, 모델 매니저 및 OpenAI 호환 REST/SSE 서버
Production-Grade GGUF LLM Runtime Utilizing Device Resources, Model Manager & OpenAI Server for Android ARM64


Architecture & Overview

검증된 Android Bionic ARM64 네이티브 바이너리와 필수 공유 라이브러리, 디바이스 리소스 최적화 연동을 통해 제로 컴파일 로컬 실행과 안정적인 OpenAI 호환 REST/SSE 스트리밍 서버를 제공합니다.

Ships verified Android ARM64 native binaries with bundled shared libraries and device resource optimization, enabling instant zero-compilation local inference and a robust OpenAI-compatible REST/SSE server.


Installation & Quickstart

Python (PyPI)

pip install termux-llamacpp
from termux_llamacpp import LlamaRuntime, RuntimeConfig

# 1. Initialize Runtime with Vulkan GPU Acceleration
config = RuntimeConfig(
    model_path="models/Llama-3.2-1B-Instruct-Q4_K_M.gguf",
    device="auto",
    threads=4,
    context_size=2048
)
runtime = LlamaRuntime(config)

# 2. Synchronous or Streaming Text Generation
output = runtime.generate("한국의 사계절 중 가을의 매력에 대해 설명해줘.", max_tokens=256)
print(output.text)
print(f"Speed: {output.metrics.eval_tokens_per_sec:.2f} t/s (Prompt: {output.metrics.prompt_tokens_per_sec:.2f} t/s)")

Node.js / TypeScript (npm)

npm install termux-llamacpp
import { LlamaRuntime } from "termux-llamacpp";

// 1. Initialize ESM Runtime
const runtime = new LlamaRuntime({
  modelPath: "models/Llama-3.2-1B-Instruct-Q4_K_M.gguf",
  device: "auto",
  threads: 4
});

// 2. Generate Completion
const result = await runtime.generate({
  prompt: "Explain the architectural philosophy of edge computing.",
  maxTokens: 256
});
console.log("Response:", result.text);
console.log(`Generation Speed: ${result.metrics.evalTokensPerSec} t/s`);

Distributed Clustering & Memory Pooling (AMEVA-Cluster)

Termux-LlamaCpp natively integrates with AMEVA-Cluster (pip install ameva-cluster) for distributed RAM pooling across heterogeneous mobile fleets.

1. Install Cluster Runtime

pip install ameva-cluster
# or Node.js:
npm install @ameva/cluster

2. Launch Worker Node on Remote Phone

# On worker device (e.g. Galaxy A53):
ameva-cluster worker --port 50052

3. Distributed Inference via Master Node

# Master CLI with automatic memory guardband balancing:
termux-llama run -m Qwen2.5-7B-Instruct-Q4_K_M.gguf \
  --rpc 192.0.2.10:50052,192.0.2.11:50052 \
  -ts auto \
  -p "Explain quantum computing in simple terms."
from termux_llamacpp import LlamaRuntime, RuntimeConfig

# Python SDK Distributed Memory Pooling
config = RuntimeConfig(
    model_path="models/Qwen2.5-7B-Instruct-Q4_K_M.gguf",
    cluster_rpc_servers="192.0.2.10:50052,192.0.2.11:50052",
    tensor_split="auto"
)
runtime = LlamaRuntime(config)
res = runtime.generate("Distributed clustering active across multiple mobile devices.")
print(res.text)
  • --rpc / cluster_rpc_servers: Comma-separated RPC worker addresses.
  • -ts auto: Dynamic tensor-split automatically balanced by hardware VRAM/RAM capacity.
  • Protocol Guard: Cryptographically authenticated via ClusterGuardProxy with HMAC-SHA256 handshake.


License

Licensed under the Apache-2.0 License. Copyright (c) 2026 Eunho Kim (@uno-km).

Metadata

Release files for termux-llamacpp 1.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for termux-llamacpp 1.4.0
File Size Uploaded
termux_llamacpp-1.4.0.tar.gz 301.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for termux-llamacpp 1.4.0
File Interpreter ABI Platform
termux_llamacpp-1.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 500.3 kB

Release files / termux_llamacpp-1.4.0.tar.gz

Download URL termux_llamacpp-1.4.0.tar.gz
Size 301.7 kB
Tags Source
SHA-256 checksum
How to use checksums
e5012e269fbf5b4e20524c9b1e6134c450966ce3ba427e82f3d788fb754b704b
BLAKE2b-256 checksum
How to use checksums
092f97fc49701f8ff64cb41849744e4421ce4c281dff17d794593faa0d8c015f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.0

Release files / termux_llamacpp-1.4.0-py3-none-any.whl

Download URL termux_llamacpp-1.4.0-py3-none-any.whl
Size 198.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
beaea6ef878a6871ce85bd62809b06a0426deebf0bd48074adeffcc5d521509f
BLAKE2b-256 checksum
How to use checksums
cbf64518c8b97844d70c81c026e4101f0916456cc1c32fcce6c3e6d6de522ae5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.0
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page