AMEVA-Cluster
Symmetric Disaggregated Mobile RAM Pooling & On-Device AI Acceleration Runtime with Compute-Memory Decoupling
1. Executive Architecture & Threat Model
Executing frontier AI models (7B–70B parameters, 4B Flow Matching Diffusion, Multimodal Vision) on edge mobile smartphones inevitably triggers immediate Android Low Memory Killer (LMK) SIGKILL terminations and extreme battery thermal throttling.
Conventional distributed inference frameworks force every connected node to execute computational passes, causing lower-tier edge devices to overheat, desynchronize, and crash.
AMEVA-Cluster resolves this architectural limitation through Compute-Memory Decoupling:
- Symmetric Single-Binary Architecture: Both
masterandworkerroles run from a unified binary (ameva-cluster), dynamically configurable via CLI flags. - Pure RAM Pooling: Worker nodes execute zero GPU/NPU compute passes, serving exclusively as remote high-bandwidth LPDDR memory pools to eliminate worker thermal load.
- Android LMK Safety Guard-Band (300MB): Automatically deducts safety buffers from active RAM metrics (
MemAvailable) to prevent Android kernel out-of-memory kills. - Protocol Guard Proxy & MasterTunnel: Intercepts all incoming connections with a 15-byte magic preamble and 60-second time-windowed HMAC-SHA256 mutual authentication. Drops unauthorized RPC probes in
< 0.001s(HTTP 403 Forbidden). - Storage-Assisted Hybrid Failover: Automatically recovers from worker dropouts by streaming remaining tensor weights directly from local high-speed UFS 3.1 flash memory via kernel
mmap.
[ Master Node (Flagship GPU) ]
│
├── Local Compute: Flagship Adreno / Mali GPU (100% of forward passes)
│
├── MasterTunnel (Loopback: 127.0.0.1:5105X)
│ └── Transparently injects HMAC-SHA256 handshake token
│
├── USB 3.0 ADB / Tailscale WireGuard Transport (Encrypted)
│ │
│ ├──▶ [ Worker 1 (Galaxy S20 - 12GB LPDDR5) ] ──▶ ClusterGuardProxy ──▶ Loopback rpc-server
│ ├──▶ [ Worker 2 (Galaxy A35 - 8GB LPDDR4X) ] ──▶ ClusterGuardProxy ──▶ Loopback rpc-server
│ └──▶ [ Worker 3 (Galaxy A53 - 6GB LPDDR4X) ] ──▶ ClusterGuardProxy ──▶ Loopback rpc-server
│
└── Storage Failover: Local UFS 3.1 mmap fallback upon node disconnect
2. Installation & Quickstart
Python (PyPI)
pip install ameva-cluster
Node.js / TypeScript (npm)
npm install -g @ameva/cluster
Termux One-Liner (Android ARM64)
pkg update && pkg install python nodejs-lts git -y
pip install ameva-cluster
3. Command Line Interface (CLI) Reference
1) Starting a Worker Daemon (Edge Phone)
Launch the worker daemon with automatic Termux WakeLock acquisition and Protocol Guard protection:
# Protected worker on public port 50052 (isolates rpc-server on loopback 50055)
ameva-cluster worker --port 50052 --guard-band 300
Key Flags:
--port <int>: External listening port (default:50052).--isolated-port <int>: Internal isolated raw RPC loopback port (default:50055).--guard-band <int>: Memory safety guardband in MB (default:300).--no-guard: Disables HMAC authentication proxy (raw debugging only).
2) Running the Master Orchestrator (Workstation / Host Phone)
Coordinate distributed inference across multiple edge smartphones:
# Auto tensor sharding across Galaxy A35 and Galaxy A53
ameva-cluster master \
-m models/Qwen2.5-7B-Instruct-Q4_K_M.gguf \
--workers 100.106.251.21:50052,100.77.47.37:50052 \
-ts auto \
-p "Explain the physics of gravitational lensing."
Key Flags:
-m, --model <path>: Path to target model (.gguf,.safetensors,.bin).--workers <csv>: Comma-separated list of worker endpoints (ip:port).-ts, --tensor-split <ratio>: Sharding ratio across master and workers (e.g.50,50orauto).--engine <name>: Backend engine:llamacpp,diffusion,vision,stt,tts,bitnet,train.-p, --prompt <text>: Inference prompt.
3) Pre-Flight Diagnostics & Fleet Probing
Inspect network round-trip time (RTT), packet jitter, and compute node availability:
ameva-cluster probe --fleet 100.106.251.21:50052,100.77.47.37:50052
4. 6-Modality Distributed Integration
AMEVA-Cluster natively interconnects with the six official on-device AI modality runtimes:
1) Large Language Models: termux-llamacpp
termux-llama run -m Qwen2.5-7B-Instruct-Q4_K_M.gguf \
--rpc 100.106.251.21:50052,100.77.47.37:50052 \
--tensor-split auto \
-p "Synthesize distributed systems theory."
2) Image Diffusion: termux-diffusion
termux-diffusion generate \
--model z_image_turbo-q4.gguf \
--vae taef1.gguf \
--prompt "Hyper-detailed cybernetic macro lens photograph" \
--rpc 100.77.47.37:50052 \
--tensor-split 50,50 \
--steps 8
3) Multimodal Vision: termux-vision
termux-vision analyze \
--model Qwen2.5-VL-7B-Instruct-Q4_K_M.gguf \
--mmproj mmproj-Qwen2.5-VL-7B-Instruct-f16.gguf \
--image telemetry.png \
--rpc 100.106.251.21:50052 \
--prompt "Identify all memory leaks in this timeline graph."
4) Speech-to-Text: termux-stt
termux-stt transcribe conference_call.wav \
--model ggml-large-v3-turbo.bin \
--rpc 100.106.251.21:50052 \
--tensor-split 50,50
5) Text-to-Speech: termux-tts
termux-tts synthesize "Distributed edge clustering operational." \
--voice en_us_neutral \
--rpc 100.77.47.37:50052 \
--output alert.wav
6) 1-Bit LLM & LoRA Training: termux-bitnet & termux-train
# BitNet 1.58-bit ternary inference
termux-bitnet run -m bitnet_b1_58-3B.gguf --rpc 100.106.251.21:50052
# Distributed LoRA fine-tuning
termux-train lora --base-model Qwen2.5-3B.gguf --dataset train.jsonl --rpc 100.106.251.21:50052
5. Software Development Kit (SDK) API
Python API
from ameva_cluster import (
AmevaCluster,
ClusterMaster,
ClusterWorker,
ClusterHardwareProbe,
ClusterSecurityVerifier,
ClusterConfig,
)
# 1. Probe local hardware and subtract 300MB safety buffer
mem = ClusterHardwareProbe.get_memory_info(guard_band_mb=300)
print(f"Usable RAM: {mem['usable_mb']}MB / Total: {mem['total_mb']}MB")
# 2. Acquire Termux system WakeLock
ClusterHardwareProbe.ensure_wakelock()
# 3. Execute disaggregated cluster inference
AmevaCluster.run_master(
model_path="Qwen2.5-7B-Instruct-Q4_K_M.gguf",
workers=["100.106.251.21:50052", "100.77.47.37:50052"],
tensor_split="40,30,30",
prompt="Explain quantum entanglement."
)
Node.js / TypeScript API
const { AmevaCluster } = require('@ameva/cluster');
// 1. Launch protected worker daemon
const worker = AmevaCluster.startWorker({
port: 50052,
guardBandMb: 300,
enableGuard: true
});
// 2. Launch master orchestrator
const master = AmevaCluster.runMaster({
modelPath: "Qwen2.5-7B-Instruct-Q4_K_M.gguf",
workers: ["100.106.251.21:50052", "100.77.47.37:50052"],
tensorSplit: "auto",
prompt: "Demonstrate distributed Node.js tensor coordination."
});
master.stdout.on('data', (chunk) => process.stdout.write(chunk));
6. Security Gating & License Verification (Fail-Fast E403)
Multi-device distributed clustering is an exclusive capability governed by the AMEVA-Cluster protocol.
When standalone execution is attempted without official licensing or package presence, the runtime halts with code E403_CLUSTER_LICENSE_REQUIRED:
================================================================================
[AMEVA-CLUSTER] CLUSTER LICENSE REQUIRED (E403)
================================================================================
Multi-device distributed clustering is an exclusive capability of 'ameva-cluster'.
Standalone distributed execution without the official AMEVA-Cluster package is prohibited.
================================================================================
Community Tier Inclusion
The official package distribution comes pre-configured with the default community authorization key (AMEVA-COMMUNITY-CLUSTER-AUTH-TOKEN-2026). Developers can immediately link consumer phones without manual credential injection, while external unauthorized clients are blocked at the network barrier.
7. Official Documentation & Portal
- Web Documentation: https://uno-km.vercel.app/lib/cluster/
- Guard Protocol Specification: https://uno-km.vercel.app/lib/cluster/guard-protocol.html
- 6-Modality Sharding Guide: https://uno-km.vercel.app/lib/cluster/modalities-guide.html
- Fleet Topology & Tuning: https://uno-km.vercel.app/lib/cluster/fleet-optimization.html
8. License
Licensed under the Apache-2.0 License. Copyright (c) 2026 Eunho Kim (@uno-km). All rights reserved.
Metadata
Release files for ameva-cluster 1.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ameva_cluster-1.0.0.tar.gz | 46.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ameva_cluster-1.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 89.1 kB
Release files / ameva_cluster-1.0.0.tar.gz
| Download URL | ameva_cluster-1.0.0.tar.gz |
|---|---|
| Size | 46.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
10ecce5d8f46bbe2e2752739c24b68ac84c4a3531c2ca33984fd93483e116c1b
|
|
BLAKE2b-256 checksum How to use checksums |
9b376c29384af9c97e8ba47346d6964cf96841b7a200b1e6006f427f36ccee89
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.0
|
Release files / ameva_cluster-1.0.0-py3-none-any.whl
| Download URL | ameva_cluster-1.0.0-py3-none-any.whl |
|---|---|
| Size | 42.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
da9139464cdc61f4aeeeebf70ad66ccb917ac51f4d90d14ad2839291323be0b0
|
|
BLAKE2b-256 checksum How to use checksums |
2baa7c3d26903a8aa8347cfd65b0a4f56941630bd72b43d1a57d604ac82118ac
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.0
|