Skip to main content

Termux-Train (v2.0.0)

PyPI Python npm npm downloads License

Unified Multimodal On-Device Deep Learning & LoRA Training Framework for Android Termux with 6-Modality Adapters (LLM, Diffusion, VLM, STT, TTS, BitNet), GPU Slicing, and 44GB Disaggregated Cluster Virtual RAM Pooling


Executive Summary & Architecture Overview

Termux-Train v2.0.0 is an industrial-grade, zero-dependency deep learning framework engineered specifically for constrained ARM64 Bionic environments (Android Termux, mobile Linux, embedded edge devices). By combining a C-vectorized Directed Acyclic Graph (DAG) Autograd runtime with mobile hardware acceleration (Vulkan Compute, OpenCL, ARM64 NEON), Termux-Train enables full backward gradient propagation, parameter-efficient fine-tuning (LoRA / DoRA), and reinforcement learning directly on mobile devices.

+---------------------------------------------------------------------------------------------------+
|                                        TERMUX-TRAIN v2.0.0                                        |
|                           Unified Multimodal On-Device Deep Learning Engine                       |
+---------------------------------------------------------------------------------------------------+
                                                  |
        +-----------------------------------------+-----------------------------------------+
        |                                                                                   |
        v                                                                                   v
+-------------------------------+                                           +-------------------------------+
|     6-Modality Adapters       |                                           |     Memory & Scaling Subsys   |
|  - LLM / DoRA / TinyLM        |                                           |  - Single-Device GPU Slicing  |
|  - Image Diffusion LoRA       |                                           |    * Vocab Slicing            |
|  - Vision Multimodal (VLM)    |                                           |    * Chunked Layer Dispatch   |
|  - Whisper STT Acoustic LoRA  |                                           |    * Layer Streaming          |
|  - Speech TTS Style LoRA      |                                           |  - 44GB Cluster Virtual RAM   |
|  - BitNet 1.58-bit Ternary    |                                           |  - 300MB Guard-Band Memory    |
+-------------------------------+                                           +-------------------------------+
        |                                                                                   |
        +-----------------------------------------+-----------------------------------------+
                                                  |
                                                  v
+---------------------------------------------------------------------------------------------------+
|                             Target-Aware Runtime Export & Compatibility                           |
|    ComfyUI / Diffusers   |   llama.cpp / GGUF   |   termux-vision   |   termux-stt   |   BitNet  |
+---------------------------------------------------------------------------------------------------+

Core Architectural Pillars

  1. Zero External Heavy Dependencies: Executes autonomously without PyTorch, LibTorch, CMake, or heavy BLAS runtimes. Pure Python DAG autograd paired with optional C-accelerator ABI (libtermux_train_accel.so) or Vulkan compute shaders.
  2. 6-Modality Target-Aware Adapters: Generates standard SafeTensors weights and companion _config.json manifests natively recognized by major desktop and mobile inference runtimes without conversion tooling.
  3. Single-Device GPU Slicing: Bypasses mobile Unified Memory Architecture (UMA) spikes and watchdog timeouts on ARM Mali and Qualcomm Adreno GPUs through granular vocabulary chunking and layer streaming.
  4. 44GB Disaggregated Cluster Virtual RAM Pooling: Binds physical memory across heterogeneous mobile nodes (e.g., Galaxy S25, S21, S20, A53, A35) into an aggregated pipeline training session protected by a 300MB memory Guard-Band against the Android Low Memory Killer (LMK).
  5. 100% Dual-Engine Parity: Complete feature and interface symmetry across Python CLI (termux-train), Node.js Global CLI (termux-train), Python Library (import termux_train), and Node.js SDK (import { ... } from 'termux-train').

Installation

1. Via Python Package Index (PyPI)

pip install termux-train

2. Via Node Package Manager (npm)

npm install -g termux-train

3. Standalone Bootstrap Script (Termux One-Liner)

pkg update && pkg install -y python nodejs clang
bash <(curl -sSL https://raw.githubusercontent.com/uno-km/termux-train/main/install.sh)

6-Modality Training Pipelines & CLI Reference

1. Image Diffusion LoRA Training (diffusion-train)

Trains latent diffusion cross-attention LoRA layers. Integrates a deterministic DDPM noise scheduler, sinusoidal timestep MLP, and SHA-256 disk-based VAE latent caching to eliminate redundant image decoding across epochs. Outputs ComfyUI- and Diffusers-compatible SafeTensors adapters.

termux-train diffusion-train \
  --image-dir ./training_images \
  --prompt "high quality product photograph, 4k" \
  --output ./adapter_diffusion.safetensors \
  --resolution 512 \
  --epochs 5 \
  --batch-size 1 \
  --lr 0.0001 \
  --rank 8 \
  --alpha 16.0 \
  --backend vulkan

2. Vision Multimodal VLM LoRA Training (vision-train)

Trains multimodal projection adapters and visual cross-attention layers. Operates on image-text pairs formatted in LLaVA or Qwen2-VL visual question answering schemas. Directly bridges into the termux-vision mobile inference runtime.

termux-train vision-train \
  --data ./vlm_dataset.jsonl \
  --output ./adapter_vision.safetensors \
  --epochs 3 \
  --batch-size 2 \
  --lr 0.0002 \
  --rank 8 \
  --alpha 16.0 \
  --backend vulkan

3. Speech-to-Text Whisper LoRA Training (stt-train)

Performs acoustic adaptation on Whisper encoder-decoder attention blocks. Features automated Mel filterbank spectrogram extraction and persistent caching. Natively consumable by termux-stt.

termux-train stt-train \
  --data ./speech_corpus.jsonl \
  --output ./adapter_stt.safetensors \
  --epochs 4 \
  --batch-size 4 \
  --lr 0.0005 \
  --rank 4 \
  --backend vulkan

4. Text-to-Speech Style Adaptation LoRA Training (tts-train)

Fine-tunes speaker embedding and acoustic style parameters for neural speech synthesis. Generates low-rank voice profile adapters directly compatible with termux-tts.

termux-train tts-train \
  --data ./voice_samples.jsonl \
  --output ./adapter_tts.safetensors \
  --epochs 5 \
  --batch-size 2 \
  --lr 0.0003 \
  --rank 4 \
  --backend cpu

5. Large Language Model PEFT (train / peft)

Trains LoRA / DoRA weight matrices over multi-head attention and feed-forward projections with RoPE positional encodings. Supports single-device GPU slicing flags to eliminate out-of-memory crashes on mobile GPUs. Exports standard GGUF PEFT adapters for llama.cpp.

termux-train train \
  --model tiny-transformer \
  --data ./training_corpus.txt \
  --output ./adapter_llm.safetensors \
  --epochs 3 \
  --lr 0.0002 \
  --lora-rank 8 \
  --lora-alpha 16.0 \
  --vocab-slice 4096 \
  --chunk-layers 2 \
  --stream-layers \
  --backend vulkan

6. BitNet 1.58-bit Ternary Quantization-Aware Training

Applies dynamic activation scaling and ternary weight quantization ${-1, 0, +1}$ during forward and backward passes. Directly exports quantized kernels compatible with bitnet.cpp and termux-bitnet.

termux-train train \
  --model bitnet \
  --quantize-bits 1.58 \
  --data ./corpus.txt \
  --output ./bitnet_adapter.safetensors \
  --backend opencl

Complete Parameter Reference Matrix

Parameter Flag Type Default Applicable Subcommands Engineering Semantics & Usage Description
--image-dir String Required diffusion-train Absolute or relative path to directory containing source training images (.png, .jpg).
--prompt String "" diffusion-train Conditioning text prompt applied across image diffusion training instances.
--resolution Integer 512 diffusion-train Input image dimension for latent VAE spatial encoding ($H=W$).
--data String Required vision-train, stt-train, tts-train, train, rl Path to dataset file (.jsonl, .txt, .csv).
--output String Required All training subcommands Destination filepath for exported .safetensors adapter file.
--epochs Integer 3 All training subcommands Total complete passes over the dataset.
--batch-size Integer 1 All training subcommands Micro-batch size per training step.
--lr Float 0.0001 All training subcommands Optimization learning rate for AdamW / SGD gradient update.
--rank / --lora-rank Integer 8 All training subcommands Inner rank dimension $r$ for LoRA projection matrices $A \in \mathbb{R}^{r \times d_{in}}$ and $B \in \mathbb{R}^{d_{out} \times r}$.
--alpha / --lora-alpha Float 16.0 All training subcommands Scaling hyperparameter $\alpha$ for LoRA weight contribution ($\Delta W = \frac{\alpha}{r} BA$).
--backend String auto All subcommands Compute acceleration backend: auto, vulkan, opencl, cpu.
--vocab-slice Integer 0 (off) train, peft Granular vocabulary partition size to prevent GPU out-of-memory spikes during softmax cross-entropy.
--chunk-layers Integer 0 (off) train, peft Number of sequential transformer layers evaluated per GPU dispatch to prevent OS watchdog timeout resets.
--stream-layers Boolean false train, peft Dynamically streams inactive layer weights between system RAM and GPU buffer memory.
--virtual-ram-pool String "" train, diffusion-train, peft Cluster connection endpoint (host:port) binding local training session to aggregated virtual RAM pool.
--guard-band Integer 300 cluster-worker, train Reserved unallocated system RAM buffer (in MB) preventing Android Low Memory Killer (LMK) aborts.
--method String grpo rl Reinforcement learning optimization algorithm: grpo, dpo, ppo.
--group-size Integer 4 rl Generation group candidate size for Group Relative Policy Optimization (GRPO).
--beta Float 0.1 rl KL divergence regularization coefficient for DPO and GRPO reference policy stabilization.
--host String 0.0.0.0 cluster-worker, cluster-probe Network host address for disaggregated virtual RAM clustering daemon.
--port Integer 50052 cluster-worker, cluster-probe Network TCP port for cluster worker coordination.

Single-Device GPU Slicing Architecture

On mobile platforms, the GPU and CPU share a Unified Memory Architecture (UMA). Large batch matrix multiplications during autograd activation accumulation trigger sudden memory surges that prompt the Android system to issue SIGKILL signals. Furthermore, ARM Mali GPU driver stacks enforce a hardware watchdog timeout: any single compute command buffer executing longer than 2.0 seconds is terminated with VK_ERROR_DEVICE_LOST.

                        MOBILE UNIFIED RAM (LPDDR5)
+-------------------------------------------------------------------------+
| OS Kernel / Termux Base | Active Layer Buffer | Guard-Band Buffer (300MB)|
+-------------------------------------------------------------------------+
                                   ^
                                   |  Layer Streaming Swap
                                   v
                        VULKAN / OPENCL GPU ENGINE
+-------------------------------------------------------------------------+
| [Vocab Chunk 1] -> [Vocab Chunk 2] | [Layer N] -> [Layer N+1] (Chunked) |
+-------------------------------------------------------------------------+

To neutralize these constraints:

  • Vocab Slicing (--vocab-slice 4096): Partitions vocabulary projection matrices into slices of 4,096 tokens, performing sequential cross-entropy reductions with minimal temporary tensor overhead.
  • Chunked Layer Dispatch (--chunk-layers 2): Segments multi-layer transformer forward-backward passes into 2-layer dispatches, returning control to the Vulkan queue to reset watchdog timers.
  • Layer Streaming (--stream-layers): Maintains only active weight layers inside device GPU memory buffers, paging inactive weights into host RAM.

44GB Disaggregated Cluster Virtual RAM Pooling (Powered by AMEVA Cluster)

When training deep learning models that exceed physical mobile device memory (e.g. 7B+ parameter models, high-resolution diffusion unets, or multimodal projectors), Termux-Train binds heterogeneous smartphone clusters into a unified virtual memory pool. Idle mobile devices contribute unallocated LPDDR RAM via lightweight worker daemons managed by AMEVA Cluster.

+-------------------------------------------------------------------------+
|              COORDINATOR / PRIMARY NODE (Galaxy S25 - 12GB)             |
|                Directs Training Pipeline & Weight Updates               |
+-------------------------------------------------------------------------+
           |                           |                           |
  gRPC/HTTP RPC               gRPC/HTTP RPC               gRPC/HTTP RPC
  (MasterTunnel)              (MasterTunnel)              (MasterTunnel)
           v                           v                           v
+--------------------+      +--------------------+      +--------------------+
| WORKER 1 (S21)     |      | WORKER 2 (S20)     |      | WORKER 3 (A53/A35) |
| RAM: 8GB LPDDR5    |      | RAM: 12GB LPDDR5   |      | RAM: 12GB LPDDR4X  |
| Dedicated: 6.8GB   |      | Dedicated: 10.5GB  |      | Dedicated: 9.2GB   |
+--------------------+      +--------------------+      +--------------------+
           \                           |                           /
            +--------------------------+--------------------------+
                                       |
                                       v
                       TOTAL AGGREGATED VIRTUAL POOL: ~44GB

Complete Cluster Step-by-Step Operation Manual

Step 1: Network & Environment Setup

Ensure all participating smartphones are connected to the same local Wi-Fi network, router, or tethered hotspot. Verify TCP port 50052 is accessible between nodes:

pkg update && pkg install -y python clang netcat-openbsd
pip install --upgrade ameva-cluster termux-train

Step 2: Launching Cluster Workers on Contributor Nodes

On each memory-contributor smartphone (e.g., Galaxy S21, S20, A53), start the worker daemon. The --guard-band 300 flag reserves 300MB of unallocated system RAM at all times to prevent the Android Low Memory Killer (LMK) from terminating processes:

# Option A: Directly through ameva-cluster
ameva-cluster worker --port 50052 --guard-band 300

# Option B: Through termux-train CLI bridge
termux-train cluster-worker --port 50052 --guard-band 300

Step 3: Probing and Discovering Fleet Nodes

On your primary coordinator smartphone (e.g., Galaxy S25), run a cluster probe to inspect node reachability, available LPDDR memory headroom, and network round-trip ping times:

termux-train cluster-probe --fleet 192.168.1.101:50052,192.168.1.102:50052,192.168.1.103:50052

Step 4: Initiating Distributed Virtual RAM Pipeline Training

Train across the aggregated 44GB virtual RAM pool with zero changes to model architecture:

# Distributed LLM LoRA Training with GPU Slicing
termux-train train \
  --model tiny-transformer \
  --data ./large_corpus.txt \
  --output ./cluster_adapter.safetensors \
  --virtual-ram-pool 192.168.1.101:50052,192.168.1.102:50052,192.168.1.103:50052 \
  --epochs 5 \
  --vocab-slice 4096 \
  --chunk-layers 2 \
  --backend vulkan

# Distributed Image Diffusion LoRA Training
termux-train diffusion-train \
  --image-dir ./dataset \
  --output ./diff_cluster_adapter.safetensors \
  --virtual-ram-pool 192.168.1.101:50052,192.168.1.102:50052 \
  --epochs 5 \
  --backend vulkan

Step 5: MasterTunnel Cryptographic Loopback Security

When passing sensitive weights and intermediate activations across untrusted Wi-Fi subnets, termux-train and ameva-cluster automatically spawn an authenticated MasterTunnel that wraps all inter-device gRPC/HTTP RPC traffic with an ephemeral SHA-256 cryptographic handshake.


Programmatic SDK Usage

Python SDK

import termux_train as tt

# High-level training interface
trainer = tt.TermuxTrainer(
    model_type="lora",
    lora_rank=8,
    lora_alpha=16.0,
    backend="vulkan",
    vocab_slice=4096,
    chunk_layers=2
)

metrics = trainer.fit(
    data_path="./corpus.txt",
    epochs=3,
    learning_rate=0.0002,
    output_path="./model_lora.safetensors"
)
print(f"Final training loss: {metrics['final_loss']:.4f}")

Node.js / TypeScript SDK

import { trainDiffusion, trainVision, trainSTT, TermuxTrainer } from 'termux-train';

// Asynchronous image diffusion fine-tuning
const result = await trainDiffusion({
  imageDir: './training_images',
  output: './adapter_diffusion.safetensors',
  prompt: 'photorealistic macro nature photography',
  resolution: 512,
  epochs: 5,
  backend: 'vulkan'
});
console.log(`Diffusion training completed with status: ${result.status}`);

Target Runtime Compatibility & Adapter Hub

All adapters generated by Termux-Train are zero-conversion artifacts accompanied by a standardized _config.json sidecar.

Target Runtime Adapter Format Companion Sidecar Validation Status
ComfyUI / Diffusers .safetensors (UNet CrossAttention) adapter_diffusion_config.json Validated by automated checks
llama.cpp / GGUF .safetensors / GGUF LoRA adapter_llm_config.json Validated by automated checks
termux-vision .safetensors (Multimodal Projector) adapter_vision_config.json Validated by automated checks
termux-stt .safetensors (Whisper Attention) adapter_stt_config.json Validated by automated checks
termux-tts .safetensors (Speaker Style) adapter_tts_config.json Validated by automated checks
bitnet.cpp .safetensors (Ternary 1.58-bit) bitnet_config.json Validated by automated checks

Official Documentation & Foundation Resources


Open-Source Attribution & License

Licensed under the Apache-2.0 License. Copyright (c) 2026 Eunho Kim (@uno-km) and the AMEVA Open-Source Foundation.

Metadata

Release files for termux-train 2.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for termux-train 2.0.1
File Size Uploaded
termux_train-2.0.1.tar.gz 230.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for termux-train 2.0.1
File Interpreter ABI Platform
termux_train-2.0.1-py3-none-any.whl Python 3 none any Details

Total release size: 417.1 kB

Release files / termux_train-2.0.1.tar.gz

Download URL termux_train-2.0.1.tar.gz
Size 230.5 kB
Tags Source
SHA-256 checksum
How to use checksums
92f23e85b4fe5935fabc4d12abb0614f3167e197208062196d663bac94599261
BLAKE2b-256 checksum
How to use checksums
7da787bcb7af01dfea7d99720ca9823110e79641cf20b279335144d2120e383a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.0

Release files / termux_train-2.0.1-py3-none-any.whl

Download URL termux_train-2.0.1-py3-none-any.whl
Size 186.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
688b662ec6283297d172d1896d4e735d52ee56c07a9a579d135502f7e4b278d0
BLAKE2b-256 checksum
How to use checksums
716f4393e3b1a7f2f2311ea3160e410336a14bc246a874de2542942477f3be2a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.0

Release history Release notifications | RSS feed

This release

2.0.1 This release

2 release files

2.0.0

2 release files

1.1.9

2 release files

1.1.8

2 release files

1.1.7

2 release files

1.1.6

2 release files

1.1.5

2 release files

1.1.3

2 release files

1.1.2

2 release files

1.1.1

2 release files

1.1.0

2 release files

1.0.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page