Skip to main content

Termux-STT

PyPI Python npm npm downloads License

디바이스 리소스를 활용한 통합 온디바이스 음성인식(STT) 및 순수 Python 128d X-Vector 화자 분리 프레임워크
Unified On-Device Speech-to-Text Utilizing Device Resources & Pure Python 128d X-Vector Speaker Diarization


Architecture & Overview

Whisper.cpp, Vosk, Sherpa-ONNX 3대 네이티브 엔진을 통합하고, 닫힌 형태 순수 Python 128차원 X-Vector 클러스터링을 결합하여 80MB 미만의 초경량 메모리로 100% 온디바이스 실시간 음성인식과 화자 분리를 구현합니다.

Integrates Whisper.cpp, Vosk, and Sherpa-ONNX with a closed-form pure-Python 128-dimensional X-Vector clustering algorithm that operates in under 80MB RAM with zero cloud egress.


Asymmetric Hybrid Architecture Benchmark (Galaxy S25 - Snapdragon 8 Elite)

Test Audio: JFK Inaugural Address (60.59s mono 16kHz WAV)
Silicon: Qualcomm Snapdragon 8 Elite (Adreno 830 GPU + Oryon CPU)

Model Pure CPU (4T) Pure GPU (Vulkan) Hybrid GPU-CPU (4T) Hybrid Efficiency (1T) Speedup / Advantage
Whisper Tiny 4.59s 9.56s 4.38s 5.48s +4.6% vs CPU (Dispatch overhead minimized)
Whisper Small 15.48s 16.36s 12.93s 23.02s +21.0% vs GPU, +16.5% vs CPU
Large-v3-Turbo 115.05s 87.87s 86.52s 90.39s -28.5s vs CPU, CPU Load 0% during 80s encode

Multi-SoC Verification: Tested across Qualcomm Adreno 830, ARM Mali-G78 (Galaxy S21), and ARM Mali-G77 (Galaxy S20).
Upstream Contribution: Official PR submitted to ggml-org/whisper.cpp (PR #4089).


Installation & Quickstart

Python (PyPI)

pip install termux-stt
from termux_stt import create_engine

# 1. Initialize Engine with Hybrid GPU-Encoder / CPU-Decoder Acceleration
engine = create_engine("whisper", model="small", lang="en", threads=4, split_mode=True)

# 2. Transcribe Audio directly into Subtitles
result = engine.transcribe("samples/jfk_1min.wav")
print("Transcript:\n", result.text)
print("SRT Subtitles:\n", result.to_srt())

# 3. 2-Speaker Diarization without PyTorch
hybrid = create_engine("hybrid", lang="en", num_speakers=2)
diar_result = hybrid.diarize("samples/jfk_1min.wav")
for seg in diar_result.segments:
    print(f"[{seg.speaker}] ({seg.start:.1f}s -> {seg.end:.1f}s): {seg.text}")

CLI Command Line Usage

# 1. High-Performance Hybrid Transcribe (Vulkan GPU Encoder + 4 CPU Threads Decoder)
termux-stt transcribe samples/jfk_1min.wav --model small --split-mode

# 2. Ultra-Low Power Background Transcribe (GPU Encoder + 1 CPU Thread Decoder)
termux-stt transcribe meeting.wav --model small --optimize-1

# 3. Force Pure CPU Execution
termux-stt transcribe samples/jfk_1min.wav --device cpu --threads 4

Node.js / TypeScript (npm)

npm install termux-stt
const { createEngine } = require("termux-stt");

async function main() {
  // 1. Initialize Whisper Engine
  const engine = createEngine("whisper", { model: "tiny", lang: "en", threads: 4 });

  // 2. Transcribe Audio
  const result = await engine.transcribe("samples/jfk_1min.wav");
  console.log("Transcript:", result.text);
  console.log("SRT Subtitles:\n", result.toSrt());
}
main();

Official Documentation & Benchmarks


License

Licensed under the Apache-2.0 License. Copyright (c) 2026 Eunho Kim (@uno-km).

Metadata

Release files for termux-stt 1.3.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for termux-stt 1.3.1
File Size Uploaded
termux_stt-1.3.1.tar.gz 1.7 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for termux-stt 1.3.1
File Interpreter ABI Platform
termux_stt-1.3.1-py3-none-any.whl Python 3 none any Details

Total release size: 1.8 MB

Release files / termux_stt-1.3.1.tar.gz

Download URL termux_stt-1.3.1.tar.gz
Size 1.7 MB
Tags Source
SHA-256 checksum
How to use checksums
4c0efa0317d51577415228c0347d766b74191a18d36feb4dc8253d3a05ead1b5
BLAKE2b-256 checksum
How to use checksums
e92aa58e8a617e4c42093c1e391a33bf11aa3c24b93c6d5e0925127745fd663c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.0

Release files / termux_stt-1.3.1-py3-none-any.whl

Download URL termux_stt-1.3.1-py3-none-any.whl
Size 78.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4c1fa845a045d99ca1b22c824c33855235c4949f58f6e703a329298739bb2c0e
BLAKE2b-256 checksum
How to use checksums
e67b92477a9df4428e4071cd8aab5dcdb8b40eaa926afc52f13faf24aae9a74e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.0

Release history Release notifications | RSS feed

2.0.4

2 release files

2.0.3

2 release files

2.0.2

2 release files

2.0.1

2 release files

2.0.0

2 release files

1.5.1

2 release files

1.5.0

2 release files

1.4.0

2 release files

1.3.3

2 release files

1.3.2

2 release files

This release

1.3.1 This release

2 release files

1.3.0

2 release files

1.2.15

2 release files

1.2.14

2 release files

1.2.13

2 release files

1.2.12

2 release files

1.2.11

2 release files

1.2.10

2 release files

1.2.9

2 release files

1.2.8

2 release files

1.2.7

2 release files

1.2.6

2 release files

1.2.5

2 release files

1.2.4

2 release files

1.2.3

2 release files

1.2.2

2 release files

1.2.1

2 release files

1.2.0

2 release files

1.1.8

1 release file

1.1.7

2 release files

1.1.5

2 release files

1.1.4

2 release files

1.1.3

2 release files

1.1.2

2 release files

1.1.1

2 release files

1.0.9

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page