Skip to main content

Termux-STT

PyPI Python npm npm downloads License

디바이스 리소스를 활용한 통합 온디바이스 음성인식(STT) 및 순수 Python 128d X-Vector 화자 분리 프레임워크
Unified On-Device Speech-to-Text Utilizing Device Resources & Pure Python 128d X-Vector Speaker Diarization


Architecture & Overview

Whisper.cpp, Vosk, Sherpa-ONNX 3대 네이티브 엔진을 통합하고, 닫힌 형태 순수 Python 128차원 X-Vector 클러스터링을 결합하여 80MB 미만의 초경량 메모리로 100% 온디바이스 실시간 음성인식과 화자 분리를 구현합니다.

Integrates Whisper.cpp, Vosk, and Sherpa-ONNX with a closed-form pure-Python 128-dimensional X-Vector clustering algorithm that operates in under 80MB RAM with zero cloud egress.


Asymmetric Hybrid Architecture Benchmark (Galaxy S25 - Snapdragon 8 Elite)

Test Audio: JFK Inaugural Address (60.59s mono 16kHz WAV)
Silicon: Qualcomm Snapdragon 8 Elite (Adreno 830 GPU + Oryon CPU)

Model Pure CPU (4T) Pure GPU (Vulkan) Hybrid GPU-CPU (4T) Hybrid Efficiency (1T) Speedup / Advantage
Whisper Tiny 4.59s 9.56s 4.38s 5.48s +4.6% vs CPU (Dispatch overhead minimized)
Whisper Small 15.48s 16.36s 12.93s 23.02s +21.0% vs GPU, +16.5% vs CPU
Large-v3-Turbo 115.05s 87.87s 86.52s 90.39s -28.5s vs CPU, CPU Load 0% during 80s encode

Multi-SoC Verification: Tested across Qualcomm Adreno 830, ARM Mali-G78 (Galaxy S21), and ARM Mali-G77 (Galaxy S20).
Upstream Contribution: Official PR submitted to ggml-org/whisper.cpp (PR #4089).


Installation & Quickstart

Python (PyPI)

pip install termux-stt
from termux_stt import create_engine

# 1. Initialize Engine with Hybrid GPU-Encoder / CPU-Decoder Acceleration
engine = create_engine("whisper", model="small", lang="en", threads=4, split_mode=True)

# 2. Transcribe Audio directly into Subtitles
result = engine.transcribe("samples/jfk_1min.wav")
print("Transcript:\n", result.text)
print("SRT Subtitles:\n", result.to_srt())

# 3. 2-Speaker Diarization without PyTorch
hybrid = create_engine("hybrid", lang="en", num_speakers=2)
diar_result = hybrid.diarize("samples/jfk_1min.wav")
for seg in diar_result.segments:
    print(f"[{seg.speaker}] ({seg.start:.1f}s -> {seg.end:.1f}s): {seg.text}")

CLI Command Line Usage

# 1. High-Performance Hybrid Transcribe (Vulkan GPU Encoder + 4 CPU Threads Decoder)
termux-stt transcribe samples/jfk_1min.wav --model small --split-mode

# 2. Ultra-Low Power Background Transcribe (GPU Encoder + 1 CPU Thread Decoder)
termux-stt transcribe meeting.wav --model small --optimize-1

# 3. Force Pure CPU Execution
termux-stt transcribe samples/jfk_1min.wav --device cpu --threads 4

Node.js / TypeScript (npm)

npm install termux-stt
const { createEngine } = require("termux-stt");

async function main() {
  // 1. Initialize Whisper Engine
  const engine = createEngine("whisper", { model: "tiny", lang: "en", threads: 4 });

  // 2. Transcribe Audio
  const result = await engine.transcribe("samples/jfk_1min.wav");
  console.log("Transcript:", result.text);
  console.log("SRT Subtitles:\n", result.toSrt());
}
main();

Official Documentation & Benchmarks


License

Licensed under the Apache-2.0 License. Copyright (c) 2026 Eunho Kim (@uno-km).

Metadata

Release files for termux-stt 1.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for termux-stt 1.3.0
File Size Uploaded
termux_stt-1.3.0.tar.gz 1.7 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for termux-stt 1.3.0
File Interpreter ABI Platform
termux_stt-1.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.8 MB

Release files / termux_stt-1.3.0.tar.gz

Download URL termux_stt-1.3.0.tar.gz
Size 1.7 MB
Tags Source
SHA-256 checksum
How to use checksums
483db6202fc86d97b6cb33e640b261573bcd17b2148d4a1ebdd7efc8b051133a
BLAKE2b-256 checksum
How to use checksums
4076292e0b540aaeef1498bd2f4596619fc587d258f660d3300f9753a50c9820
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release files / termux_stt-1.3.0-py3-none-any.whl

Download URL termux_stt-1.3.0-py3-none-any.whl
Size 78.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
377a5bec5343fcb7a345effa55af2f3500eb99cc60aa3cbfe97758bd6187eaea
BLAKE2b-256 checksum
How to use checksums
db260a6e0a171fccffaa88267a04146c824f6e79857bb1c725740e6b231a4f53
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release history Release notifications | RSS feed

2.0.4

2 release files

2.0.3

2 release files

2.0.2

2 release files

2.0.1

2 release files

2.0.0

2 release files

1.5.1

2 release files

1.5.0

2 release files

1.4.0

2 release files

1.3.3

2 release files

1.3.2

2 release files

1.3.1

2 release files

This release

1.3.0 This release

2 release files

1.2.15

2 release files

1.2.14

2 release files

1.2.13

2 release files

1.2.12

2 release files

1.2.11

2 release files

1.2.10

2 release files

1.2.9

2 release files

1.2.8

2 release files

1.2.7

2 release files

1.2.6

2 release files

1.2.5

2 release files

1.2.4

2 release files

1.2.3

2 release files

1.2.2

2 release files

1.2.1

2 release files

1.2.0

2 release files

1.1.8

1 release file

1.1.7

2 release files

1.1.5

2 release files

1.1.4

2 release files

1.1.3

2 release files

1.1.2

2 release files

1.1.1

2 release files

1.0.9

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page