Skip to main content

Termux-STT

PyPI Python npm npm downloads License

디바이스 리소스를 활용한 통합 온디바이스 음성인식(STT) 및 순수 Python 128d X-Vector 화자 분리 프레임워크
Unified On-Device Speech-to-Text Utilizing Device Resources & Pure Python 128d X-Vector Speaker Diarization


Architecture & Overview

Whisper.cpp, Vosk, Sherpa-ONNX 3대 네이티브 엔진을 통합하고, 닫힌 형태 순수 Python 128차원 X-Vector 클러스터링을 결합하여 80MB 미만의 초경량 메모리로 100% 온디바이스 실시간 음성인식과 화자 분리를 구현합니다.

Integrates Whisper.cpp, Vosk, and Sherpa-ONNX with a closed-form pure-Python 128-dimensional X-Vector clustering algorithm that operates in under 80MB RAM with zero cloud egress.


Asymmetric Hybrid Architecture Benchmark (Galaxy S25 - Snapdragon 8 Elite)

Test Audio: JFK Inaugural Address (60.59s mono 16kHz WAV)
Silicon: Qualcomm Snapdragon 8 Elite (Adreno 830 GPU + Oryon CPU)

Model Pure CPU (4T) Pure GPU (Vulkan) Hybrid GPU-CPU (4T) Hybrid Efficiency (1T) Speedup / Advantage
Whisper Tiny 4.59s 9.56s 4.38s 5.48s +4.6% vs CPU (Dispatch overhead minimized)
Whisper Small 15.48s 16.36s 12.93s 23.02s +21.0% vs GPU, +16.5% vs CPU
Large-v3-Turbo 115.05s 87.87s 86.52s 90.39s -28.5s vs CPU, CPU Load 0% during 80s encode

Multi-SoC Verification: Tested across Qualcomm Adreno 830, ARM Mali-G78 (Galaxy S21), and ARM Mali-G77 (Galaxy S20).
Upstream Contribution: Official PR submitted to ggml-org/whisper.cpp (PR #4089).


Installation & Quickstart

Python (PyPI)

pip install termux-stt
from termux_stt import create_engine

# 1. Initialize Engine with Hybrid GPU-Encoder / CPU-Decoder Acceleration
engine = create_engine("whisper", model="small", lang="en", threads=4, split_mode=True)

# 2. Transcribe Audio directly into Subtitles
result = engine.transcribe("samples/jfk_1min.wav")
print("Transcript:\n", result.text)
print("SRT Subtitles:\n", result.to_srt())

# 3. 2-Speaker Diarization without PyTorch
hybrid = create_engine("hybrid", lang="en", num_speakers=2)
diar_result = hybrid.diarize("samples/jfk_1min.wav")
for seg in diar_result.segments:
    print(f"[{seg.speaker}] ({seg.start:.1f}s -> {seg.end:.1f}s): {seg.text}")

CLI Command Line Usage

# 1. High-Performance Hybrid Transcribe (Vulkan GPU Encoder + 4 CPU Threads Decoder)
termux-stt transcribe samples/jfk_1min.wav --model small --split-mode

# 2. Ultra-Low Power Background Transcribe (GPU Encoder + 1 CPU Thread Decoder)
termux-stt transcribe meeting.wav --model small --optimize-1

# 3. Force Pure CPU Execution
termux-stt transcribe samples/jfk_1min.wav --device cpu --threads 4

Node.js / TypeScript (npm)

npm install termux-stt
const { createEngine } = require("termux-stt");

async function main() {
  // 1. Initialize Whisper Engine
  const engine = createEngine("whisper", { model: "tiny", lang: "en", threads: 4 });

  // 2. Transcribe Audio
  const result = await engine.transcribe("samples/jfk_1min.wav");
  console.log("Transcript:", result.text);
  console.log("SRT Subtitles:\n", result.toSrt());
}
main();

Official Documentation & Benchmarks


License

Licensed under the Apache-2.0 License. Copyright (c) 2026 Eunho Kim (@uno-km).

Metadata

Release files for termux-stt 1.3.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for termux-stt 1.3.2
File Size Uploaded
termux_stt-1.3.2.tar.gz 1.7 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for termux-stt 1.3.2
File Interpreter ABI Platform
termux_stt-1.3.2-py3-none-any.whl Python 3 none any Details

Total release size: 1.8 MB

Release files / termux_stt-1.3.2.tar.gz

Download URL termux_stt-1.3.2.tar.gz
Size 1.7 MB
Tags Source
SHA-256 checksum
How to use checksums
7c150dba5b674541747ac4b4505bb77b7a4f0226e29ba24665507dd71075b749
BLAKE2b-256 checksum
How to use checksums
d959654b2c12ae763cd8cf37acfb4cda4b279cd9f920499b534ad06d3c41fc6d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.0

Release files / termux_stt-1.3.2-py3-none-any.whl

Download URL termux_stt-1.3.2-py3-none-any.whl
Size 71.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0631a377959c036d234e553d5e4f7b4d704ca2c2e8027647a91a059bc4b196c9
BLAKE2b-256 checksum
How to use checksums
cc32f3c4d18eaab04c14f5a0964ec7d6210e10946b53b9ca2dfab28cb74c0b76
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.0

Release history Release notifications | RSS feed

2.0.4

2 release files

2.0.3

2 release files

2.0.2

2 release files

2.0.1

2 release files

2.0.0

2 release files

1.5.1

2 release files

1.5.0

2 release files

1.4.0

2 release files

1.3.3

2 release files

This release

1.3.2 This release

2 release files

1.3.1

2 release files

1.3.0

2 release files

1.2.15

2 release files

1.2.14

2 release files

1.2.13

2 release files

1.2.12

2 release files

1.2.11

2 release files

1.2.10

2 release files

1.2.9

2 release files

1.2.8

2 release files

1.2.7

2 release files

1.2.6

2 release files

1.2.5

2 release files

1.2.4

2 release files

1.2.3

2 release files

1.2.2

2 release files

1.2.1

2 release files

1.2.0

2 release files

1.1.8

1 release file

1.1.7

2 release files

1.1.5

2 release files

1.1.4

2 release files

1.1.3

2 release files

1.1.2

2 release files

1.1.1

2 release files

1.0.9

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page