Termux-STT
디바이스 리소스를 활용한 통합 온디바이스 음성인식(STT) 및 차세대 신경망 화자 분리(PyAnnote 3.0 + CAM++ 192d + TS-VAD) 프레임워크
Unified On-Device Speech-to-Text & Neural Multi-Speaker Diarization (PyAnnote 3.0 + CAM++ 192d + TS-VAD Overlap Resolver)
Architecture & Overview
Whisper.cpp(네이티브 Vulkan GPU / NEON CPU) 및 Sherpa-ONNX 듀얼 모던 엔진을 통합하고, PyAnnote 3.0 신경망 세그멘테이션, 3D-Speaker CAM++ 192차원 성문 임베딩, TS-VAD 동시 발화(마이크 물림) 분리기를 결합하여 완전 온디바이스 실시간 다중 화자 음성인식을 구현합니다.
Integrates Whisper.cpp (Native Vulkan GPU / NEON CPU) and Sherpa-ONNX with PyAnnote 3.0 neural segmentation, 3D-Speaker CAM++ 192-dim embeddings, and TS-VAD overlapped speech resolution with zero cloud egress.
Installation & Quickstart
Python (PyPI)
pip install termux-stt
from termux_stt import create_engine
# 1. Initialize Whisper Engine with Vulkan GPU / CPU Hybrid Acceleration
engine = create_engine("whisper", model="small", lang="ko", threads=4, split_mode=True)
# 2. Transcribe Audio directly into Subtitles
result = engine.transcribe("samples/jfk_1min.wav")
print("Transcript:\n", result.text)
print("SRT Subtitles:\n", result.to_srt())
# 3. Multi-Speaker Neural Diarization (PyAnnote 3.0 + CAM++ 192d + TS-VAD)
hybrid = create_engine("hybrid", lang="ko", num_speakers=2)
diar_result = hybrid.diarize("samples/kor_diarization.wav")
for seg in diar_result.segments:
print(f"[{seg.speaker}] ({seg.start:.1f}s -> {seg.end:.1f}s): {seg.text}")
Node.js / TypeScript (npm)
npm install termux-stt
const { createEngine } = require("termux-stt");
async function main() {
// 1. Initialize Whisper Engine
const engine = createEngine("whisper", { model: "tiny", lang: "ko", threads: 4 });
// 2. Transcribe Audio
const result = await engine.transcribe("samples/jfk_1min.wav");
console.log("Transcript:", result.text);
console.log("SRT Subtitles:\n", result.toSrt());
// 3. Neural Diarization via Hybrid Engine
const hybrid = createEngine("hybrid", { lang: "ko", numSpeakers: 2 });
const diarResult = await hybrid.diarize("samples/kor_diarization.wav");
for (const seg of diarResult.segments) {
console.log(`[${seg.speaker}] (${seg.start.toFixed(1)}s -> ${seg.end.toFixed(1)}s): ${seg.text}`);
}
}
main();
Official Documentation & Benchmarks
- Official Architecture & API Reference
- Ecosystem Metrics & Registry Stats
- AMEVA Open-Source Foundation Portal
License
Licensed under the Apache-2.0 License. Copyright (c) 2026 Eunho Kim (@uno-km).
Metadata
Release files for termux-stt 2.0.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| termux_stt-2.0.2.tar.gz | 3.6 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| termux_stt-2.0.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 3.7 MB
Release files / termux_stt-2.0.2.tar.gz
| Download URL | termux_stt-2.0.2.tar.gz |
|---|---|
| Size | 3.6 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
b58416161c6607f9d80d7a20685ac383f5cdbb149c66976a046c3af448f9445a
|
|
BLAKE2b-256 checksum How to use checksums |
fd0b896525354d936a5b695840f9dc536f1c64c4ff270be4e6440b7f42e2ccff
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.17
|
Release files / termux_stt-2.0.2-py3-none-any.whl
| Download URL | termux_stt-2.0.2-py3-none-any.whl |
|---|---|
| Size | 91.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ff3adb732536a2398001e1a844d4c836f7f21300a4771422d11d2a2815c31d29
|
|
BLAKE2b-256 checksum How to use checksums |
a2cfdb693efb6384fba1c03cd2aa2a03026926df1118455e10e965f7f9ac0f74
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.17
|