stream-translator-gpt
Command line utility to transcribe or translate audio from livestreams in real time. Uses yt-dlp to get livestream URLs from various services and Whisper / Faster-Whisper for transcription.
This fork optimized the audio slicing logic based on VAD, introduced GPT API / Gemini API to support language translation beyond English, and supports input from the audio devices.
Prerequisites
Linux or Windows:
- Python >= 3.8 (Recommend >= 3.10)
- Install CUDA on your system. You can check the installed CUDA version with
nvcc --version. - Install cuDNN to your CUDA dir if you want to use Faseter-Whisper.
- Install PyTorch (with CUDA) to your Python.
- Create a Google API key if you want to use Gemini API for translation. (Recommend, Free 60 requests / minute)
- Create a OpenAI API key if you want to use Whisper API for transcription or GPT API for translation.
If you are in Windows, you also need to:
- Install and add ffmpeg to your PATH.
- Install yt-dlp and add it to your PATH.
Installation
Install release version from PyPI (Recommend):
pip install stream-translator-gpt
stream-translator-gpt
or
Clone master version code from Github:
git clone https://github.com/ionic-bond/stream-translator-gpt.git
pip install -r ./stream-translator-gpt/requirements.txt
python3 ./stream-translator-gpt/translator.py
Usage
-
Transcribe live streaming (default use Whisper):
stream-translator-gpt {URL} --model large --language {input_language} -
Transcribe by Faster Whisper:
stream-translator-gpt {URL} --model large --language {input_language} --use_faster_whisper -
Transcribe by Whisper API:
stream-translator-gpt {URL} --language {input_language} --use_whisper_api --openai_api_key {your_openai_key} -
Translate to other language by Gemini:
stream-translator-gpt {URL} --model large --language ja --gpt_translation_prompt "Translate from Japanese to Chinese" --google_api_key {your_google_key} -
Translate to other language by GPT:
stream-translator-gpt {URL} --model large --language ja --gpt_translation_prompt "Translate from Japanese to Chinese" --openai_api_key {your_openai_key} -
Using Whisper API and Gemini at the same time:
stream-translator-gpt {URL} --model large --language ja --use_whisper_api --openai_api_key {your_openai_key} --gpt_translation_prompt "Translate from Japanese to Chinese" --google_api_key {your_google_key} -
Local video/audio file as input:
stream-translator-gpt /path/to/file --model large --language {input_language} -
Computer microphone as input:
stream-translator-gpt device --model large --language {input_language}Will use the system's default audio device as input.
If you want to use another audio input device,
stream-translator-gpt device --print_all_devicesget device index and then run the CLI with--device_index {index}.If you want to use the audio output of another program as input, you need to enable stereo mix.
-
Sending result to Cqhttp:
stream-translator-gpt {URL} --model large --language {input_language} --cqhttp_url {your_cqhttp_url} --cqhttp_token {your_cqhttp_token} -
Sending result to Discord:
stream-translator-gpt {URL} --model large --language {input_language} --discord_webhook_url {your_discord_webhook_url}
Release files for stream-translator-gpt 2024.3.22
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| stream-translator-gpt-2024.3.22.tar.gz | 1.1 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| stream_translator_gpt-2024.3.22-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 2.3 MB
Release files / stream-translator-gpt-2024.3.22.tar.gz
| Download URL | stream-translator-gpt-2024.3.22.tar.gz |
|---|---|
| Size | 1.1 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
640b7053ec799519687fd5f3bbb9865440f3d1517928c5d9c5789e278c40a92d
|
|
BLAKE2b-256 checksum How to use checksums |
ef5d10ceeac5c3643a3f8b71592b41a3ca3dbbab5c5b53b0a3bd5d4def27f592
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/5.0.0 CPython/3.10.12
|
Release files / stream_translator_gpt-2024.3.22-py3-none-any.whl
| Download URL | stream_translator_gpt-2024.3.22-py3-none-any.whl |
|---|---|
| Size | 1.1 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0271a4de5c4f5cc152f95c731c9844829fc04a5beebbd367fc642e17256945e7
|
|
BLAKE2b-256 checksum How to use checksums |
bd75e3a0f3a7daa142c5864b91121aa02974676e8d380d94dfdd180afd4cd19f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/5.0.0 CPython/3.10.12
|