nanovt
Transcribe video or audio files with OpenAI or Volcengine speech-to-text.
The script extracts audio with ffmpeg, converts it to mono 16 kHz, transcribes it, and writes a single text file. The OpenAI provider splits the audio into short chunks; the Volcengine provider submits the whole file to an asynchronous job.
Requirements
- Python 3.13+
ffmpegOPENAI_API_KEYin your environment (OpenAI provider, the default)VOLC_ASR_API_KEYin your environment (Volcengine provider)
Installation
From PyPI:
pip install nanovt
Or install it as an isolated uv tool:
uv tool install nanovt
Usage
nanovt input.mp4
By default, input.mp4 writes input.txt.
To force a language, pass an optional language code. For example, English:
nanovt input.mp4 --language en
If --language is omitted, the model detects the language automatically.
Other useful options:
nanovt input.mp4 --chunk-seconds 180 --retries 3
nanovt input.mp4 --output transcript.txt
nanovt input.mp4 --keep-temp
For speaker-labeled dialogue transcription, enable diarization:
nanovt input.mp4 --diarize
This uses gpt-4o-transcribe-diarize by default and writes dialogue lines such
as A: ..., B: ..., and C: ....
Volcengine provider
Volcengine's big-model ASR handles Chinese and code-switched speech well. Set VOLC_ASR_API_KEY and select the provider:
nanovt input.mp4 --provider volc --language zh --diarize
The audio is compressed to a mono 16 kHz MP3 and submitted inline, so no public URL is needed. --model defaults to bigmodel; override the resource id with VOLC_ASR_RESOURCE_ID (default volc.seedasr.auc). --chunk-seconds is unused for this provider. Diarized output uses the same A: ... / B: ... lines.
Development
uv sync --group dev
uv run pytest
uv run nanovt input.mp4
Release files for nanovt 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| nanovt-0.3.0.tar.gz | 47.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| nanovt-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 59.7 kB
Release files / nanovt-0.3.0.tar.gz
| Download URL | nanovt-0.3.0.tar.gz |
|---|---|
| Size | 47.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ed89c8a1812330e073712961617900b6978c2fe2ce7387b6795c2e71932847c6
|
|
BLAKE2b-256 checksum How to use checksums |
502ce84a54077d5cc1fd4b12924122bb0e948204e3b75ece84e50161ca757e32
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 27, 2026.
Transparency logRelease files / nanovt-0.3.0-py3-none-any.whl
| Download URL | nanovt-0.3.0-py3-none-any.whl |
|---|---|
| Size | 12.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
3f4855bc520f543da22366f0f1a979a6ef2f225c870f4d9863cb5848eaa812a5
|
|
BLAKE2b-256 checksum How to use checksums |
e5e3732f99cdb80360883ec2f09cd4be586fdd712afcd9f3f4ee5fdf47b23f7f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 27, 2026.
Transparency log