Wenbi
Wenbi is a CLI-first toolkit that turns media (video/audio/URL) and text into structured Markdown, then rewrites or translates the result. It is built around the ollama/qwen3.5:cloud rewrite/translation model by default, with DeepL as the preferred translator and an LLM fallback.
It supports:
- Video/audio/URL transcription to VTT/Markdown
- Text rewriting (
rewrite,academicstyle) - English interview rewriting (
en-en) with speaker-separated output, now bilingual by default (English → Chinese) - Chinese interview rewriting (
zh-zh) with speaker-separated output - English/Chinese bilingual audio extraction (
en-zh) — keep English, translate to Chinese - Single-language multi-speaker diarization (
speaker) with rewrite + translate - Optional slide-combine (
--ppt/-pon any subcommand) — align slides with speech by timestamp - Batch directory processing (
wenbi-batch)
How it works
Every subcommand follows the same pipeline; the only difference is which stages run and the default options:
input (media / URL / text / subtitles)
│
1. Download & extract ── yt-dlp for URLs, ffmpeg/pydub for media → audio
│
2. ASR / transcribe ── auto | gladia | funasr | whisper
│ gladia is the default when GLADIA_API_KEY is set;
│ funasr (paraformer-zh, local) is the fallback;
│ whisper is the third fallback.
│ → VTT + raw transcript
│
3. Diarize (optional) ── gladia / funasr cam++ → speaker labels
│
4. Rewrite (optional) ── LLM (default ollama/qwen3.5:cloud)
│ rewrites oral speech into written style,
│ preserves speaker turns for interview flows
│
5. Translate (optional) ── DeepL first (needs DEEPL_API_KEY),
│ LLM fallback if DeepL unavailable/fails.
│ EN→ZH glossary applied by default
│ (DeepL glossary API + LLM prompt).
│
6. Slide combine (--ppt only) ── frame extraction + optional OCR → slides aligned
│ with speech by timestamp.
│ Bare --ppt: embed frames as base64 (TYPE 1, no OCR).
│ --ppt <file>: OCR the slides file, match to frames (TYPE 3).
│ Bilingual subcommands: timestamps recovered from VTT
│ via rapidfuzz fuzzy matching before alignment.
│
└── outputs in --output-dir: *_rewritten.md, *_translated.md,
*_bilingual.md, *_zh.md, *_en.md,
*_combine.md, *_diagnostics.json, *.vtt, *.csv
Subcommand → pipeline mapping:
| Command | Stages |
|---|---|
rewrite / rw |
1 → 2 → 4 |
en-en / enen |
1 → 2 → 3 → 4 → 5 (interview rewrite + translate defaults) |
zh-zh / zhzh |
1 → 2 → 3 → 4 (interview rewrite defaults) |
en-zh / enzh |
1 → 2 → 3 (keep EN, drop ZH) → 5 |
speaker / sp |
1 → 2 → 3 → 4 → 5 |
any + --ppt |
adds stage 6 (slide-combine) on top of the subcommand's own stages |
wenbi-batch |
runs one of the above over every media file in a directory |
Install
Prerequisites:
- Python 3.11+
ffmpegin PATH- Gladia API key (optional but recommended): set
GLADIA_API_KEYfor cloud ASR (the default for all audio-capable subcommands). Free tier is 10h/month — see https://www.gladia.io/pricing. Without the key, ASR auto-falls-back to FunASR (local, no key needed, requires the heavy ML deps).
Install:
pip install wenbi
Or from a local checkout:
# from the project directory
pip install -e .
Quick Start
Rewrite:
wenbi rewrite input.mp4 --lang Chinese --llm ollama/qwen3.5:cloud
English interview rewrite:
wenbi en-en interview.mp4 --gladia-key "$GLADIA_API_KEY"
Chinese interview rewrite:
wenbi zh-zh interview.mp4 --gladia-key "$GLADIA_API_KEY"
English/Chinese bilingual audio (keep English, translate to Chinese):
wenbi en-zh bilingual.mp4 --gladia-key "$GLADIA_API_KEY"
Single-language multi-speaker (diarize, rewrite, translate):
wenbi speaker panel.mp4 --source-lang en --gladia-key "$GLADIA_API_KEY"
Slide-combine (any subcommand + --ppt):
# TYPE 1: embed video frames as base64, combine with rewritten speech
wenbi rewrite lecture.mp4 --ppt
# TYPE 3: OCR an external slides file, match to video frames, combine
wenbi rewrite lecture.mp4 --ppt slides.pdf
Commands
rewrite (rw)
Rewrite spoken/transcribed text into written style.
wenbi rewrite <input> [options]
Key options:
--style rewrite|academic--lang--llm--asr-provider auto|gladia|funasr|whisper--cite-timestamps--start-time,--end-time(media/URL)
en-en (enen)
Transcribe an English interview, separate speaker turns, rewrite it as polished written English, and translate to Chinese by default (DeepL first, LLM fallback). Runs stages 1 → 2 → 3 → 4 → 5.
wenbi en-en <input> [options]
Key options:
--speaker-count(default:2)--asr-provider auto|gladia|funasr|whisper--gladia-key(orGLADIA_API_KEYenv var)--llm(default:ollama/qwen3.5:cloud)--lang— target translation language (default:Chinese; set toEnglishto skip translation and keep English-only output)--deepl-key(orDEEPL_API_KEYenv var)--glossary/--no-glossary(default: enabled, EN→ZH term consistency)--glossary-file path.json(custom{english: chinese}glossary)--start-time,--end-time(media/URL)
The rewrite preserves speaker labels and adds a ## Questions for Clarification section when speaker roles, names, terms, or ambiguous ASR phrases need human confirmation.
zh-zh (zhzh)
Transcribe a Chinese interview, separate speaker turns, and rewrite it as polished written Chinese using ollama/qwen3.5:cloud by default.
wenbi zh-zh <input> [options]
Key options:
--speaker-count(default:2)--asr-provider auto|gladia|funasr|whisper--gladia-key(orGLADIA_API_KEYenv var)--llm(default:ollama/qwen3.5:cloud)--start-time,--end-time(media/URL)
The rewrite preserves speaker labels and adds a ## 需要确认的问题 section when speaker roles, names, terms, or ambiguous ASR phrases need human confirmation.
en-zh (enzh)
Extract English from English/Chinese bilingual audio (e.g. interpreted interviews), drop the interpreter language, and translate the kept English into Chinese using DeepL first with LLM fallback.
wenbi en-zh <input> [options]
Key options:
--asr-provider auto|gladia|funasr|whisper(default:auto)--source-lang(default:en) — language to keep--interpreter-lang(default:zh) — language to drop--gladia-key(orGLADIA_API_KEYenv var)--lang— target translation language (default:Chinese)--glossary/--no-glossary(default: enabled, EN→ZH term consistency)--glossary-file path.json(custom{english: chinese}glossary)--no-speaker-labels— disable speaker diarization--save-json— write segment diagnostics and raw provider JSON--start-time,--end-time(media/URL)
Outputs include the kept-language VTT/Markdown, a bilingual Markdown side-by-side, and (optionally) a rewritten English Markdown and diagnostics JSON.
speaker (sp)
Transcribe single-language multi-speaker audio with diarization, then rewrite and translate it. Same engine as en-en/zh-zh but without the interview-style rewrite defaults — use it for panels, podcasts, and any multi-speaker source where you want to keep the source language.
wenbi speaker <input> [options]
Key options:
--asr-provider auto|gladia|funasr|whisper(default:auto)--source-lang(default:en)--speaker-count(default: provider decides)--gladia-key(orGLADIA_API_KEYenv var)--lang— target translation language (default:Chinese)--glossary/--no-glossary(default: enabled, EN→ZH term consistency)--glossary-file path.json(custom{english: chinese}glossary)--no-speaker-labels— disable speaker diarization--save-json— write segment diagnostics and raw provider JSON--start-time,--end-time(media/URL)
Outputs a transcript VTT, transcript Markdown, rewritten Markdown, and (when translation is requested) a bilingual Markdown.
Slide-Combine (--ppt / -p)
Any subcommand accepts --ppt / -p to add stage 6 (slide-combine) on top of its own stages. Useful for lectures and talks where slides should be aligned with the transcript.
wenbi <subcommand> <input> --ppt [slides_file] [options]
Two modes:
- Bare
--ppt/-p(TYPE 1): extract video frames, deduplicate, embed as base64 images with timestamps. No OCR — frames are the slides. --ppt <file>/-p <file>(TYPE 3): OCR the given PPT/PDF/image file, match pages to video frames via SSIM, combine. Higher quality when a clean slides file is available.
Slide options (apply to both modes):
--frame-interval— extract a frame every N seconds (default: 60)--max-slides,--no-deduplicate,--dedup-method,--similarity-threshold,--ssim-threshold,--hist-threshold--no-clean— keep timestamps and image references in combined output--no-ocr— TYPE 3 only: embed slide images as base64 instead of OCR'ing them
Outputs {base}_combine.md (and {base}_combine_clean.md unless --no-clean).
Timestamp alignment
Slide-combine aligns slides to speech by timestamp (### **HH:MM:SS** headers). Two paths:
rewritesubcommand: produces timestamped speech directly (forces--cite-timestampswhen--pptis set). Slides align precisely.- Bilingual subcommands (
en-zh,en-en,zh-zh,speaker): the rewritten output strips timestamps (clean prose separated by---).--pptrecovers timestamps by fuzzy-matching each rewritten paragraph against the VTT file from stage 2 ASR usingrapidfuzz. Sincegroup_into_topicspreserves 100% of transcript text andrewrite_englishkeeps ~97% wording, the match scores high and the correct VTT segment window is found for each paragraph. The recovered timestamps are then used for slide alignment.
Supported Inputs
- Media:
.mp4 .avi .mov .mkv .flv .wmv .m4v .webm .mp3 .flac .aac .ogg .m4a .opus - Text/subtitles:
.vtt .srt .ass .ssa .sub .smi .txt .md .markdown .docx - URL inputs are supported for media flows.
Common Global Options
Used by subcommands:
--output-dir--lang--llm--translation-engine auto|deepl|ollama|openai—openaiselectsopenai/gpt-5.6-terraand disables DeepL--chunk-length--max-tokens--timeout--temperature--asr-provider auto|gladia|funasr|whisper--transcribe-lang--multi-language--verbose
OpenAI GPT-5.6-terra translation
Use the OpenAI Platform API key as an environment variable; do not put it in a command, config file, or source code. A ChatGPT/Codex subscription does not itself supply an API key or API credits.
export OPENAI_API_KEY="sk-..."
uv run wenbi speaker "https://www.youtube.com/watch?v=tFX_IVRAcXI&t=18183s" \
--translation-engine openai --lang Chinese --asr-provider gladia -v
With --translation-engine openai, the topic grouping, English cleanup, and
Chinese translation all use GPT-5.6-terra; DeepL and Ollama are not used.
GPT-5 models only accept temperature=1, so Wenbi applies that value even
when the general CLI default is 0.1.
Glossary
Translation-capable subcommands (en-en, en-zh, speaker) apply a built-in EN→ZH patristic glossary by default for term consistency. This covers ~800 Orthodox Christian / patristic terms (e.g. theosis → 神化, theoria → 静观, Origen → 奥利金).
--glossary/--no-glossary— toggle the glossary (default: enabled). No-op when the target language is not Chinese.--glossary-file path.json— supply a custom glossary JSON ({english: chinese}dict). Overrides the built-in patristic glossary.
The glossary is wired into both translation paths:
- DeepL: creates a DeepL server-side glossary from the term pairs and passes it to
translate_text. - LLM fallback: injects the glossary as a
glossaryfield on the DSPyTranslateSignatureprompt.
Output Files
Typical outputs:
*_rewritten.md*_translated.md*_academic.md*_en.md,*_en.vtt(English interview transcripts)*_zh.md,*_zh.vtt(Chinese interview transcripts)*_bilingual.md(en-zh and speaker translated output)*_diagnostics.json(when--save-jsonis used)*_combine.md/*_combine_clean.md(PPT workflows)*.vtt,*.csv(depending on flow)
Batch Processing
Process a directory of media files:
wenbi-batch <input_dir> --output-dir <dir> --md
Optional config:
wenbi-batch <input_dir> --config config.yaml
YAML Config (CLI)
wenbi supports YAML via --config.
Example:
input: lecture.mp4
output_dir: ./out
llm: ollama/qwen3.5:cloud
lang: Chinese
chunk_length: 20
Multi-input format is also supported using inputs:.
Python API
from wenbi.main import process_input
text, md_file, csv_file, base_name = process_input(
file_path="input.mp4",
subcommand="rewrite",
lang="Chinese",
llm="ollama/qwen3.5:cloud",
)
Troubleshooting
- No DeepL translation output:
- Set
DEEPL_API_KEYor--deepl-key - Run with
--verboseto confirm DeepL connectivity
- Set
- Fallback LLM not working:
- Ensure your provider is reachable (for example, Ollama running locally for
ollama/...)
- Ensure your provider is reachable (for example, Ollama running locally for
- PPT OCR issues:
- Ensure
marker_singleand OCR dependencies are installed correctly
- Ensure
License
Apache-2.0
Release files for wenbi 0.142.10
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| wenbi-0.142.10.tar.gz | 6.3 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| wenbi-0.142.10-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 6.4 MB
Release files / wenbi-0.142.10.tar.gz
| Download URL | wenbi-0.142.10.tar.gz |
|---|---|
| Size | 6.3 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
2fb281a4a38b01a68169043b7ad482707a52a123a6486d4b2687babe65d4c086
|
|
BLAKE2b-256 checksum How to use checksums |
506299cd74e3651529ba63965e2125c9e5c65c20ebe93799c7a92f8a799fbbaa
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/5.1.1 CPython/3.13.1
|
Release files / wenbi-0.142.10-py3-none-any.whl
| Download URL | wenbi-0.142.10-py3-none-any.whl |
|---|---|
| Size | 118.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
024421fe8e1b57918376e7b608bc044d8bbec16fba7f8dca22be1010eca1f818
|
|
BLAKE2b-256 checksum How to use checksums |
3dadeda829dd7f29477d14eef1540a6e7b0e4fac9472809023f922e3a3b6842a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/5.1.1 CPython/3.13.1
|