gemini-transcribe-wrapper
Transcribe hours of audio for free with Google AI model gemini-3.5-transcribe — auto-chunked, rate-limit aware, and exported straight to SRT and TXT (with optional speaker diarization). GitHub PyPI
Quick Start
Prerequisites
- Get an API key from https://aistudio.google.com/api-keys/ for free.
- Install uv (https://docs.astral.sh/uv/getting-started/installation/).
Install the tool
Linux / macOS:
export GEMINI_API_KEY=your_key_here
uv tool install --python 3.13 gemini-transcribe-wrapper@latest
gtw -v
Windows (Command Prompt):
set GEMINI_API_KEY=your_key_here
uv tool install --python 3.13 gemini-transcribe-wrapper@latest
gtw -v
Transcribe for free
gtw sample.mp4 # with `GEMINI_API_KEY` set
# or
gtw --gemini-api-key YOUR_API_KEY sample.mp4
Output (default: --no-diarize — no speaker labels, fewer API calls):
sample.transcript.json
sample.srt
sample.txt
For speaker-diarized output, opt in with --diarize:
gtw --diarize sample.mp4
Output (with --diarize):
sample.diarized.transcript.json
sample.diarized.srt
sample.srt
sample.txt
What are improved by this project?
This wrapper automatically overcomes Google Gemini API's free tier limits and constraints (as of Aug 2026):
- Audio Length Limit — depends on whether you ask for speaker diarization:
--no-diarize(default): up to ~60 min per API call. The wrapper cuts the file into 59-min logical units, each sent as one API call.--diarize: up to ~30 min per API call. The wrapper cuts into 29-min chunks to stay under the limit with a 1-min safety margin.
- Rate Limit (Max 2 RPM): Applies a built-in delay (default 30s) between sequential requests to prevent rate-limit errors.
- Daily Quota (Max 25 RPD or input-token quota): Tracks daily Pacific-time API usage locally (
~/.cache/gemini-transcribe-wrapper/usage-<sha256(key)[:12]>.json) and warns clearly when limits approach. - Raw Response to Ready-to-Use Subtitles: Converts AI transcription output directly into
.srt,.txt, and (with--diarize).diarized.srtfiles in a single run. - Free-Tier-Friendly Defaults:
--no-diarizeis the default to minimize API calls; the wrapper packs each call to the per-mode maximum. - 429 = Stop the Batch: On HTTP 429 (rate limit or quota exhausted), the wrapper prints retry suggestions and aborts the batch immediately. The CLI exits with code
2to distinguish quota errors from other failures (code1).
On Quota / Rate-Limit (HTTP 429)
The wrapper does not retry automatically on 429. When a 429 is detected it:
- Logs the raw error and the quota category (free-tier daily quota vs short-term rate limit).
- Prints retry suggestions: "wait about 1 minute for a short-term 429, or wait Xh Ym (sleep Zs) until PST midnight for the daily quota to reset, then re-run."
- Aborts the rest of the batch (no point burning more calls on a guaranteed 429).
- Returns exit code
2.
Sample log on a quota hit:
ERROR Rate limit / quota exceeded (429): Error code: 429 - You exceeded your current quota ...
ERROR It looks like you hit the free tier daily quota (25 calls/day).
ERROR You hit the Gemini API rate limits:
- max 2 API calls per minute
- max 30 minutes of audio per call
- max 25 API calls per day (free tier)
ERROR To retry: wait about 1 minute for a short-term 429, or wait 5h 40m (sleep 20400s) until PST midnight for the daily quota to reset, then re-run.
ERROR Switching to a paid tier (enable billing) removes the free-tier limits.
ERROR Aborting batch: quota / rate limit hit while processing <file>. Remaining files will not be processed.
Batch Bulk Transcribing Tip
The free tier caps 25 calls/day per PST day. For multi-day batches, either:
- Run in shorter bursts so you naturally pause before hitting the quota, or
- Swap in a fresh API key (each key has its own counter under
~/.cache/gemini-transcribe-wrapper/usage-<sha256(key)[:12]>.json), or - Wait for the PST reset — on a 429 the wrapper prints the exact sleep seconds needed; wrap your batch in a shell loop:
# Run, on quota hit wait until PST midnight, then re-run.
while ! gtw '*.mp4' --diarize; do
echo "Batch hit a quota; waiting for PST reset..."
# Read the sleep seconds from the previous log, or sleep 1h and retry.
sleep 3600
done
Tip: --no-diarize halves (or better) the number of API calls vs --diarize for the same audio, since each call covers ~2× the audio length.
Diarizing Tip (Speaker Labels)
--diarize produces .diarized.srt with raw speaker ids (spk:0, spk:1, ...). Map them to real names one pass at a time with --speakers: delete .diarized.srt, re-run with a more complete map, repeat until no spk: string remains.
# Pass 0: no --speakers — every cue keeps its raw spk:# tag.
gtw --diarize interview.mp4
# Wrapper logs the unmapped speakers and prints the re-render recipe:
# WARNING Some speakers are not covered by --speakers mapping.
# Speaker map: spk:0 spk:1 spk:2
# Unmapped: spk:0, spk:1, spk:2
# WARNING To re-render with names, delete the .diarized.srt and re-run with the
# option, editing the Name# entries: rm 'interview.diarized.srt' &&
# gtw 'interview.mp4' --diarize --speakers 'spk:0=Name0;spk:1=Name1;spk:2=Name2;'
# Pass 1: rename spk:0 → Host. Edit the recipe, delete the old .diarized.srt, re-run.
rm interview.diarized.srt
gtw --diarize interview.mp4 --speakers 'spk:0=Host;'
# Pass 2: add spk:1 → Guest. Same drill.
rm interview.diarized.srt
gtw --diarize interview.mp4 --speakers 'spk:0=Host;spk:1=Guest;'
# Pass 3: add spk:2 → Interpreter. Now every tag is a real name.
rm interview.diarized.srt
gtw --diarize interview.mp4 --speakers 'spk:0=Host;spk:1=Guest;spk:2=Interpreter;'
Simulation of the iterative renaming (.diarized.srt excerpt after each pass):
# Pass 0 (no --speakers): all raw ids
[spk:0] Hello everyone, welcome to the show.
[spk:1] Today we have a special guest with us.
[spk:2] Thanks for having me, it's great to be here.
# Pass 1 (--speakers 'spk:0=Host;')
[Host] Hello everyone, welcome to the show.
[spk:1] Today we have a special guest with us.
[spk:2] Thanks for having me, it's great to be here.
# Pass 2 (add spk:1=Guest)
[Host] Hello everyone, welcome to the show.
[Guest] Today we have a special guest with us.
[spk:2] Thanks for having me, it's great to be here.
# Pass 3 (add spk:2=Interpreter) — done, no spk: left
[Host] Hello everyone, welcome to the show.
[Guest] Today we have a special guest with us.
[Interpreter] Thanks for having me, it's great to be here.
Tip: each pass is a no-API re-render — the wrapper reads <base>.diarized.transcript.json (kept by default) and rewrites .diarized.srt from it, so iterating is fast and free.
Note: --speakers is ignored when --diarize is off (the default). Pass --diarize to enable speaker mapping.
Relevant Repositories
License
MIT
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file gemini_transcribe_wrapper-0.0.18.tar.gz.
File metadata
- Download URL: gemini_transcribe_wrapper-0.0.18.tar.gz
- Upload date:
- Size: 93.4 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.11.8 {"installer":{"name":"uv","version":"0.11.8","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7c9fb11343ea5bd107f0d062470089ffe3fbc824b4ee8434bb37d96d9d59ff87
|
|
| MD5 |
5e53d3ee10280895ba725bfd3c1c6823
|
|
| BLAKE2b-256 |
0887a6050fd5fbc755ab1123a0428e8fd7a62eeb7e62a308ead0359ea98b237f
|
File details
Details for the file gemini_transcribe_wrapper-0.0.18-py3-none-any.whl.
File metadata
- Download URL: gemini_transcribe_wrapper-0.0.18-py3-none-any.whl
- Upload date:
- Size: 37.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.11.8 {"installer":{"name":"uv","version":"0.11.8","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0ecffbe8c18d4bfe774ef4001abb4e4b8c5fe6e60fa831b332749785f67c5098
|
|
| MD5 |
2d5eff4b98571a3ebc6f3e87336bf398
|
|
| BLAKE2b-256 |
879f6003b3f4e99547c50040bfb52d4d68d5729f9a24b01bb8b42df8d2b7b8c5
|