Skip to main content

gemini-transcribe-wrapper

Transcribe hours of audio for free with Google AI model gemini-3.5-transcribe — auto-chunked, rate-limit aware, and exported straight to SRT and TXT (with optional speaker diarization). GitHub PyPI

Quick Start

Prerequisites

Install the tool

Linux / macOS:

export GEMINI_API_KEY=your_key_here
uv tool install --python 3.13 gemini-transcribe-wrapper@latest
gtw -v

Windows (Command Prompt):

set GEMINI_API_KEY=your_key_here
uv tool install --python 3.13 gemini-transcribe-wrapper@latest
gtw -v

Transcribe for free

gtw sample.mp4   # with `GEMINI_API_KEY` set
# or
gtw --gemini-api-key YOUR_API_KEY sample.mp4

Output (default: --no-diarize — no speaker labels, fewer API calls):

sample.transcript.json
sample.srt
sample.txt

For speaker-diarized output, opt in with --diarize:

gtw --diarize sample.mp4

Output (with --diarize):

sample.diarized.transcript.json
sample.diarized.srt
sample.srt
sample.txt

What are improved by this project?

This wrapper automatically overcomes Google Gemini API's free tier limits and constraints (as of Aug 2026):

  • Audio Length Limit — depends on whether you ask for speaker diarization:
    • --no-diarize (default): up to ~60 min per API call. The wrapper cuts the file into 59-min logical units, each sent as one API call.
    • --diarize: up to ~30 min per API call. The wrapper cuts into 29-min chunks to stay under the limit with a 1-min safety margin.
  • Rate Limit (Max 2 RPM): Applies a built-in delay (default 30s) between sequential requests to prevent rate-limit errors.
  • Daily Quota (Max 25 RPD or input-token quota): Tracks daily Pacific-time API usage locally (~/.cache/gemini-transcribe-wrapper/usage-<sha256(key)[:12]>.json) and warns clearly when limits approach.
  • Raw Response to Ready-to-Use Subtitles: Converts AI transcription output directly into .srt, .txt, and (with --diarize) .diarized.srt files in a single run.
  • Free-Tier-Friendly Defaults: --no-diarize is the default to minimize API calls; the wrapper packs each call to the per-mode maximum.
  • 429 = Stop the Batch: On HTTP 429 (rate limit or quota exhausted), the wrapper prints retry suggestions and aborts the batch immediately. The CLI exits with code 2 to distinguish quota errors from other failures (code 1).

On Quota / Rate-Limit (HTTP 429)

The wrapper does not retry automatically on 429. When a 429 is detected it:

  1. Logs the raw error and the quota category (free-tier daily quota vs short-term rate limit).
  2. Prints retry suggestions: "wait about 1 minute for a short-term 429, or wait Xh Ym (sleep Zs) until PST midnight for the daily quota to reset, then re-run."
  3. Aborts the rest of the batch (no point burning more calls on a guaranteed 429).
  4. Returns exit code 2.

Sample log on a quota hit:

ERROR Rate limit / quota exceeded (429): Error code: 429 - You exceeded your current quota ...
ERROR It looks like you hit the free tier daily quota (25 calls/day).
ERROR You hit the Gemini API rate limits:
  - max 2 API calls per minute
  - max 30 minutes of audio per call
  - max 25 API calls per day (free tier)
ERROR To retry: wait about 1 minute for a short-term 429, or wait 5h 40m (sleep 20400s) until PST midnight for the daily quota to reset, then re-run.
ERROR Switching to a paid tier (enable billing) removes the free-tier limits.
ERROR Aborting batch: quota / rate limit hit while processing <file>. Remaining files will not be processed.

Batch Bulk Transcribing Tip

The free tier caps 25 calls/day per PST day. For multi-day batches, either:

  • Run in shorter bursts so you naturally pause before hitting the quota, or
  • Swap in a fresh API key (each key has its own counter under ~/.cache/gemini-transcribe-wrapper/usage-<sha256(key)[:12]>.json), or
  • Wait for the PST reset — on a 429 the wrapper prints the exact sleep seconds needed; wrap your batch in a shell loop:
# Run, on quota hit wait until PST midnight, then re-run.
while ! gtw '*.mp4' --diarize; do
  echo "Batch hit a quota; waiting for PST reset..."
  # Read the sleep seconds from the previous log, or sleep 1h and retry.
  sleep 3600
done

Tip: --no-diarize halves (or better) the number of API calls vs --diarize for the same audio, since each call covers ~2× the audio length.

Diarizing Tip (Speaker Labels)

--diarize produces .diarized.srt with raw speaker ids (spk:0, spk:1, ...). Map them to real names one pass at a time with --speakers: delete .diarized.srt, re-run with a more complete map, repeat until no spk: string remains.

# Pass 0: no --speakers — every cue keeps its raw spk:# tag.
gtw --diarize interview.mp4

# Wrapper logs the unmapped speakers and prints the re-render recipe:
# WARNING Some speakers are not covered by --speakers mapping.
#   Speaker map: spk:0 spk:1 spk:2
#   Unmapped: spk:0, spk:1, spk:2
# WARNING To re-render with names, delete the .diarized.srt and re-run with the
#   option, editing the Name# entries: rm 'interview.diarized.srt' &&
#   gtw 'interview.mp4' --diarize --speakers 'spk:0=Name0;spk:1=Name1;spk:2=Name2;'

# Pass 1: rename spk:0 → Host. Edit the recipe, delete the old .diarized.srt, re-run.
rm interview.diarized.srt
gtw --diarize interview.mp4 --speakers 'spk:0=Host;'

# Pass 2: add spk:1 → Guest. Same drill.
rm interview.diarized.srt
gtw --diarize interview.mp4 --speakers 'spk:0=Host;spk:1=Guest;'

# Pass 3: add spk:2 → Interpreter. Now every tag is a real name.
rm interview.diarized.srt
gtw --diarize interview.mp4 --speakers 'spk:0=Host;spk:1=Guest;spk:2=Interpreter;'

Simulation of the iterative renaming (.diarized.srt excerpt after each pass):

# Pass 0 (no --speakers): all raw ids
[spk:0] Hello everyone, welcome to the show.
[spk:1] Today we have a special guest with us.
[spk:2] Thanks for having me, it's great to be here.

# Pass 1 (--speakers 'spk:0=Host;')
[Host]     Hello everyone, welcome to the show.
[spk:1]    Today we have a special guest with us.
[spk:2]    Thanks for having me, it's great to be here.

# Pass 2 (add spk:1=Guest)
[Host]     Hello everyone, welcome to the show.
[Guest]    Today we have a special guest with us.
[spk:2]    Thanks for having me, it's great to be here.

# Pass 3 (add spk:2=Interpreter) — done, no spk: left
[Host]       Hello everyone, welcome to the show.
[Guest]      Today we have a special guest with us.
[Interpreter] Thanks for having me, it's great to be here.

Tip: each pass is a no-API re-render — the wrapper reads <base>.diarized.transcript.json (kept by default) and rewrites .diarized.srt from it, so iterating is fast and free.

Note: --speakers is ignored when --diarize is off (the default). Pass --diarize to enable speaker mapping.

Relevant Repositories

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

gemini_transcribe_wrapper-0.0.18.tar.gz (93.4 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

gemini_transcribe_wrapper-0.0.18-py3-none-any.whl (37.1 kB view details)

Uploaded Python 3

File details

Details for the file gemini_transcribe_wrapper-0.0.18.tar.gz.

File metadata

  • Download URL: gemini_transcribe_wrapper-0.0.18.tar.gz
  • Upload date:
  • Size: 93.4 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.8 {"installer":{"name":"uv","version":"0.11.8","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for gemini_transcribe_wrapper-0.0.18.tar.gz
Algorithm Hash digest
SHA256 7c9fb11343ea5bd107f0d062470089ffe3fbc824b4ee8434bb37d96d9d59ff87
MD5 5e53d3ee10280895ba725bfd3c1c6823
BLAKE2b-256 0887a6050fd5fbc755ab1123a0428e8fd7a62eeb7e62a308ead0359ea98b237f

See more details on using hashes here.

File details

Details for the file gemini_transcribe_wrapper-0.0.18-py3-none-any.whl.

File metadata

  • Download URL: gemini_transcribe_wrapper-0.0.18-py3-none-any.whl
  • Upload date:
  • Size: 37.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.8 {"installer":{"name":"uv","version":"0.11.8","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for gemini_transcribe_wrapper-0.0.18-py3-none-any.whl
Algorithm Hash digest
SHA256 0ecffbe8c18d4bfe774ef4001abb4e4b8c5fe6e60fa831b332749785f67c5098
MD5 2d5eff4b98571a3ebc6f3e87336bf398
BLAKE2b-256 879f6003b3f4e99547c50040bfb52d4d68d5729f9a24b01bb8b42df8d2b7b8c5

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

0.0.18 This release

2 files

0.0.17

2 files

0.0.16

2 files

0.0.15

2 files

0.0.14

2 files

0.0.13

2 files

0.0.12

2 files

0.0.11

2 files

0.0.10

2 files

0.0.9

2 files

0.0.8

2 files

0.0.7

2 files

0.0.6

2 files

0.0.5

2 files

0.0.4

2 files

0.0.3

2 files

0.0.2

2 files

0.0.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page