Skip to main content

transcribe-it

A lightweight CLI for ingesting meeting transcripts (Gmail or Slack), enriching them with an LLM, and storing the results as local files.

Source -> Extract -> LLM Enrich -> Local Files

Prerequisites

  • Python 3.12+
  • A Google Cloud project with Gmail API + Google Drive API enabled (for the Gmail source), or a Slack bot token (for the Slack source)
  • An API key for one of the supported LLM providers (Anthropic, OpenAI, or Groq) — only needed if you want LLM enrichment

Install

uv tool install transcribe-it

Or with pipx:

pipx install transcribe-it

Setup

Setup is split into two steps: a one-off global step for credentials, and a per-project step for what to ingest.

Step 1: Configure credentials (once per machine)

transcribe-it setup

Pick which credentials to set up — Google OAuth (for Gmail), Slack bot token, and/or LLM provider — and the values are written to ~/.config/transcript/env. Re-run any time to add or rotate values; existing values are preserved unless you confirm overwrite (or pass --force).

Step 2: Initialise a project (per directory)

From the directory where you want transcripts to land:

transcribe-it init

This asks which sources to enable, source-specific config (sender filter, subject filter, channel ID, etc.), output path, and lookback window. Writes .transcripts/config.yaml. No secrets prompts — it'll warn if the credentials a chosen source needs aren't set yet.

Gmail credentials

setup asks for GOOGLE_OAUTH_CLIENT_ID and GOOGLE_OAUTH_CLIENT_SECRET. Two options:

  1. Reuse someone else's OAuth client — ask a teammate for the values and have them add your Google account as a Test user on their OAuth consent screen.
  2. Create your own — in Google Cloud Console, create an OAuth 2.0 Client ID of type Desktop app, then copy the client ID and secret from the resulting credentials.

After setup and init, authenticate:

transcribe-it auth gmail

Slack credentials

setup asks for SLACK_BOT_TOKEN (xoxb-...). The bot needs to be a member of the channels you want to ingest from. The channel ID itself is configured per-project in init.

Usage

By default, ingestion only extracts the raw transcript — no LLM call, no API key required. Pass --enrich to also generate a summary, topics, and participants via LLM.

# Last N days, raw extraction only (default)
transcribe-it ingest gmail --days 7

# With LLM enrichment
transcribe-it ingest gmail --days 7 --enrich

# Enrichment + cleaned transcript variant (--clean implies --enrich)
transcribe-it ingest gmail --days 7 --clean

# Specific date range
transcribe-it ingest gmail --from 2026-04-01 --to 2026-04-05

# Preview matching emails without fetching or writing
transcribe-it ingest gmail --days 1 --dry-run

# Ingest a single transcript file directly
transcribe-it ingest file path/to/transcript.txt

Output

Raw mode (default) writes a single .txt file per transcript:

transcripts/
  2026-04-09-ai-labs-daily.txt

With --enrich, each transcript becomes a folder:

transcripts/
  2026-04-09-ai-labs-daily/
    raw.txt          # Original transcript (immutable)
    metadata.json    # Source, date, participants, topics, summary

With --clean, an additional clean.md is written (structured: title, summary, topics, cleaned transcript).

Prompts

LLM prompts are bundled with the package under transcribe_it/prompts/. To customise, fork the repo and edit prompts/enrich.md.

Commands

Command Description
transcribe-it setup Configure global credentials (OAuth, LLM, Slack token)
transcribe-it init Initialise project config (sources, output path, lookback)
transcribe-it auth gmail Authenticate with Gmail (OAuth); --profile NAME for a second account
transcribe-it ingest gmail Ingest transcripts from Gmail
transcribe-it ingest file PATH Ingest a single transcript file

Ingest options (Gmail)

Flag Description
--days N How many days back to search
--from YYYY-MM-DD Start date
--to YYYY-MM-DD End date
--subject TEXT Only emails whose subject matches TEXT; repeatable, */? wildcards (overrides config)
--profile NAME Restrict the run to one Gmail account (default: all configured)
--dry-run List matching emails without processing
--enrich Run LLM enrichment (summary, topics, participants)
--clean Also generate a cleaned version of the transcript (implies --enrich)

Subject filters

By default every email from the configured sender is ingested. To narrow it down, set subject under sources.gmail — a single string or a list:

sources:
  gmail:
    sender: gemini-notes@google.com
    subject:
      - AI Labs
      - Weekly*Review

A pattern with no wildcards matches anywhere in the subject, so AI Labs matches "Notes: AI Labs daily". A pattern containing * or ? is matched against the whole subject instead, so Weekly*Review matches "Weekly Team Review" but not "Notes: Weekly Team Review" — write *Weekly*Review* if you want both. Matching is case-insensitive, and an email is kept if any pattern matches.

--subject overrides the config for one run and can be repeated:

transcribe-it ingest gmail --subject "AI Labs" --subject "Standup*"

Multiple Gmail accounts

Each account is a profile — a name for its own OAuth token, stored at ~/.config/transcript/credentials/<profile>.json. To pull from more than one account into the same project, list them under sources.gmail, each with its own sender and subject filters:

sources:
  gmail:
    - profile: personal
      sender: gemini-notes@google.com
    - profile: work
      sender: notes@zoom.us
      subject:
        - Weekly*Review

Authenticate each one separately:

transcribe-it auth gmail --profile personal
transcribe-it auth gmail --profile work

transcribe-it ingest gmail then searches every configured account and writes the results to the same output directory; a transcript that reaches both accounts is only ingested once. --profile NAME restricts the run to one account.

Both profiles use the OAuth client from ~/.config/transcript/env, so each Google account must be a Test user on that OAuth consent screen.

The single-account form stays valid:

sources:
  gmail:
    profile: default
    sender: gemini-notes@google.com

Configuration files

Path Purpose
.transcripts/config.yaml Per-project: sources, lookback, output destinations
~/.config/transcript/env Global: API keys and OAuth credentials

Metadata

Release files for transcribe-it 0.5.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for transcribe-it 0.5.1
File Size Uploaded
transcribe_it-0.5.1.tar.gz 142.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for transcribe-it 0.5.1
File Interpreter ABI Platform
transcribe_it-0.5.1-py3-none-any.whl Python 3 none any Details

Total release size: 163.5 kB

Release files / transcribe_it-0.5.1.tar.gz

Download URL transcribe_it-0.5.1.tar.gz
Size 142.1 kB
Tags Source
SHA-256 checksum
How to use checksums
821138e7a5e42ae4aa3153db6adc5da5d639decf15040fefc5155a4149fd7f64
BLAKE2b-256 checksum
How to use checksums
ac2a466be7e7ca45cd3887b2a0ce59ad99a4befee397249ae19329723efeed98
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / transcribe_it-0.5.1-py3-none-any.whl

Download URL transcribe_it-0.5.1-py3-none-any.whl
Size 21.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
cc0a2a4fb2d55bf1958083b9b721b578dcb7f5cbf9ba7a42534abb5d83a40da6
BLAKE2b-256 checksum
How to use checksums
0961a64d05c1731a9b5856db3815ba29a68ec36e6df78ca1f7d0d9d538e31ced
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

This release

0.5.1 This release

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page