Local-first meeting recorder with AI transcription and speaker diarization

These details have not been verified by PyPI

Project description

Bee Recorder

Local-first meeting recorder for Linux. Captures system audio (Google Meet, Teams, Zoom, Discord, etc.) and your microphone in parallel, then transcribes everything with Whisper and identifies who said what with SpeechBrain — all on your machine. No audio leaves your computer, no API keys, no HuggingFace token.

The product is the transcript itself: a Markdown file you can read and a JSON file you can feed to any LLM later for analysis.

Features

Fully local — Whisper for transcription, SpeechBrain ECAPA-TDNN for speaker diarization. Nothing uploaded.
Zero-friction setup — bee setup detects your hardware (RAM, GPU) and picks a Whisper model that will run well, then downloads it (~1.6 GB).
Channel-aware diarization — microphone audio is diarized into Mic_00, Mic_01, ... (works for in-person meetings where multiple people share one mic), and remote participants from system audio are diarized into Speaker_00, Speaker_01, ...
Interactive labeling — after processing, Bee shows representative quotes per speaker so you can match Mic_00/Speaker_00 to real names.
Two outputs per recording — <id>.md (human-readable) and <id>.json (structured for automation/LLM consumption).
Auto language detection — works out of the box in any language Whisper supports.

Requirements

Linux with PulseAudio or PipeWire (Ubuntu, Fedora, Debian, Arch, Mint, Pop!_OS, etc.)
Python 3.10+
ffmpeg and pulseaudio-utils (for pactl)

# Ubuntu / Debian
sudo apt install ffmpeg pulseaudio-utils python3 python3-venv

# Fedora
sudo dnf install ffmpeg-free pulseaudio-utils python3

# Arch
sudo pacman -S ffmpeg libpulse python

macOS and Windows are not supported today (the audio capture layer relies on PulseAudio/PipeWire).

Install

pipx install bee-recorder
bee setup       # detects hardware, downloads Whisper + SpeechBrain models
bee doctor      # verifies ffmpeg, audio server, models, config

Or with plain pip inside a venv:

python3 -m venv ~/.venvs/bee
~/.venvs/bee/bin/pip install bee-recorder
~/.venvs/bee/bin/bee setup

Recording a meeting

1. Start the recording before joining

bee start -n weekly-2026-05-08

The name is optional — without -n, Bee uses a timestamp ID. Add --no-mic if you want to capture only the system audio.

2. Join the meeting normally

Bee runs in the background and captures whatever your audio output and microphone produce. It auto-detects new audio sinks during the call (e.g. you plug in AirPods mid-meeting), so you don't need to restart anything.

3. Stop and process

bee stop

This stops the capture and runs:

Whisper transcription (mic + system audio)
SpeechBrain diarization on both audio streams (mic + system)
Merging — mic speakers become Mic_XX, remote participants become Speaker_XX

How long it takes depends on the meeting length and your hardware. With a CPU and the medium model, expect roughly 0.5–1× real time (a 30-minute meeting takes 15–30 minutes to process). With a GPU, much faster.

If you want to stop now and process later:

bee stop --no-process
bee process weekly-2026-05-08    # transcribe + diarize on demand

4. Label the participants

After processing, Bee prints a table of representative quotes per speaker. To re-print it any time:

bee speakers weekly-2026-05-08

Then assign real names:

bee label weekly-2026-05-08 Mic_00="Alice" Speaker_00="Bob (Acme Corp)" Speaker_01="Carol (Acme Corp)"

The label rewrites both the JSON and Markdown transcripts in place.

5. Read or share the transcript

bee show weekly-2026-05-08       # prints metadata + transcript
bee show weekly-2026-05-08 -t    # transcript only
bee list                          # all recordings

Output files

Transcripts are written to ~/Documents/bee/transcripts/ by default (configurable):

File	Contents
`<id>.md`	Human-readable transcript with timestamps and speaker labels
`<id>.json`	Structured transcript with metadata header (duration, language, speakers, segments) for automation and LLM consumption

Raw audio (mic.wav, system.wav) and a metadata.json are kept under ~/Documents/bee/recordings/<id>/ so you can re-process a recording any time with bee process <id>.

Optional: AI summary

Bee does not generate summaries by default — the transcript is the deliverable. If you want a quick summary, the summarize command shells out to the Claude Code CLI (must be installed and authenticated separately):

bee summarize ~/Documents/bee/transcripts/weekly-2026-05-08.md

This produces weekly-2026-05-08_summary.md next to the transcript.

For richer analyses (action items, sentiment, decisions), feed the JSON transcript into your LLM of choice. The structured format is designed for that.

Command reference

Command	What it does
`bee start -n <name>`	Start recording
`bee start --no-mic`	Capture system audio only
`bee stop`	Stop and process (transcribe + diarize + ask to label)
`bee stop --no-process`	Just stop the recording
`bee stop --no-diarize`	Skip diarization (faster, no speaker labels)
`bee stop --no-identify`	Don't prompt for participant names interactively
`bee process <id>`	(Re)process a stopped recording
`bee speakers <id>`	Show preview quotes per detected speaker
`bee label <id> SPEAKER_00=Name ...`	Rename speakers in a transcript
`bee show <id>`	Print transcript and metadata
`bee list`	List all recordings
`bee summarize <file>`	Generate a Claude summary (requires Claude CLI)
`bee watch`	Watch the transcripts folder and auto-summarize new files
`bee setup`	First-time setup: hardware detection + model download
`bee doctor`	Health check (ffmpeg, audio server, models, config)
`bee models list`	Show available and installed models
`bee models download <name>`	Download a specific Whisper or diarization model
`bee models usage`	Show disk usage of downloaded models

Configuration

Bee stores its config at ~/.config/bee/config.toml. Defaults are written by bee setup:

[whisper]
model = "medium"        # tiny, base, small, medium, large-v3
device = "auto"         # auto, cpu, cuda
language = "auto"       # auto-detect, or "pt", "en", "es", ...

[audio]
sample_rate = 16000
channels = 1

You can change the model later (bee models download large-v3 and edit the config), or change the language to skip detection. Run bee doctor after editing.

Privacy

Everything runs locally. Audio files, transcripts, and models all live on your disk. The only network traffic is the initial model download from HuggingFace (anonymous, no token required) when you run bee setup.

License

MIT

Project details

These details have not been verified by PyPI

Release history Release notifications | RSS feed

0.1.0b9 pre-release

May 14, 2026

0.1.0b8 pre-release

May 14, 2026

This version

0.1.0b7 pre-release

May 14, 2026

0.1.0b6 pre-release

May 14, 2026

0.1.0b5 pre-release

May 14, 2026

0.1.0b4 pre-release

May 13, 2026

0.1.0b3 pre-release

May 8, 2026

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

bee_recorder-0.1.0b7.tar.gz (33.0 kB view details)

Uploaded May 14, 2026 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

bee_recorder-0.1.0b7-py3-none-any.whl (35.0 kB view details)

Uploaded May 14, 2026 Python 3

File details

Details for the file bee_recorder-0.1.0b7.tar.gz.

File metadata

Download URL: bee_recorder-0.1.0b7.tar.gz
Upload date: May 14, 2026
Size: 33.0 kB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/6.2.0 CPython/3.10.12

File hashes

Hashes for bee_recorder-0.1.0b7.tar.gz
Algorithm	Hash digest
SHA256	`49b868ae172165434fbc6badfd4343922662e190d2eb340c17cbad3862da77ec`
MD5	`ff02be72595d3922fcff834f05297951`
BLAKE2b-256	`ac00574b4b418a7eb1173cea78170d1872dfb480d8d23da197b1f2e15263fed2`

See more details on using hashes here.

File details

Details for the file bee_recorder-0.1.0b7-py3-none-any.whl.

File metadata

Download URL: bee_recorder-0.1.0b7-py3-none-any.whl
Upload date: May 14, 2026
Size: 35.0 kB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: twine/6.2.0 CPython/3.10.12

File hashes

Hashes for bee_recorder-0.1.0b7-py3-none-any.whl
Algorithm	Hash digest
SHA256	`6408b662559853e2774824cab242f50b954953d50f6d14d54c3d91170db93e12`
MD5	`429a896c374069149ed1d1424dc06675`
BLAKE2b-256	`678b32b94dc31b5a06108eaaac748574713c0569508d451c361b31e19ef2eba9`

See more details on using hashes here.

bee-recorder 0.1.0b7

Navigation

Verified details

Maintainers

Unverified details

Meta

Classifiers

Project description

Bee Recorder

Features

Requirements

Install

Recording a meeting

1. Start the recording before joining

2. Join the meeting normally

3. Stop and process

4. Label the participants

5. Read or share the transcript

Output files

Optional: AI summary

Command reference

Configuration

Privacy

License

Project details

Verified details

Maintainers

Unverified details

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes