Skip to main content

voxmd

PyPI Python 3.13+ License: MIT Platform: macOS | Linux

Turn voice memos into linked Obsidian notes, entirely on your own machine. voxmd transcribes a recording, pulls out a summary, key points, decisions, actions, people and topics, and writes a note into your vault. Run it on one file, or let it watch the folder your phone syncs into.

Setup

Works on macOS and Linux (not Windows). The commands use Homebrew; on Linux, install ffmpeg, whisper.cpp and Ollama with your package manager.

  1. Install voxmd and its tools. voxmd installs with uv, which also fetches Python 3.13 if you don't have it.

    uv tool install voxmd
    brew install ffmpeg whisper-cpp
    
  2. Download the speech model (1.6 GB).

    mkdir -p ~/.local/share/voxmd/models
    curl -L -o ~/.local/share/voxmd/models/ggml-large-v3-turbo.bin \
      https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-large-v3-turbo.bin
    
  3. Pull a language model. Start Ollama first (open the app, or run ollama serve).

    ollama pull qwen3:8b
    
  4. Create your config, then set at least whisper.model and vault.path (and watch.dir to use the watcher).

    mkdir -p ~/.config/voxmd
    curl -L -o ~/.config/voxmd/config.yaml \
      https://raw.githubusercontent.com/rvnztolentino/voxmd/main/voxmd.example.yaml
    
  5. Check everything. Every line should say ok; anything that fails tells you how to fix it. It changes nothing.

    voxmd doctor
    

Usage

One recording:

voxmd process memo.m4a

Prints the path of the new note. Useful flags: --model (another Ollama model), --vault PATH, --date, --no-archive (leave the audio where it is), --force, and -v for timings.

A folder, automatically:

voxmd watch      # runs until you press Ctrl-C or close the terminal
voxmd status     # in another terminal: running?, last activity, notes made today

The watcher waits for each file to finish syncing, handles one memo at a time, picks up anything already in the folder when it starts, and logs everything it does to ~/.local/state/voxmd/voxmd.log.

Other commands: voxmd models lists your Ollama models. transcribe, extract and render run one step at a time and pipe into each other:

voxmd transcribe memo.m4a | voxmd extract | voxmd render --source memo.m4a

Configuration

One YAML file, read from $VOXMD_CONFIG, then ./voxmd.yaml, then ~/.config/voxmd/config.yaml (first found wins). voxmd.example.yaml documents every setting.

whisper:
  model: ~/.local/share/voxmd/models/ggml-large-v3-turbo.bin
  language: auto                    # or a code like en, tl

ollama:
  model: qwen3:8b                   # any model you've pulled

vault:
  path: ~/Documents/Obsidian/Main   # must already exist
  folder: Voice memos
  transcripts: true                 # save each transcript as a linked note
  duplicates: copy                  # same audio again: copy (Title 2.md) or skip

archive:
  dir: ~/Documents/Voice memos archive   # omit to leave recordings in place

watch:
  dir: ~/Documents/Voice memos inbox     # must exist; keep the archive outside it

Paths must be absolute or start with ~, and unknown keys are rejected, so a typo shows up as an error.

Choosing a model

Pick a language model by your machine's memory, then set ollama.model or pass --model:

  • 8 GB: qwen3:4b or llama3.2:3b. Fast, but less accurate.
  • 16 GB: qwen3:8b (default) or qwen2.5:7b.
  • 32 GB+: qwen3:14b or gemma3:12b. Better notes, slower to load.

Smaller models make more mistakes and are more easily misled by instructions spoken inside a memo. The model must support Ollama's structured output.

For transcription, whisper.model can be any ggml file from the same Hugging Face repo, from ggml-base.bin (142 MB, rougher) to the default ggml-large-v3-turbo.bin (1.6 GB, most accurate). Whisper understands about 99 languages and detects each memo's language by default.

What a note looks like

  • File name: date and title, like 2026-09-17 Launch Plan.md, dated when the memo was recorded.
  • Contents: a summary (longer for longer recordings), key points, decisions, and actions as checkboxes.
  • Links: people and topics in entities.json become [[links]], and new ones are added for next time.
  • Transcript: the full transcript is saved in Transcripts/ and linked from the note.
  • Repeats: processing the same audio again creates Title 2.md and never changes the first note.
  • Templates: the layout comes from note.md.j2. Copy it and set render.template to customise.

Good to know

  • Local only. voxmd talks to nothing but Ollama on your own machine.
  • Nothing runs in the background. The watcher only runs while you have it open, and never starts on its own.
  • Nothing is overwritten or deleted. A recording is only archived after its note is saved.
  • It can make mistakes. Every note says so; check names and actions against the transcript or the recording.
  • Transcripts are private. Notes and transcripts are only readable by your user, but they're stored unencrypted. If your vault syncs, they sync too; set vault.transcripts: false to skip them.
  • Updating. uv tool upgrade voxmd. See the changelog for what changed.
  • No speaker labels. voxmd can't tell who said what in a meeting.
  • Exit codes: 0 done · 2 config · 3 missing tool · 4 bad input · 5 tool failed · 6 timeout · 7 note not written · 8 note written, a later step failed.

Tech stack

Development

git clone https://github.com/rvnztolentino/voxmd.git
cd voxmd
uv sync
uv run voxmd --help
uv run pytest            # no models, tools or network needed
uv run ruff check
uv run ruff format --check

License

MIT · Changelog

Metadata

Release files for voxmd 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for voxmd 0.1.0
File Size Uploaded
voxmd-0.1.0.tar.gz 108.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for voxmd 0.1.0
File Interpreter ABI Platform
voxmd-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 181.9 kB

Release files / voxmd-0.1.0.tar.gz

Download URL voxmd-0.1.0.tar.gz
Size 108.2 kB
Tags Source
SHA-256 checksum
How to use checksums
9e7f5a98b09f768ed64e1c2639eac741b2664d22043855ddcd547e2c8fcc6d9a
BLAKE2b-256 checksum
How to use checksums
539770fcafec3ce880f81bd579b2f1b837c0a69739bbe48764af17f6446f8134
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.

Transparency log

Release files / voxmd-0.1.0-py3-none-any.whl

Download URL voxmd-0.1.0-py3-none-any.whl
Size 73.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3f2d778faffbdf0142f80fb6f7f6c812941c95a3595c964e76acb2937416ce78
BLAKE2b-256 checksum
How to use checksums
c8b75656c42ebb6054e8238b26e55a07cc2b366a54a216994443e7eb4f7292b1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page