voxmd
Turn voice memos into linked Obsidian notes, entirely on your own machine. voxmd transcribes a recording, pulls out a summary, key points, decisions, actions, people and topics, and writes a note into your vault. Run it on one file, or let it watch the folder your phone syncs into.
Setup
Works on macOS and Linux (not Windows). The commands use Homebrew; on Linux, install ffmpeg, whisper.cpp and Ollama with your package manager.
-
Install voxmd and its tools. voxmd installs with uv, which also fetches Python 3.13 if you don't have it.
uv tool install voxmd brew install ffmpeg whisper-cpp
-
Download the speech model (1.6 GB).
mkdir -p ~/.local/share/voxmd/models curl -L -o ~/.local/share/voxmd/models/ggml-large-v3-turbo.bin \ https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-large-v3-turbo.bin
-
Pull a language model. Start Ollama first (open the app, or run
ollama serve).ollama pull qwen3:8b
-
Create your config, then set at least
whisper.modelandvault.path(andwatch.dirto use the watcher).mkdir -p ~/.config/voxmd curl -L -o ~/.config/voxmd/config.yaml \ https://raw.githubusercontent.com/rvnztolentino/voxmd/main/voxmd.example.yaml
-
Check everything. Every line should say
ok; anything that fails tells you how to fix it. It changes nothing.voxmd doctor
Usage
One recording:
voxmd process memo.m4a
Prints the path of the new note. Useful flags: --model (another Ollama model), --vault PATH, --date, --no-archive (leave the audio where it is), --force, and -v for timings.
A folder, automatically:
voxmd watch # runs until you press Ctrl-C or close the terminal
voxmd status # in another terminal: running?, last activity, notes made today
The watcher waits for each file to finish syncing, handles one memo at a time, picks up anything already in the folder when it starts, and logs everything it does to ~/.local/state/voxmd/voxmd.log.
Other commands: voxmd models lists your Ollama models. transcribe, extract and render run one step at a time and pipe into each other:
voxmd transcribe memo.m4a | voxmd extract | voxmd render --source memo.m4a
Configuration
One YAML file, read from $VOXMD_CONFIG, then ./voxmd.yaml, then ~/.config/voxmd/config.yaml (first found wins). voxmd.example.yaml documents every setting.
whisper:
model: ~/.local/share/voxmd/models/ggml-large-v3-turbo.bin
language: auto # or a code like en, tl
ollama:
model: qwen3:8b # any model you've pulled
vault:
path: ~/Documents/Obsidian/Main # must already exist
folder: Voice memos
transcripts: true # save each transcript as a linked note
duplicates: copy # same audio again: copy (Title 2.md) or skip
archive:
dir: ~/Documents/Voice memos archive # omit to leave recordings in place
watch:
dir: ~/Documents/Voice memos inbox # must exist; keep the archive outside it
Paths must be absolute or start with ~, and unknown keys are rejected, so a typo shows up as an error.
Choosing a model
Pick a language model by your machine's memory, then set ollama.model or pass --model:
- 8 GB:
qwen3:4borllama3.2:3b. Fast, but less accurate. - 16 GB:
qwen3:8b(default) orqwen2.5:7b. - 32 GB+:
qwen3:14borgemma3:12b. Better notes, slower to load.
Smaller models make more mistakes and are more easily misled by instructions spoken inside a memo. The model must support Ollama's structured output.
For transcription, whisper.model can be any ggml file from the same Hugging Face repo, from ggml-base.bin (142 MB, rougher) to the default ggml-large-v3-turbo.bin (1.6 GB, most accurate). Whisper understands about 99 languages and detects each memo's language by default.
What a note looks like
- File name: date and title, like
2026-09-17 Launch Plan.md, dated when the memo was recorded. - Contents: a summary (longer for longer recordings), key points, decisions, and actions as checkboxes.
- Links: people and topics in
entities.jsonbecome[[links]], and new ones are added for next time. - Transcript: the full transcript is saved in
Transcripts/and linked from the note. - Repeats: processing the same audio again creates
Title 2.mdand never changes the first note. - Templates: the layout comes from
note.md.j2. Copy it and setrender.templateto customise.
Good to know
- Local only. voxmd talks to nothing but Ollama on your own machine.
- Nothing runs in the background. The watcher only runs while you have it open, and never starts on its own.
- Nothing is overwritten or deleted. A recording is only archived after its note is saved.
- It can make mistakes. Every note says so; check names and actions against the transcript or the recording.
- Transcripts are private. Notes and transcripts are only readable by your user, but they're stored unencrypted. If your vault syncs, they sync too; set
vault.transcripts: falseto skip them. - Updating.
uv tool upgrade voxmd. See the changelog for what changed. - No speaker labels. voxmd can't tell who said what in a meeting.
- Exit codes:
0done ·2config ·3missing tool ·4bad input ·5tool failed ·6timeout ·7note not written ·8note written, a later step failed.
Tech stack
- Python 3.13+ with uv
- whisper.cpp for transcription, Ollama for extraction, ffmpeg for audio
- typer, pydantic, Jinja2, RapidFuzz, watchdog
Development
git clone https://github.com/rvnztolentino/voxmd.git
cd voxmd
uv sync
uv run voxmd --help
uv run pytest # no models, tools or network needed
uv run ruff check
uv run ruff format --check
License
Metadata
Release files for voxmd 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| voxmd-0.1.0.tar.gz | 108.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| voxmd-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 181.9 kB
Release files / voxmd-0.1.0.tar.gz
| Download URL | voxmd-0.1.0.tar.gz |
|---|---|
| Size | 108.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
9e7f5a98b09f768ed64e1c2639eac741b2664d22043855ddcd547e2c8fcc6d9a
|
|
BLAKE2b-256 checksum How to use checksums |
539770fcafec3ce880f81bd579b2f1b837c0a69739bbe48764af17f6446f8134
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.
Transparency logRelease files / voxmd-0.1.0-py3-none-any.whl
| Download URL | voxmd-0.1.0-py3-none-any.whl |
|---|---|
| Size | 73.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
3f2d778faffbdf0142f80fb6f7f6c812941c95a3595c964e76acb2937416ce78
|
|
BLAKE2b-256 checksum How to use checksums |
c8b75656c42ebb6054e8238b26e55a07cc2b366a54a216994443e7eb4f7292b1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.
Transparency log