Skip to main content

Local MCP server that transcribes audio files with OpenAI — point it at a path, get the text back

Project description

jackai-stt-mcp

tests PyPI License: MIT

Transcribe audio by asking for it in plain language:

You: transcribe ~/Downloads/voice.opus

Claude: تمام تمام، موضوع اللغة العربية إن شاء الله محلول…

An MCP server that runs on your own machine, so the assistant can open the file directly. Nothing to run, no commands to memorise — you mention a file and the assistant does the rest.

No uploading, no copying, no base64. Your OpenAI key never leaves your computer.

Arabic works well, including dialect.

Install

Add this to your MCP client's config. uvx fetches and runs the package on first use — nothing to install by hand.

{
  "servers": {
    "jackai-stt": {
      "command": "uvx",
      "args": ["jackai-stt-mcp"],
      "env": {
        "OPENAI_API_KEY": "sk-proj-..."
      }
    }
  }
}

Where the config lives:

Client File
VS Code ~/Library/Application Support/Code/User/mcp.json (macOS)
Claude Desktop ~/Library/Application Support/Claude/claude_desktop_config.json
Claude Code .mcp.json in the project, or claude mcp add
Cursor ~/.cursor/mcp.json

Restart the client, then ask it to transcribe something.

Get an API key at platform.openai.com/api-keys. The account needs credit; transcription runs about $0.0045 per minute.

Usage

Nothing to run, no syntax to learn. Talk to the assistant the way you normally would and mention the file:

transcribe ~/Downloads/voice.opus

what does the voice note on my desktop say?

read ./recordings/call.m4a and summarise what the client is asking for

transcribe meeting.mp3 and tell me who said what

The assistant recognises that it needs this tool and fills in the arguments from what you said. You never call it yourself or write out its parameters.

That last example needs speaker labels, which only one model produces. Mention it in passing and the assistant picks it:

transcribe meeting.mp3 with the diarize model

Same for a language it keeps mishearing ("it's Egyptian Arabic"), or names it should spell correctly ("the speakers are Ahmad and Sara"). Ordinary words — the table below is just what those words map onto.

What the assistant fills in

You don't set these by hand; this is here so you know what it can control.

Argument Default Notes
file_path Path on this machine. Absolute, relative, or ~/....
audio_url Public http(s) URL, downloaded then transcribed.
audio_base64 For short clips. Prefer file_path.
model gpt-transcribe See the table below.
language auto-detect ISO-639-1 (ar, en). Set only if detection is wrong.
prompt Names or jargon likely in the audio, to steer spelling.

Exactly one audio source per request.

Models

Model Cost/min Speaker labels
gpt-transcribe (default) $0.0045 no
gpt-4o-mini-transcribe $0.003 no
gpt-4o-transcribe $0.006 no
gpt-4o-transcribe-diarize $0.006 yes
whisper-1 $0.006 no

gpt-transcribe is OpenAI's recommended model: cheaper than whisper-1 and more accurate. There is no reason to pick whisper-1 unless you need it specifically.

Formats

flac m4a mp3 mp4 mpeg mpga oga ogg wav webm — up to 25 MB.

.opus files work too. OpenAI rejects that extension even though the bytes are what it accepts as .ogg, so this server renames it in flight. WhatsApp voice notes are all .opus, which is exactly the case that would otherwise fail.

Why it runs locally

A remote MCP server cannot read a file you attached in chat. There is no mechanism in the protocol that carries attachments to a server — it has been an open issue since early 2025, and the proposals to add one are still drafts. Base64 through a tool argument works in theory but dies on client argument-size limits after a second or two of audio.

Running on your machine avoids all of it. The server is a process you own, reading a file you own, with a key you hold.

Development

git clone https://github.com/jack-ai-net/stt-mcp
cd stt-mcp
python3 -m venv .venv && .venv/bin/pip install -e ".[dev]"
.venv/bin/python -m pytest

Tests stub the OpenAI call, so they need no API key and cost nothing.

License

MIT — see LICENSE.

Built by JackAI.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

jackai_stt_mcp-0.1.1.tar.gz (11.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

jackai_stt_mcp-0.1.1-py3-none-any.whl (9.3 kB view details)

Uploaded Python 3

File details

Details for the file jackai_stt_mcp-0.1.1.tar.gz.

File metadata

  • Download URL: jackai_stt_mcp-0.1.1.tar.gz
  • Upload date:
  • Size: 11.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.14.6

File hashes

Hashes for jackai_stt_mcp-0.1.1.tar.gz
Algorithm Hash digest
SHA256 da4e490214b4871a3b8bb862b2eb27ea3297041ac2730644063a0d4e2a33d20f
MD5 50707edda32eac413f9db1ae31cec748
BLAKE2b-256 f8d9d0a53f7acb8c5d4ddc85da685a1f83f72e5b1bbaabe7d3d669a2c78077aa

See more details on using hashes here.

File details

Details for the file jackai_stt_mcp-0.1.1-py3-none-any.whl.

File metadata

  • Download URL: jackai_stt_mcp-0.1.1-py3-none-any.whl
  • Upload date:
  • Size: 9.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.14.6

File hashes

Hashes for jackai_stt_mcp-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 514a0622b7c0de0034d4b874ca1d5a0f02dab0c2a1d164da1bf7f41c17ebc597
MD5 0492217e74e6667c97039a928f7fdfa4
BLAKE2b-256 610c09691551f46c87e2865f25dc18aef7a6f1edfb024feb226ae401287051ea

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page