MCP server for Sonilo AI music generation API
Project description
Sonilo MCP Server
An MCP (Model Context Protocol) server that exposes Sonilo's licensed music and sound-effects API to MCP-compatible clients (Claude Code, Claude Desktop, Codex).
The flagship tool is video_to_music: hand it your finished video and it composes an original soundtrack matched to the cut — the music follows the pacing, emotion, and edits because the model saw them. Length matches the video automatically. Every track is licensed and safe for commercial use (terms apply). text_to_music is also available for fixed-length tracks with no video to match.
For sound design, video_to_sfx watches your video and generates matching sound effects, returned as a standalone audio file. text_to_sfx generates a standalone effect from a description.
▶ Example result — an AI-generated trailer with its soundtrack composed by video_to_music from the assembled cut. For recipes covering any AI-video pipeline (stitch → grade → add music → mux), see the Sonilo video-to-music cookbook.
Quickstart with Claude Code
claude mcp add sonilo --env SONILO_API_KEY=sks_... -- uvx sonilo-mcp
Get your API key from the Sonilo dashboard, then start a session and ask, e.g. "Make background music that matches this video: ~/Desktop/promo.mp4."
Why Sonilo
- Video-to-music — give it a video and Sonilo composes a full-length score matched to its pacing, motion, and emotion. Transitions and beat drops align to your cut points, and the track matches the video's duration exactly — no prompts or manual syncing required.
- Text-to-music — generate tracks from a text description (genre, mood, tempo, instrumentation) at an exact duration (1–360s).
- Video-to-SFX — Sonilo watches the video and generates sound effects for what it sees. You get the SFX as a standalone audio file. Optional
segmentslet you script effects to specific time ranges ([{start, end, prompt}]). - Text-to-SFX — generate a standalone sound effect from a description (1–180s), in
wav,mp3,aac, orflac. - Fully licensed, commercial-safe — music licensed via Shutterstock; every generated track is cleared for commercial use on social, brand content, and advertising, with no Content ID worries.
- Video-to-sound — generate music and sound effects for the same clip in one call, mixed into a single balanced soundtrack. Get back the mixed audio, or a new video with it muxed in.
- Multiple variants per call —
text_to_music,video_to_music,video_to_video_music,video_to_sound, andvideo_to_video_soundacceptvariants_num(1–10, default 1): generate several distinct creative directions in one request instead of re-rolling one at a time. Cost scales linearly with N, and N > 1 is never covered by the free trial. - Pay as you go — billed only for the seconds of music you generate. Self-serve accounts start with free runs on every endpoint except
dubbing, no card required: 2 each on text-to-music, text-to-sfx and audio-ducking, and 1 each on video-to-music, video-to-sfx, video-to-video-music, video-to-video-sfx, video-to-sound and video-to-video-sound. After that, calls bill at the normal rate.dubbinghas zero free runs and is billed from the first call — it chargesvideo duration × number of languages, so a free run on it would be worth far more than on any other endpoint.
Audio Playback Dependencies
The play_audio tool requires PortAudio at runtime (for sounddevice). On macOS/Linux, install via:
- macOS:
brew install portaudio - Debian/Ubuntu:
sudo apt-get install libportaudio2
uvx sonilo-mcp and pip install will pull the Python bindings, but the system PortAudio library must be installed separately. The other tools (text_to_music, video_to_music, text_to_sfx, video_to_sfx, audio_ducking, get_sfx_task, get_account_services, get_usage) work without PortAudio.
Quickstart with Claude Desktop
-
Get your API key from the Sonilo dashboard.
-
Install the
uvpackage manager (providesuvx):curl -LsSf https://astral.sh/uv/install.sh | sh
See the uv repo for other install methods.
-
Go to Claude > Settings > Developer > Edit Config > claude_desktop_config.json to include the following:
{ "mcpServers": { "sonilo": { "command": "uvx", "args": ["sonilo-mcp"], "env": { "SONILO_API_KEY": "sks_...", "SONILO_API_URL": "https://api.sonilo.com", "TIME_OUT_SECONDS": "600" } } } }
-
Restart Claude Desktop. You should see the Sonilo tools available in the tool menu.
Quickstart with Codex
-
Get your API key from the Sonilo dashboard.
-
Install the
uvpackage manager (providesuvx):curl -LsSf https://astral.sh/uv/install.sh | sh
-
**Go to Codex > Settings > MCP servers to fill out the following:
Or you can add the server** to ~/.codex/config.toml:
[mcp_servers.sonilo]
command = "uvx"
args = ["sonilo-mcp"]
[mcp_servers.sonilo.env]
SONILO_API_KEY = "sk_..."
SONILO_API_URL = "https://api.sonilo.com"
TIME_OUT_SECONDS = "600"
- Restart Codex (or start a new session), then run
/mcpto confirmsonilois connected and its tools are listed.
Example usage
Once the server is connected, just ask your assistant in natural language. For example:
- "Make background music that matches this video:
~/Desktop/promo.mp4." - "Compose music for
https://example.com/clip.mp4with a calm, ambient style." - "I stitched my AI-generated clips into
~/Desktop/trailer.mp4— add a soundtrack that matches the cut." - "Use Sonilo mcp to generate 30 seconds of upbeat lo-fi hip-hop for a study playlist and save it to my Desktop."
- "Use Sonilo to write an epic orchestral cinematic track, about 60 seconds long."
- "What Sonilo services and limits does my account have?"
- "Show my Sonilo usage for the last 7 days."
- "Play the track you just generated."
The assistant will call the matching tool (text_to_music, video_to_music, text_to_sfx, video_to_sfx, audio_ducking, get_sfx_task, get_account_services, get_usage, or play_audio) and save generated audio to your configured output directory.
Configuration
Environment Variables
| Variable | Default | Description |
|---|---|---|
SONILO_API_KEY |
(required) | Bearer token. |
SONILO_API_URL |
https://api.sonilo.com |
Public API base URL. |
SONILO_MCP_BASE_PATH |
~/Desktop |
Default output directory and base for relative input paths. Also the confinement boundary (see below). |
SONILO_MCP_ALLOW_ANY_PATH |
false |
Set to true to let tools read/write files outside SONILO_MCP_BASE_PATH. |
TIME_OUT_SECONDS |
600 |
Generation timeout, in seconds. Aligned with the backend's read timeout. |
File access & confinement
By default, the file tools (video_to_music input, play_audio, and any
output_directory) are confined to SONILO_MCP_BASE_PATH. Paths that
resolve outside it (after symlink resolution) are rejected. This limits the
blast radius if a client is tricked into reading or exfiltrating arbitrary
files. To opt out — e.g. to read a video from elsewhere on disk — set
SONILO_MCP_ALLOW_ANY_PATH=true.
Tools
| Tool | Description | Cost |
|---|---|---|
text_to_music(prompt, duration, output_format?, variants_num?, output_directory?) |
Generate music from a text prompt. variants_num (1–10, default 1) generates that many distinct variants in one request, each with its own title; N > 1 forces the backend's async mode automatically and is never covered by the free trial. |
✅ |
video_to_music(video_path? | video_url?, prompt?, preserve_speech?, variants_num?, output_directory?) |
Generate music matched to a video. Max duration 360s (6 min); subject to the account's upload-size cap (typically 300 MB). preserve_speech (default false) keeps the source speech in the result — you also get a vocals speech stem and a ready-to-use mux; when set, the call runs the backend's async generation mode internally instead of streaming. variants_num (1–10, default 1) works the same as text_to_music's. |
✅ |
text_to_sfx(prompt, duration, audio_format?, output_directory?) |
Generate a sound effect from text. Duration 1–180s; formats wav/mp3/aac/flac (default aac). | ✅ |
video_to_sfx(video_path? | video_url?, prompt?, segments?, audio_format?, output_directory?) |
Generate SFX for a video; saves the generated SFX audio. Max video duration 180s (3 min). | ✅ |
video_to_video_music(video_path? | video_url?, prompt?, preserve_speech?, variants_num?, output_directory?) |
Generate music for a video and return a new .mp4 with it muxed in. Max video duration 360s (6 min). variants_num (1–10, default 1) saves one .mp4 per variant. |
✅ |
video_to_video_sfx(video_path? | video_url?, prompt?, segments?, output_directory?) |
Generate SFX for a video and return a new .mp4 with them muxed in. Max video duration 180s (3 min). |
✅ |
video_to_sound(video_path? | video_url?, music_prompt?, sfx_prompt?, segments?, preserve_speech?, ducking?, variants_num?, output_directory?) |
Generate music and SFX for a video in one call and save the mixed audio track. One charge instead of two, with the two layers balanced by the backend. ducking (default true) dips the music under the source speech. Max video duration 180s (3 min). variants_num (1–10, default 1) saves one mixed track per variant. |
✅ |
video_to_video_sound(video_path? | video_url?, music_prompt?, sfx_prompt?, segments?, preserve_speech?, ducking?, variants_num?, output_directory?) |
Same as video_to_sound, but returns a new .mp4 with the mixed soundtrack muxed in. Max video duration 180s (3 min). variants_num (1–10, default 1) saves one .mp4 per variant. |
✅ |
dubbing(video_path? | video_url?, languages?, output_directory?) |
Dub a video into other languages and save one dubbed .mp4 per language. languages is a list of codes (en, zh_cn, ja, ko, pt, es, de, fr, it, ru); omit it for the default ["zh_cn", "es", "fr"]. video_url must be https — the dubbing pipeline fetches the source itself. Max video duration 180s (3 min). Billed per language, with no free trial. |
✅ |
audio_ducking(voice_path? | voice_url?, music_path? | music_url?, output_directory?) |
Duck a music bed under a voice track. The voice input may be a video — the ducked mix is muxed back into a new .mp4. Each input max 360s (6 min); subject to the account's upload-size cap. |
✅ |
get_sfx_task(task_id, output_directory?) |
Check an SFX, audio-ducking, video-to-video, video-to-sound, dubbing, or async video-to-music task and download its result — recovery for timed-out text_to_sfx, video_to_sfx, audio_ducking, video_to_video_music, video_to_video_sfx, video_to_sound, video_to_video_sound, dubbing, and video_to_music(preserve_speech=true) calls. |
❌ |
get_account_services() |
List available services, limits, and the free-trial allowance left per service. | ❌ |
get_usage(days=30) |
Show usage summary + per-day breakdown. | ❌ |
play_audio(input_file_path) |
Play a local audio file. | ❌ |
Tools marked ✅ make API calls that incur charges on your Sonilo account.
Free trial: self-serve accounts start with a few free runs per service — no card required.
get_account_services()reports what is left astrial[service] = {granted, used, remaining}; check it before calling a ✅ tool so you can warn the user instead of failing on them. When a service'sremaininghits0, calls to it fail withtrial_exhausteduntil a payment method is added. Dubbing has no free runs and bills from the first call. Accounts without a free-trial allowance simply have notrialkey.
Optional: if
ffprobe(part of FFmpeg) is installed,video_to_musicchecks a video's duration locally and rejects anything over 360s before uploading.video_to_sfxperforms the same local check with its 180s cap.audio_duckingdoes the same for both of its inputs against its 360s cap. Without it, the same limits are still enforced by the backend.
Sound effects and ducking run as tasks
The music tools stream their result and finish in one call — with one exception: video_to_music(preserve_speech=true) submits a task and polls it internally instead, because keeping the source speech is only available in the backend's async mode. You still get the saved file paths back from a single call; you just don't see the task, and — like the SFX tools — it uses the same get_sfx_task recovery path if the call times out. The SFX tools submit a task, then poll it until it completes — text_to_sfx and video_to_sfx do this for you and return the saved file paths, so you normally never see the task. audio_ducking uses the same submit-then-poll flow and the same get_sfx_task recovery path.
If a call times out, the generation keeps running (and is already charged). The error message carries the task id, and get_sfx_task("<id>") retrieves the result once it's ready. The task id is also printed to stderr the moment a task is submitted, so it survives even a cancelled call. get_sfx_task is safe to call repeatedly: if the file is already on disk it reports that instead of downloading a second copy.
Output Format
Music is saved as .m4a (AAC in MP4 container). File names use the title returned by the backend (slugified), or a sonilo-<timestamp>.m4a fallback. When multiple parallel streams are returned, a -<index> suffix is appended.
video_to_music(preserve_speech=true) saves up to three kinds of file, each labeled in the returned text: the generated music audio (same naming as above, based on prompt or falling back to music-<first 8 chars of the task id>), the preserved speech stem as <base>-vocals.<ext>, and the mux — speech and music already mixed together, the ready-to-use combined result — as <base>-mux.<ext> (or <base>-mux-<index>.<ext> for multiple streams). The speech stem and mux extensions come from the backend's reported content_type (typically .m4a).
Sound effects are saved in the requested audio_format — wav, mp3, flac, or aac (the default, written as .m4a); video_to_sfx saves audio only, not the source video.
Combined sound (video_to_sound / video_to_video_sound) is saved as a single file: a .wav for video_to_sound, a .mp4 for video_to_video_sound. The name comes from music_prompt, falling back to sfx_prompt and then to sound-<first 8 chars of the task id> / v2v-sound-<first 8 chars of the task id>. The separate music and SFX stems stay on the backend — only the mixed result is downloaded.
Multiple variants (variants_num > 1 on text_to_music, video_to_music, video_to_video_music, video_to_sound, video_to_video_sound) save one file per variant instead of one, suffixed -0, -1, … in request order (e.g. <base>-0.m4a, <base>-1.m4a). For the music tools, each variant is its own creative direction and the returned text names it when the backend provides a title. Cost scales linearly with variants_num, and any value above 1 is billed in full even on a free-trial account. A timed-out multi-variant call recovers all of its variants through get_sfx_task, the same as a single-variant one.
Dubbing (dubbing) is saved as one .mp4 per requested language, named dubbing-<first 8 chars of the task id>.<language>.mp4 — there is no prompt to name the files after. Asking for four languages writes four files and costs four times as much as one. Dubbing polls for at least two hours regardless of TIME_OUT_SECONDS, matching the backend's own ceiling for the job.
File names come from the prompt (slugified, truncated to 80 characters). When there is no prompt to name a file after — video_to_sfx without one, or an SFX/ducking file recovered via get_sfx_task — the name is sfx-<first 8 chars of the task id> instead. A video_to_music(preserve_speech=true) task recovered via get_sfx_task is named music-<first 8 chars of the task id> instead (get_sfx_task detects the music envelope shape and saves audio/vocals/mux the same way as a direct video_to_music call). Existing files are never overwritten: a -1, -2, … suffix is added instead.
Ducking results are saved as a single file: a .wav, or a .mp4 when the voice input was a video (the ducked mix is muxed back into it). The file name is the voice input's name plus -ducked (e.g. interview.mp4 → interview-ducked.mp4), falling back to ducked-<first 8 chars of the task id> when there is no usable name. A ducking result recovered via get_sfx_task is named sfx-<first 8 chars of the task id> instead, since that tool has no voice file name to work from.
Common Errors
| Message | What to do |
|---|---|
Invalid SONILO_API_KEY |
Verify the key at https://platform.sonilo.com/dashboard/api-keys. |
Insufficient minutes / Credit limit exceeded |
Top up at https://platform.sonilo.com/dashboard/billing. |
You've used your N free trial calls for <service> |
That service's free trial is spent. Add a payment method at https://platform.sonilo.com/dashboard/billing — retrying can't help. get_account_services shows what is left on the other services. |
Rate limit exceeded |
Check get_account_services for your rpm/concurrency limits. |
Generation timed out (music, text_to_music/video_to_music without preserve_speech) |
Raise TIME_OUT_SECONDS. Check get_usage to confirm whether the backend completed and charged. |
Timed out … waiting for task <id> (SFX, ducking, or video_to_music(preserve_speech=true)) |
The generation is still running. Call get_sfx_task("<id>") to retrieve the result — nothing is lost, including for a timed-out preserve_speech music task (get_sfx_task recognizes its result shape and saves audio/vocals/mux the same way video_to_music itself would). |
Task not found |
The task id doesn't exist, or belongs to a purely streaming generation (text_to_music, or video_to_music without preserve_speech), which isn't pollable. Check the id. |
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file sonilo_mcp-0.11.0.tar.gz.
File metadata
- Download URL: sonilo_mcp-0.11.0.tar.gz
- Upload date:
- Size: 83.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
92dc52ee77e1072b054d9f5de773566821240f67fb34495ba761da972969c1a9
|
|
| MD5 |
518dd1ea3157dd6920fdd114a2dd3510
|
|
| BLAKE2b-256 |
7bb3716a9e9334d1fb16dfca851f60be929e27a9c8d2374afbcb0f245a3a5b4e
|
Provenance
The following attestation bundles were made for sonilo_mcp-0.11.0.tar.gz:
Publisher:
release.yml on sonilo-ai/sonilo-mcp
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
sonilo_mcp-0.11.0.tar.gz -
Subject digest:
92dc52ee77e1072b054d9f5de773566821240f67fb34495ba761da972969c1a9 - Sigstore transparency entry: 2305970758
- Sigstore integration time:
-
Permalink:
sonilo-ai/sonilo-mcp@df61ef2a208473fc6e06285a82e4feba9b6c5b78 -
Branch / Tag:
refs/tags/v0.11.0 - Owner: https://github.com/sonilo-ai
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@df61ef2a208473fc6e06285a82e4feba9b6c5b78 -
Trigger Event:
push
-
Statement type:
File details
Details for the file sonilo_mcp-0.11.0-py3-none-any.whl.
File metadata
- Download URL: sonilo_mcp-0.11.0-py3-none-any.whl
- Upload date:
- Size: 44.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2352803a10aae1d612b5713566b1e7de59dfd26a22a90fd9f8cf2426ecc91667
|
|
| MD5 |
cc5ff11b751655a5c127b216725f81f1
|
|
| BLAKE2b-256 |
5d2acb9c4107ed64f1a7f6d9dce2c2c3c79ed37b9f9d14d41b1e89fd545caa25
|
Provenance
The following attestation bundles were made for sonilo_mcp-0.11.0-py3-none-any.whl:
Publisher:
release.yml on sonilo-ai/sonilo-mcp
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
sonilo_mcp-0.11.0-py3-none-any.whl -
Subject digest:
2352803a10aae1d612b5713566b1e7de59dfd26a22a90fd9f8cf2426ecc91667 - Sigstore transparency entry: 2305970822
- Sigstore integration time:
-
Permalink:
sonilo-ai/sonilo-mcp@df61ef2a208473fc6e06285a82e4feba9b6c5b78 -
Branch / Tag:
refs/tags/v0.11.0 - Owner: https://github.com/sonilo-ai
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@df61ef2a208473fc6e06285a82e4feba9b6c5b78 -
Trigger Event:
push
-
Statement type: