Skip to main content

Moonshine Voice Python Package

A fast, accurate, on-device AI library for building interactive voice applications. Join our Discord to get help and support.

Installation

pip install moonshine-voice

Quick Start

# Listens to the microphone, logging to the console when there are 
# speech updates.
moonshine-voice mic

Installing the package adds a moonshine-voice command (with a shorter moonshine alias) that groups the built-in tools as subcommands: mic, transcribe, tts, intent, download, and g2p. Run moonshine-voice --help, or moonshine-voice <command> --help for a specific tool. Each subcommand is equivalent to python -m moonshine_voice.<module>, so either invocation style works.

Example

"""Transcribes live audio from the default microphone"""
import time
from moonshine_voice import (
    MicTranscriber,
    TranscriptEventListener,
    get_model_for_language,
)

# This will download the model files and cache them.
model_path, model_arch = get_model_for_language("en")

# MicTranscriber handles connecting to the microphone, capturing
# the audio data, detecting voice activity, breaking the speech
# up into segments, transcribing the speech, and sending events
# as the results are updated over time.
mic_transcriber = MicTranscriber(
    model_path=model_path, model_arch=model_arch)

# We use an event-driven interface to respond in real time
# as speech is detected.
class TestListener(TranscriptEventListener):
    def on_line_started(self, event):
        print(f"Line started: {event.line.text}")

    def on_line_text_changed(self, event):
        print(f"Line text changed: {event.line.text}")

    def on_line_completed(self, event):
        print(f"Line completed: {event.line.text}")

listener = TestListener()
mic_transcriber.add_listener(listener)
mic_transcriber.start()
print("Listening to the microphone, press Ctrl+C to stop...")

while True:
    time.sleep(0.1)

Other Sources

If you have a different source you're capturing audio from you can supply it directly to a transcriber.

"""Transcribes live audio from an arbitrary audio source."""
from moonshine_voice import (
    Transcriber,
    TranscriptEventListener,
    get_model_for_language,
    load_wav_file,
    get_assets_path,
)
import os
from typing import Iterator, Tuple


def audio_chunk_generator(
    wav_file_path: str, chunk_duration: float = 0.1
) -> Iterator[Tuple[list, int]]:
    """
    Example function that loads a WAV file and yields audio chunks.

    This demonstrates how you can integrate your own proprietary
    audio data capture sources. Replace this function with your own
    implementation that yields (audio_chunk, sample_rate) tuples.

    Args:
        wav_file_path: Path to the WAV file to load
        chunk_duration: Duration of each chunk in seconds

    Yields:
        Tuple of (audio_chunk, sample_rate) where:
        - audio_chunk: List of float audio samples
        - sample_rate: Sample rate in Hz
    """
    audio_data, sample_rate = load_wav_file(wav_file_path)
    chunk_size = int(chunk_duration * sample_rate)

    for i in range(0, len(audio_data), chunk_size):
        chunk = audio_data[i: i + chunk_size]
        yield (chunk, sample_rate)


model_path, model_arch = get_model_for_language("en")

transcriber = Transcriber(
    model_path=model_path, model_arch=model_arch)

stream = transcriber.create_stream(update_interval=0.5)
stream.start()


class TestListener(TranscriptEventListener):
    def on_line_started(self, event):
        print(f"{event.line.start_time:.2f}s: Line started: {event.line.text}")

    def on_line_text_changed(self, event):
        print(
            f"{event.line.start_time:.2f}s: Line text changed: {event.line.text}")

    def on_line_completed(self, event):
        print(f"{event.line.start_time:.2f}s: Line completed: {event.line.text}")


listener = TestListener()
stream.add_listener(listener)

# Feed audio chunks from the generator into the stream.
wav_file_path = os.path.join(get_assets_path(), "two_cities.wav")
for chunk, sample_rate in audio_chunk_generator(wav_file_path):
    stream.add_audio(chunk, sample_rate)

stream.stop()
stream.close()

Voice Commands

We also provide voice command recognition using the IntentRecognizer module. It captures transcribed audio from a MicTranscriber and invokes callback functions that match your programmed intents. Since it relies on an embedding model, you can use a helper function to get started:

from moonshine_voice import (
    MicTranscriber,
    IntentRecognizer,
    ModelArch,
    EmbeddingModelArch,
    get_embedding_model,
    get_model_for_language
)

# Download and load the embedding model for intent recognition
embedding_model_path, embedding_model_arch = get_embedding_model()

Next, create a recognizer and register your intent callbacks:

intent_recognizer = IntentRecognizer(
    model_path=embedding_model_path,
    model_arch=embedding_model_arch
)

def on_lights_on(trigger: str, utterance: str, similarity: float):
    """Handler for turning lights on."""
    print(f"\n💡 LIGHTS ON! (matched '{trigger}' with {similarity:.0%} confidence)")

def on_lights_off(trigger: str, utterance: str, similarity: float):
    """Handler for turning lights off."""
    print(f"\n🌑 LIGHTS OFF! (matched '{trigger}' with {similarity:.0%} confidence)")

intent_recognizer.register_intent("turn on the lights", on_lights_on)
intent_recognizer.register_intent("turn off the lights", on_lights_off)

Finally, create a MicTranscriber, connect it to your IntentRecognizer, and start the audio stream:

# Get the transcription model and initialize a MicTranscriber
model_path, model_arch = get_model_for_language("en")
mic_transcriber = MicTranscriber(model_path=model_path, model_arch=model_arch)

# The intent recognizer will process completed transcript lines and invoke trigger handlers
mic_transcriber.add_listener(intent_recognizer)

mic_transcriber.start()
try:
    while True:
        time.sleep(0.1)
except KeyboardInterrupt:
    print("\n\nStopping...", file=sys.stderr)
finally:
    intent_recognizer.close()
    mic_transcriber.stop()
    mic_transcriber.close()

Multiple Languages

The framework currently supports English, Spanish, Mandarin, Japanese, Korean, Vietnamese, Arabic, and Ukrainian. We are working on wider language support, and you can see which are supported in your version by calling supported_languages(). To use a language, request it using get_model_for_language() passing in the two-letter language code. For example get_model_for_language("es") will download the Spanish models and pass the information you need to create Transcriber objects using them.

Documentation

For more information, see the main Moonshine Voice documentation.

License

The code and English-language models are released under the MIT License - see the main project repository for details. The models used for other languages are released under the Moonshine Community License.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

moonshine_voice-0.0.73-py3-none-win_amd64.whl (50.5 MB view details)

Uploaded Python 3Windows x86-64

moonshine_voice-0.0.73-py3-none-manylinux_2_34_x86_64.whl (56.3 MB view details)

Uploaded Python 3manylinux: glibc 2.34+ x86-64

moonshine_voice-0.0.73-py3-none-manylinux_2_34_aarch64.whl (55.1 MB view details)

Uploaded Python 3manylinux: glibc 2.34+ ARM64

moonshine_voice-0.0.73-py3-none-manylinux_2_31_aarch64.manylinux_2_39_aarch64.whl (55.0 MB view details)

Uploaded Python 3manylinux: glibc 2.31+ ARM64manylinux: glibc 2.39+ ARM64

moonshine_voice-0.0.73-py3-none-macosx_15_0_arm64.whl (54.4 MB view details)

Uploaded Python 3macOS 15.0+ ARM64

File details

Details for the file moonshine_voice-0.0.73-py3-none-win_amd64.whl.

File metadata

File hashes

Hashes for moonshine_voice-0.0.73-py3-none-win_amd64.whl
Algorithm Hash digest
SHA256 fd9895bec46f929ac3747fb83bede3826f7a32c3d0c9559b4b2840fa5e6d2bcd
MD5 a975420cfa1f70cb6dbc5e16c9d3c5e9
BLAKE2b-256 421451f5813cd4cb1a172885e5b6e7f8a50ca2462cb418b67773d8bf8ab1dcf5

See more details on using hashes here.

File details

Details for the file moonshine_voice-0.0.73-py3-none-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for moonshine_voice-0.0.73-py3-none-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 6461c8097ed78d054c46b4c821143008a07b5fa1f4169674493ccfa06c1407db
MD5 f68201988a96ba26b57b4a21c626bb28
BLAKE2b-256 b67a245402af9a56879cf24f36ba86087240ef8d5f807064fdd510c58de88947

See more details on using hashes here.

File details

Details for the file moonshine_voice-0.0.73-py3-none-manylinux_2_34_aarch64.whl.

File metadata

File hashes

Hashes for moonshine_voice-0.0.73-py3-none-manylinux_2_34_aarch64.whl
Algorithm Hash digest
SHA256 c56de7614b18bbc65d8f4f2cb7432d8a258311aef38a49e9e16c8d06df8085f0
MD5 6785f1eb2c6cfe7f4e255ddc0047171c
BLAKE2b-256 e11af32fe9d27851a6c88bdea75891c1d5f0bac11ee5ab72517b9c42657c0262

See more details on using hashes here.

File details

Details for the file moonshine_voice-0.0.73-py3-none-manylinux_2_31_aarch64.manylinux_2_39_aarch64.whl.

File metadata

File hashes

Hashes for moonshine_voice-0.0.73-py3-none-manylinux_2_31_aarch64.manylinux_2_39_aarch64.whl
Algorithm Hash digest
SHA256 d01e1343bfeb436a59f7d8562566d4b2f2a526297328d776a13d1544887aec61
MD5 62f362940f84348305deb5f0dfd078e5
BLAKE2b-256 095339440fef122417b48dd94bec38749970d60cfd57b411dc810f3c8d27f8c7

See more details on using hashes here.

File details

Details for the file moonshine_voice-0.0.73-py3-none-macosx_15_0_arm64.whl.

File metadata

File hashes

Hashes for moonshine_voice-0.0.73-py3-none-macosx_15_0_arm64.whl
Algorithm Hash digest
SHA256 fb4c81f2fa96674d336530bc1cf2f86b6a4fc8b3c02e138a7a5c768c283f8bc7
MD5 01ac10427df2bf0c81a54b7164accd46
BLAKE2b-256 cca3af1e1f0aaef7175727f9fedfec6863f551c0e00159a796b11a6579c6611d

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page