Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

hivemind-audio-binary-protocol

Binary audio plugin for hivemind-core.

The plugin adds server-side WakeWord detection, VAD, STT, and TTS to a hivemind-core hub. Lightweight satellites, such as hivemind-mic-satellite, stream raw audio to the hub and receive transcriptions or synthesized speech. The satellites do not run those models locally.

Where it fits

hivemind-core
  └── hivemind-plugin-manager  (BinaryDataHandlerFactory loads plugins by entry-point)
        └── hivemind-audio-binary-protocol  ← this repo
              ├── ovos-simple-listener  (WakeWord + VAD + STT pipeline)
              └── OVOSTTSFactory / OVOSSTTFactory / OVOSVADFactory / OVOSWakeWordFactory

The plugin registers under the hivemind.binary.protocol entry-point group as hivemind-audio-binary-protocol-plugin.

Install

pip install hivemind-audio-binary-protocol

You also need OVOS STT, TTS, VAD, and WakeWord plugins. Install them as you would in a standard OVOS setup:

pip install ovos-stt-plugin-server ovos-tts-plugin-piper ovos-vad-plugin-silero \
            ovos-ww-plugin-precise-lite

Quickstart

Add the binary_protocol block to ~/.config/hivemind-core/server.json:

{
  "binary_protocol": {
    "module": "hivemind-audio-binary-protocol-plugin",
    "hivemind-audio-binary-protocol-plugin": {
      "stt": {
        "module": "ovos-stt-plugin-server",
        "ovos-stt-plugin-server": {"url": "https://stt.openvoiceos.org"}
      },
      "tts": {
        "module": "ovos-tts-plugin-piper",
        "ovos-tts-plugin-piper": {"voice": "en_US-lessac-medium"}
      },
      "vad": {
        "module": "ovos-vad-plugin-silero"
      },
      "wake_word": "hey_mycroft",
      "hotwords": {
        "hey_mycroft": {
          "module": "ovos-ww-plugin-precise-lite",
          "model": "https://github.com/OpenVoiceOS/precise-lite-models/raw/master/wakewords/en/hey_mycroft.tflite"
        }
      }
    }
  }
}

Then start hivemind-core with the listen subcommand:

hivemind-core listen

Audio streaming modes

This plugin handles three binary audio flows:

Mode Client sends Hub returns Use case
Microphone stream Raw PCM audio chunks Bus messages (wakeword/utterance events) Mic satellite. The hub runs the full pipeline.
STT transcription Raw PCM audio recognizer_loop:transcribe.response Client wants a transcription without triggering skills.
STT handle Raw PCM audio Triggers recognizer_loop:utterance on the bus Client wants the hub to handle the utterance.

The bus triggers TTS (speak:synth or speak:b64_audio) and returns binary WAV audio or a Base64-encoded string to the client.

Configuration reference

The plugin's config block mirrors the OVOS plugin config convention. Each sub-plugin (stt, tts, vad) takes its standard OVOS config:

Key Description
stt STT plugin config. module selects the OVOS STT plugin.
tts TTS plugin config. module selects the OVOS TTS plugin.
vad VAD plugin config. module selects the OVOS VAD plugin.
wake_word WakeWord name (key into hotwords).
hotwords Dict of wakeword configurations, keyed by wakeword name.
utterance_transformers List of OVOS utterance transformer plugin names.
dialog_transformers List of OVOS dialog transformer plugin names.
metadata_transformers List of OVOS metadata transformer plugin names.
audio_transformers List of OVOS audio transformer plugin names applied to raw audio before STT.
tts_transformers List of OVOS tts transformer plugin names applied to synthesized audio after TTS.

If the config block is omitted, the plugin falls back to reading mycroft.conf (the standard OVOS configuration file) to select plugins.

Access control

This plugin respects hivemind-core's per-client allowed_types whitelist. Clients must have the correct access to send binary audio or receive TTS output.

Related projects

License

Apache License 2.0. See LICENSE.

Docs

Metadata

Release files for hivemind-audio-binary-protocol 2.2.2a1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for hivemind-audio-binary-protocol 2.2.2a1
File Size Uploaded
hivemind_audio_binary_protocol-2.2.2a1.tar.gz 20.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for hivemind-audio-binary-protocol 2.2.2a1
File Interpreter ABI Platform
hivemind_audio_binary_protocol-2.2.2a1-py3-none-any.whl Python 3 none any Details

Total release size: 37.1 kB

Release files / hivemind_audio_binary_protocol-2.2.2a1.tar.gz

Download URL hivemind_audio_binary_protocol-2.2.2a1.tar.gz
Size 20.4 kB
Tags Source
SHA-256 checksum
How to use checksums
1700a0334fc0b52b9d2d9808e9861e64e00827a94e054a62cf566267ca84cfe0
BLAKE2b-256 checksum
How to use checksums
d5139b1f18f4470d1542895abf1cd4360ccf8eb0a469f7b4386d5f979c7b3652
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / hivemind_audio_binary_protocol-2.2.2a1-py3-none-any.whl

Download URL hivemind_audio_binary_protocol-2.2.2a1-py3-none-any.whl
Size 16.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
86f2d0c094f9291a5606e71fb8058316272a35c5096369c9ed3e74eab02c6804
BLAKE2b-256 checksum
How to use checksums
c3b9f0e61e80bb0c39f033f8f8ecfd2f0ec3329c409de1e947f2fdc717ea16e0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page