This release is a pre-release and may not be stable for production use.
ovos-voice-embeddings-plugin
An OVOS voice embeddings plugin. It turns a speech clip into a speaker
embedding with speakeronnx, so
the runtime needs onnxruntime and numpy only, no PyTorch. The default model
is wespeaker-resnet34, a 256-dimension L2-normalised embedding downloaded from
the OpenVoiceOS speaker-embeddings-onnx
collection on first use. Two clips of one speaker give vectors with a high
cosine similarity. Clips of different speakers give a low one.
The plugin implements the VoiceEmbedder template of
ovos-plugin-manager and
registers under the opm.embeddings.voice entry point as
ovos-voice-embeddings-plugin.
Install
pip install ovos-voice-embeddings-plugin
Usage
import numpy as np
import soundfile as sf
from ovos_voice_embeddings import SpeakerOnnxVoiceEmbedder
embedder = SpeakerOnnxVoiceEmbedder() # wespeaker-resnet34, 16 kHz input
alice_1, _ = sf.read("alice_1.wav", dtype="float32")
alice_2, _ = sf.read("alice_2.wav", dtype="float32")
bob, _ = sf.read("bob.wav", dtype="float32")
a1 = embedder.get_embeddings(alice_1)
a2 = embedder.get_embeddings(alice_2)
b = embedder.get_embeddings(bob)
print(a1.shape) # (256,)
print(float(np.dot(a1, a2))) # same speaker, about 0.8
print(float(np.dot(a1, b))) # different speakers, about 0.4
The embeddings are L2-normalised, so the dot product is the cosine similarity.
get_embeddings also accepts raw signed 16-bit PCM bytes, which is what the
OVOS listener hands to plugins.
Store and match voices
Pair the embedder with any OVOS embeddings database plugin, for example ovos-chromadb-embeddings-plugin:
from ovos_chromadb_embeddings import ChromaEmbeddingsDB
db = ChromaEmbeddingsDB({"path": "./voice_db"})
db.add_embeddings("alice", a1)
db.add_embeddings("bob", b)
print(db.query(a2, top_k=1)) # [("alice", distance)]
Configuration
| key | default | description |
|---|---|---|
model |
"wespeaker-resnet34" |
speakeronnx model alias or path to an ONNX file |
sample_rate |
16000 |
sample rate of the audio passed to get_embeddings |
Tests
pip install -e ".[test]"
pytest tests
The tests embed three real clips from LibriSpeech test-clean (CC BY 4.0) and assert the vector shape and that two clips of one speaker are closer than clips of two speakers.
License
Apache-2.0.
Metadata
Release files for ovos-voice-embeddings-plugin 0.0.1a1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ovos_voice_embeddings_plugin-0.0.1a1.tar.gz | 8.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ovos_voice_embeddings_plugin-0.0.1a1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 16.6 kB
Release files / ovos_voice_embeddings_plugin-0.0.1a1.tar.gz
| Download URL | ovos_voice_embeddings_plugin-0.0.1a1.tar.gz |
|---|---|
| Size | 8.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
cf92d080cc21f41d448a6b9618d0aaa3fdebb7e12bd6bd30cc74a69c0424629b
|
|
BLAKE2b-256 checksum How to use checksums |
2aa132eb912c934c20d756946828c5689181bb4d40da2505434e4f189bf5ef44
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / ovos_voice_embeddings_plugin-0.0.1a1-py3-none-any.whl
| Download URL | ovos_voice_embeddings_plugin-0.0.1a1-py3-none-any.whl |
|---|---|
| Size | 8.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
09c3d84aab89a8b97d1903001d65f7db0bea19a1a6cd569d2de89fa841e86f17
|
|
BLAKE2b-256 checksum How to use checksums |
0b38cefb182e307ae6916ebc1cfb6cdcd9186da966dfa67b97c16e2ba1179d3e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|