Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

ovos-gguf-plugin

A unified GGUF wrapper for OpenVoiceOS. It covers chat, summarization, dialog rewriting, translation, language detection, and text embeddings, all backed by quantized GGUF models through llama-cpp-python.

Install

pip install ovos-gguf-plugin

For GPU inference, rebuild llama-cpp-python with CUDA support first:

CMAKE_ARGS="-DGGML_CUDA=on" FORCE_CMAKE=1 pip install llama-cpp-python --force-reinstall --no-cache-dir

Plugin entry points

Entry-point group Plugin name Class Role
opm.agents.chat ovos-chat-gguf-plugin GGUFChatEngine conversational chat / question answering
opm.agents.summarizer ovos-summarizer-gguf-plugin GGUFSummarizer text summarization
opm.transformer.dialog ovos-dialog-transformer-gguf-plugin GGUFDialogTransformer dialog rewriting
opm.lang.translate ovos-translate-gguf-plugin GGUFTextTranslator machine translation
opm.lang.detect ovos-lang-detect-gguf-plugin GGUFTextLangDetector language detection
opm.embeddings.text ovos-gguf-embeddings-plugin GGUFEmbeddings text embeddings

Quickstart

Chat

from ovos_gguf_plugin.chat import GGUFChatEngine
from ovos_plugin_manager.templates.agents import AgentMessage, MessageRole

engine = GGUFChatEngine({
    "model": "afrideva/Smol-Llama-101M-Chat-v1-GGUF",
    "remote_filename": "*q2_k.gguf",
    "max_tokens": 128,
})
msgs = [AgentMessage(role=MessageRole.USER, content="Tell me a joke.")]
# stream sentence-by-sentence (suitable for TTS)
for sentence in engine.stream_sentences(msgs):
    print(sentence)
# or get the full response at once
reply = engine.continue_chat(msgs)
print(reply.content)

Summarizer

from ovos_gguf_plugin.summarizer import GGUFSummarizer

s = GGUFSummarizer({
    "model": "Qwen/Qwen2-0.5B-Instruct-GGUF",
    "remote_filename": "*q8_0.gguf",
})
print(s.summarize("Long document text goes here ... " * 20))

Dialog transformer

from ovos_gguf_plugin.dialog_transformers import GGUFDialogTransformer

dt = GGUFDialogTransformer({
    "model": "Qwen/Qwen2-0.5B-Instruct-GGUF",
    "remote_filename": "*q8_0.gguf",
})
print(dt.transform("gonna grab some food real quick"))

Translation

from ovos_gguf_plugin.translate import GGUFTextTranslator

tx = GGUFTextTranslator({
    "model": "TheBloke/TowerInstruct-7B-v0.1-GGUF",
    "remote_filename": "*Q4_K_M.gguf",
})
print(tx.translate("the easiest way to contribute is to help with translations",
                   target="es-es"))

Language detection

from ovos_gguf_plugin.translate import GGUFTextLangDetector

dt = GGUFTextLangDetector({
    "model": "Qwen/Qwen2-0.5B-Instruct-GGUF",
    "remote_filename": "*q8_0.gguf",
})
print(dt.detect("you can help without any programming knowledge"))  # → en

Text embeddings

from ovos_gguf_plugin.embeddings import GGUFEmbeddings

emb = GGUFEmbeddings({"model": "all-MiniLM-L6-v2"})
vector = emb.get_embeddings("hello world")
print(len(vector), "dims")

model accepts a friendly name from GGUFEmbeddings.DEFAULT_MODELS (e.g. labse, all-MiniLM-L6-v2, nomic-embed-text-v1.5, bge-large-en-v1.5), a bare Hugging Face repo id (with remote_filename), or a local .gguf path. Default is labse.

As an OVOS text-embeddings plugin it is selected by name (ovos-gguf-embeddings-plugin), so it is a drop-in for anything that previously used the standalone embeddings plugin.

Configuration

All wrappers share the same config keys:

Key Default Description
model required Local .gguf path, HuggingFace repo id, or friendly name (embeddings)
remote_filename *Q4_K_M.gguf Glob for selecting the file from a HF repo
n_gpu_layers 0 GPU layers to offload (-1 = all)
chat_format None llama.cpp chat format (auto-detected for most models)
verbose True llama.cpp verbosity
max_tokens 512 Maximum tokens to generate
system_prompt locale default Override the system prompt

See docs/configuration.md for the full reference, including per-wrapper options and GPU build instructions.

Localized prompts

System prompts and templates ship as .prompt resource files under ovos_gguf_plugin/locale/<lang>/. They load through OpenVoiceOS/ovos-spec-tools (OVOS-INTENT-2 §4.4). To add a language, drop translated .prompt files under a new locale/<lang>/ folder. English (en-us) ships by default and acts as the fallback. A system_prompt in config overrides the locale file.

See docs/localization.md for the full guide.

OVOS Persona Framework

{
  "name": "MyAssistant",
  "solvers": ["ovos-solver-gguf-plugin"],
  "ovos-solver-gguf-plugin": {
    "model": "TheBloke/notus-7B-v1-GGUF",
    "remote_filename": "*Q4_K_M.gguf",
    "persona": "You are a helpful assistant.",
    "verbose": false
  }
}
ovos-persona-server --persona my_persona.json

Documentation

Examples

Runnable scripts under examples/:

Testing

pip install "ovos-gguf-plugin[test]"
python -m pytest test/ -v

The test suite contains:

  • test/test_embeddings.py: hermetic unit tests (mocked llama.cpp, no downloads)
  • test/test_prompts.py: hermetic unit tests for localized prompt loading
  • test/test_e2e.py: real-model end-to-end tests (downloads tiny GGUFs once, about 70 MB total):
    • chat: afrideva/Smol-Llama-101M-Chat-v1-GGUF q2_k (~45 MB)
    • embeddings: leliuga/all-MiniLM-L6-v2-GGUF Q4_K_M (~23 MB)

Credits

Originally developed by TigreGótico for OpenVoiceOS, sponsored by VisioLab. Modernized under the NGI0 Commons Fund / NLnet.

VisioLab

This work was sponsored by VisioLab, part of Royal Dutch Visio. Royal Dutch Visio is a Dutch test, education, and research center for assistive technology for blind and visually impaired people and professionals. It explores technology such as voice, VR, and AI, and shares the resulting knowledge and expertise with everyone.

NGI0 Commons Fund

This project was funded through the NGI0 Commons Fund, a fund established by NLnet with financial support from the European Commission's Next Generation Internet programme, under the aegis of DG Communications Networks, Content and Technology under grant agreement No 101135429.

Release files for ovos-gguf-plugin 1.2.4a2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ovos-gguf-plugin 1.2.4a2
File Size Uploaded
ovos_gguf_plugin-1.2.4a2.tar.gz 25.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ovos-gguf-plugin 1.2.4a2
File Interpreter ABI Platform
ovos_gguf_plugin-1.2.4a2-py3-none-any.whl Python 3 none any Details

Total release size: 45.6 kB

Release files / ovos_gguf_plugin-1.2.4a2.tar.gz

Download URL ovos_gguf_plugin-1.2.4a2.tar.gz
Size 25.6 kB
Tags Source
SHA-256 checksum
How to use checksums
f49e5c5704b9609b6f2c5c2bc88b7050ddf9455eca90bc894a550dc521b8ff03
BLAKE2b-256 checksum
How to use checksums
b05b0fd615bafb7c867c956877aecb4335ae8c0f56f56fd9031966c5533052e9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / ovos_gguf_plugin-1.2.4a2-py3-none-any.whl

Download URL ovos_gguf_plugin-1.2.4a2-py3-none-any.whl
Size 20.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0bb95f7cb9612d042a3006a876cfcb5c03f70ea357e666d01d7aa48805caa207
BLAKE2b-256 checksum
How to use checksums
c3c910022acef99f3cb7a8ef9d4762b3ca5857b88d9fe2205e4110849f29cf85
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page