Skip to main content

LLM Voice

Library to reduce latency in voice generations from LLM chat completion streams.

This lets you generate voices using completely local LLM models, such as Ollama and TTS clients, such as Apple Say and Google Text-to-Speech with the same speed as privately created assistants such as OpenAI.

Installation and Setup

  1. Install the package from PyPI with:

    pip install llm-voice
    
  2. Copy the .env.example file to .env and fill in your OpenAI API key if you want to use OpenAI along with the model name for the Ollama/OpenAI model you want to use.

  3. Take a look at one of the examples to start generating voice responses in realtime.

Example Usage

The example below can be found in the examples directory.

# Setup output device, TTS client and LLM client.
devices: list[AudioDevice] = AudioDevices.get_list_of_devices(
   device_type=AudioDeviceType.OUTPUT
)

# Pick the first output device (usually computer builtin speakers if nothing else if connected).
output_device: AudioDevice = devices[0]

# Change to another TTS client depending on your needs and desires.
tts_client: TextToSpeechClient = OpenAITextToSpeechClient()

# Change to another LLM client depending on your needs and desires.
llm_client: LLMClient = OllamaClient(
   model_name=MODEL_NAME,
)

# Define messages to send to the LLM.
messages: list[ChatMessage] = [
   ChatMessage(
      role=MessageRole.SYSTEM,
      content="You are a helpful assistant named Alfred.",
   ),
   ChatMessage(role=MessageRole.USER, content="Hey there what is your name?"),
]

# Use the LLM to generate a response.
chat_stream: Iterator[str] = llm_client.generate_chat_completion_stream(
   messages=messages,
)

# Create the voice responder and speak the response.
voice_responder_fast = VoiceResponderFast(
   text_to_speech_client=tts_client,
   output_device=output_device,
)

# Will speak each sentence back to back as it is available.
voice_responder_fast.respond(chat_stream)

Install From Source

pip install poetry
poetry install

License

This project is licensed under the MIT License. See the LICENSE file for details.

Metadata

Release files for llm-voice 0.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llm-voice 0.0.1
File Size Uploaded
llm_voice-0.0.1.tar.gz 12.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llm-voice 0.0.1
File Interpreter ABI Platform
llm_voice-0.0.1-py3-none-any.whl Python 3 none any Details

Total release size: 33.4 kB

Release files / llm_voice-0.0.1.tar.gz

Download URL llm_voice-0.0.1.tar.gz
Size 12.7 kB
Tags Source
SHA-256 checksum
How to use checksums
7d110c0f509c5c5362797aa829432d86cd51066c2ec77acbafce42f9d12ce1e1
BLAKE2b-256 checksum
How to use checksums
e44258e7210cae570c0aad7576f56c0bb12d74859fde485108bc930d88b93a77
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.8.2 CPython/3.12.2 Darwin/23.5.0

Release files / llm_voice-0.0.1-py3-none-any.whl

Download URL llm_voice-0.0.1-py3-none-any.whl
Size 20.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4748ffdd1c84bb91d3b8ce8b60100fe6f83f4a1a6bb17bf568fb51f9a76baf21
BLAKE2b-256 checksum
How to use checksums
d98463d32108f85066fd807b67dd7bda8053b80e5e035c1c5b2cfd8cdc0c9476
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.8.2 CPython/3.12.2 Darwin/23.5.0

Release history Release notifications | RSS feed

This release

0.0.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page