Skip to main content

Inworld AI integration for Vision Agents (TTS + Realtime WebRTC)

Project description

Inworld AI Plugin

Inworld AI integration for Vision Agents. Provides both text-to-speech and a WebRTC-based Realtime speech-to-speech conversational API.

Installation

uv add "vision-agents[inworld]"
# or directly
uv add vision-agents-plugins-inworld

Get your API key from the Inworld Portal and set INWORLD_API_KEY in your environment (or pass api_key= explicitly).

TTS

High-quality text-to-speech with streaming support. The plugin now defaults to Inworld's TTS-2 model (currently in research preview), which adds natural-language steering, 100+ languages (15 GA, 90+ experimental), and high-quality instant voice cloning over the previous inworld-tts-1.5-* generation.

from vision_agents.plugins import inworld

# Defaults to model_id="inworld-tts-2", voice_id="Sarah"
tts = inworld.TTS()

# Or specify explicitly
tts = inworld.TTS(
    api_key="your_inworld_api_key",
    voice_id="Ashley",
    model_id="inworld-tts-2",
    temperature=1.1,
)

TTS options

  • api_key: Inworld AI API key (default: reads from INWORLD_API_KEY)
  • voice_id: Voice to use (default: "Sarah"; "Dennis", "Ashley", "Olivia", "Clive" and custom/cloned voices also supported)
  • model_id: "inworld-tts-2" (default), "inworld-tts-1.5-max", "inworld-tts-1.5-mini". "inworld-tts-1" and "inworld-tts-1-max" are deprecated by Inworld — migrate to inworld-tts-2 or inworld-tts-1.5-*.
  • temperature: 0–2 (default: 1.1)

The plugin requests LINEAR16 (16-bit PCM WAV) chunks from Inworld so each streamed chunk is self-contained and decodes cleanly under streaming TTS; no extra configuration needed.

Steering (TTS-2)

TTS-2 takes natural-language stage directions inline with your text. Place the instruction in square brackets before the segment it should apply to:

text = (
    "[whisper in a hushed style] I have to tell you something. "
    "[laugh] Just kidding! [say with force] Now let's get to work."
)
async for chunk in await tts.stream_audio(text):
    ...

Steering covers articulation, intonation, volume, pitch, range, speed, and vocal style — and supports non-verbal sounds like [laugh], [breathe], [clear throat], [sigh], [cough], [yawn]. Combining dimensions ([whisper in a hushed style], [say playfully and very fast]) produces better results than bare single-word tags. See Inworld's steering docs and prompting guide for the full reference.

Agent example

A complete example wiring inworld.TTS() into a Stream-edge agent with Deepgram STT, Gemini LLM, and smart-turn detection lives at example/inworld_tts_example.py. The companion example/inworld-audio-guide.md is loaded as the agent's system prompt and teaches the LLM how to emit TTS-2 steering tags so replies sound expressive out of the box.

Realtime (WebRTC)

Low-latency speech-to-speech via Inworld's Realtime API. This transport uses WebRTC (UDP, native Opus) for lower latency than the WebSocket alternative. Requires a WebRTC-capable edge transport — pair with getstream.Edge() as shown below.

from vision_agents.core import Agent, User
from vision_agents.plugins import getstream, inworld, smart_turn

agent = Agent(
    edge=getstream.Edge(),
    agent_user=User(name="My Agent", id="agent"),
    llm=inworld.Realtime(
        model="openai/gpt-4o-mini",
        voice="Dennis",
        instructions="You are a friendly voice assistant.",
    ),
    turn_detection=smart_turn.TurnDetection(),
)

Realtime options

  • model: provider-prefixed model ID. Examples: "openai/gpt-4o-mini" (default), "google-ai-studio/gemini-2.5-flash", "inworld/<router-id>" for an Inworld router
  • voice: voice for audio responses (default: "Dennis"; "Clive", "Olivia" and custom voices also supported)
  • api_key: Inworld AI API key (default: reads from INWORLD_API_KEY)
  • instructions: system prompt
  • realtime_session: advanced — pass a full RealtimeSessionCreateRequestParam for session fields not exposed by the primary args (custom turn-detection, tool_choice, etc.)

Registering tools

realtime = inworld.Realtime()

@realtime.register_function(description="Get the current weather for a city.")
async def get_weather(city: str) -> str:
    return f"It's sunny in {city}."

Tools follow the OpenAI function-calling schema. Inworld's Realtime API is protocol-compatible with OpenAI's Realtime API, so registered functions flow through the same response.function_call_arguments.done path.

Notes

  • v1 is WebRTC only; a WebSocket transport may be added later.
  • Video input is not currently supported by Inworld's Realtime API.

Requirements

  • Python 3.10+
  • httpx>=0.28, av>=10, aiortc>=1.9, openai[realtime]>=2.26,<3

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

vision_agents_plugins_inworld-0.5.7.tar.gz (15.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

vision_agents_plugins_inworld-0.5.7-py3-none-any.whl (37.7 kB view details)

Uploaded Python 3

File details

Details for the file vision_agents_plugins_inworld-0.5.7.tar.gz.

File metadata

  • Download URL: vision_agents_plugins_inworld-0.5.7.tar.gz
  • Upload date:
  • Size: 15.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.10.6 {"installer":{"name":"uv","version":"0.10.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for vision_agents_plugins_inworld-0.5.7.tar.gz
Algorithm Hash digest
SHA256 12cb458d471f504cbd7a7db6c8fccce27300a8885eb89cac30062fb24bc14a58
MD5 b16a98b80194ab54fcfece5d8ff3eae0
BLAKE2b-256 ac69bed9d29a89e9854c73ac827b15da4ae9085df34dca2f3c449342f0b6a87d

See more details on using hashes here.

File details

Details for the file vision_agents_plugins_inworld-0.5.7-py3-none-any.whl.

File metadata

  • Download URL: vision_agents_plugins_inworld-0.5.7-py3-none-any.whl
  • Upload date:
  • Size: 37.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.10.6 {"installer":{"name":"uv","version":"0.10.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for vision_agents_plugins_inworld-0.5.7-py3-none-any.whl
Algorithm Hash digest
SHA256 d19e220722667809ea5c22c5a9e10d7287bef8fc0178bc4d50c63d266fea2e8e
MD5 f02535299293171f46d1e3f73f356681
BLAKE2b-256 b043a6813ead234623b7111cc5d03d099aef9dd23d325d1f4937e14f014d81dd

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page