autourgos-live
A self-contained, async wrapper for live/realtime bidirectional-streaming LLM APIs — persistent WebSocket sessions that stream text, audio, and video both ways, as opposed to the single call-and-reply request/response wrappers in autourgos-openaichat and autourgos-responses. Part of the Autourgos agentic-AI framework, but has zero dependency on it: pip install websockets and you're ready. Talks to each provider over a raw WebSocket connection using that provider's own documented wire protocol — no vendor SDK dependency.
import asyncio
from autourgos_live import GeminiLiveSession, TranscriptDelta, TurnComplete
async def main():
async with GeminiLiveSession(
model="gemini-3.1-flash-live-preview", # reads GEMINI_API_KEY
enable_output_transcription=True, # required to get text back -- see note below
) as session:
await session.send_text("What is the capital of France?")
async for event in session:
if isinstance(event, TranscriptDelta) and event.role == "model":
print(event.text, end="", flush=True)
elif isinstance(event, TurnComplete):
break
asyncio.run(main())
# Paris
Note: every Gemini Live model currently available is native-audio-only —
response_modalitiestherefore defaults to["AUDIO"], and there is no way to get rawTextDeltaoutput today. Passenable_output_transcription=Trueand readTranscriptDelta(role="model")instead, as above. If Google ships a text-capable live model later, passresponse_modalities=["TEXT"]to getTextDeltaevents directly.
Features
- One normalized event stream, any live provider: text, audio, transcript, tool-call, turn-complete, interrupted, usage,
goAway, and session-resumption events all come back as the same small set of dataclasses — calling code never branches on provider-specific event names - Text, audio, and video input; text and audio output, with opt-in speech-to-text transcripts for both sides of the conversation
- Native tool/function calling: pass
tools=[...](acceptsautourgos_agent's@tool-decorated functions directly) and matching calls are auto-executed and replied to for you, or handleToolCallRequestedyourself withsend_tool_result() - Built-in speaker playback (
pip install autourgos-live[mic]), async-safe (doesn't block the event loop); pair with the standalone autourgos-micinput package for microphone capture (buffer-bounded, drop-oldest backpressure) — a full duplex voice conversation in ~20 lines — or bring your own audio source (telephony pipeline, browser WebRTC, file) and skip both - Handles the real duration limits of a live session (15 min audio-only / 2 min audio+video / ~10 min connection lifetime by default): context window compression to remove the cap, session resumption to pick a session back up on a new connection, and a
GoAwayevent surfaced ahead of a forced disconnect - Manual (non-VAD) turn control for push-to-talk or noisy environments, alongside the default server-side voice activity detection
- Async-only by design: a live session is a concurrent send/receive loop over one socket, not a call-and-reply — no forced blocking sync facade
- No silent reconnect: a dropped socket surfaces as a
SessionClosed/LiveErrorevent instead of automatically retrying mid-turn, which could duplicate or lose in-flight audio - Defensive parsing throughout — an unrecognized, malformed, or transport-level failure is logged and surfaced as a
LiveErrorevent rather than crashing the receive loop - Fully typed (
py.typed)
Table of Contents
Install
pip install autourgos-live
Requires Python 3.10+ and websockets>=14.0. Microphone/speaker helpers additionally need pip install "autourgos-live[mic]" (sounddevice).
Supported Providers
| Provider | Class | Model examples | Get a key |
|---|---|---|---|
| Google Gemini Live | GeminiLiveSession |
gemini-3.1-flash-live-preview, gemini-2.5-flash-native-audio-preview-09-2025 |
https://aistudio.google.com/apikey |
| OpenAI Realtime | (planned) | — | — |
Additional providers slot in as another adapter implementing the shared BaseLiveSession interface — the normalized event types and calling code above don't change.
Core Usage
Text Conversation
import asyncio
from autourgos_live import GeminiLiveSession, TranscriptDelta, TurnComplete
async def main():
async with GeminiLiveSession(
model="gemini-3.1-flash-live-preview",
api_key="...", # or set GEMINI_API_KEY / GOOGLE_API_KEY
system_instruction="You are terse.",
enable_output_transcription=True, # see note near the top of this README -- no TEXT-capable live model exists today
) as session:
await session.send_text("Name three prime numbers.")
async for event in session:
if isinstance(event, TranscriptDelta) and event.role == "model":
print(event.text, end="", flush=True)
elif isinstance(event, TurnComplete):
break
asyncio.run(main())
# 2, 3, 5.
Voice Conversation (mic input + built-in speaker output)
Microphone capture isn't bundled here — it lives in the standalone autourgos-micinput package, so import autourgos_live never forces that dependency on callers who don't do local audio input at all (telephony, browser WebRTC, text-only). Speaker output stays built in (SpeakerPlayer, below) since it's genuinely specific to playing back what a live session sends you.
pip install "autourgos-live[mic]" "autourgos-micinput[mic]"
import asyncio
from autourgos_live import AudioDelta, GeminiLiveSession, SpeakerPlayer, TranscriptDelta
from autourgos_micinput import MicrophoneStream
async def pipe_microphone(session, **kwargs):
async with MicrophoneStream(**kwargs) as mic:
async for chunk in mic:
await session.send_audio_chunk(chunk, mime_type=mic.mime_type)
async def main():
async with GeminiLiveSession(
model="gemini-3.1-flash-live-preview",
response_modalities=["AUDIO"],
enable_input_transcription=True, # required for TranscriptDelta(role="user") - opt-in, has its own cost
enable_output_transcription=True, # required for TranscriptDelta(role="model")
) as session:
mic_task = asyncio.create_task(pipe_microphone(session, sample_rate=16000))
try:
with SpeakerPlayer(sample_rate=24000) as speaker:
async for event in session:
if isinstance(event, AudioDelta):
await speaker.awrite(event.data) # not speaker.write() - that blocks the event loop
elif isinstance(event, TranscriptDelta):
print(f"[{event.role}] {event.text}")
finally:
mic_task.cancel()
asyncio.run(main())
# [user] What's the weather like on Mars?
# [model] Mars is cold and dusty, with an average temperature of about -63°C...
Custom Audio Source
No local mic/speaker? send_audio_chunk() / AudioDelta.data are plain bytes — feed audio from a telephony pipeline, browser WebRTC, or a file without depending on autourgos-micinput/SpeakerPlayer (or the mic extra) at all:
await session.send_audio_chunk(pcm_bytes, mime_type="audio/pcm;rate=16000")
Video Input
await session.send_video_frame(jpeg_bytes, mime_type="image/jpeg")
Manual Turn Control (push-to-talk)
By default the server detects speech activity itself (VAD). For push-to-talk or noisy environments, disable that and drive turns explicitly:
async with GeminiLiveSession(
model="gemini-3.1-flash-live-preview",
response_modalities=["AUDIO"],
disable_automatic_activity_detection=True,
activity_handling="START_OF_ACTIVITY_INTERRUPTS", # or "NO_INTERRUPTION"
) as session:
await session.send_activity_start()
await session.send_audio_chunk(pcm_bytes, mime_type="audio/pcm;rate=16000")
await asyncio.sleep(1.5) # see the timing note below -- don't skip this for scripted/synthetic audio
await session.send_activity_end()
await session.end_audio_stream() separately signals "no more audio for now" (e.g. the mic was turned off) — distinct from ending a turn.
⚠️ Live-tested timing requirement:
send_activity_end()needs real wall-clock time to have passed since the preceding audio was sent — not audio duration, chunk count, or chunk size. Calling it back-to-back aftersend_audio_chunk()fails 100% of the time with"Precondition check failed"(code 1007), regardless of how much audio was sent; live trials found a ~20ms gap fails consistently, ~300-600ms is flaky, and >=1.5s passes reliably (5/5 sequential trials). This is almost certainly the server needing time to process/commit the buffered activity, not a schema or ordering bug — the wire messages here are confirmed correct. In practice this rarely bites real microphone input, since a genuine push-to-talk utterance naturally spans well over a second; it mainly matters for scripted/synthetic turns (tests, demos) that send audio and immediately callsend_activity_end(). Seegemini.py's module docstring for the full investigation.
Transcripts
Opt in with enable_input_transcription=True/enable_output_transcription=True on the constructor (both default False — transcription is a genuine extra feature on Google's side, with its own cost/latency, not something to turn on silently). Without at least one of these, no TranscriptDelta events will ever arrive. Once enabled, transcripts arrive for that side of the conversation:
async with GeminiLiveSession(
model="gemini-3.1-flash-live-preview",
response_modalities=["AUDIO"],
enable_input_transcription=True,
enable_output_transcription=True,
) as session:
...
async for event in session:
if isinstance(event, TranscriptDelta):
print(f"[{event.role}] {event.text}")
# [user] What's the capital of France?
# [model] The capital of France is Paris.
Native Tool Calling
Pass tools= a list of autourgos_agent's @tool-decorated functions (or plain {"name", "description", "parameters", "func"} dicts — Tool is a dict subclass, so both work identically, no dependency on autourgos-agent required) and matching calls are auto-executed and replied to for you:
from autourgos_agent import tool
from autourgos_live import GeminiLiveSession, TurnComplete
@tool
def get_weather(city: str) -> str:
"""Get the current weather for a city."""
return f"22°C and sunny in {city}"
async def main():
async with GeminiLiveSession(
model="gemini-3.1-flash-live-preview",
tools=[get_weather],
) as session:
await session.send_text("What's the weather in Tokyo?")
async for event in session:
if isinstance(event, TurnComplete):
break
The model calls get_weather, the result is sent back over the socket, and the conversation continues — no ToolCallRequested handling needed. You still see the event go by if you want to log it (event.calls[0].result / .error / .dispatched are filled in after auto-dispatch). Turn dispatch off with auto_dispatch_tools=False, or leave a tool's name unregistered, to handle a call yourself:
from autourgos_live import ToolCallRequested
async for event in session:
if isinstance(event, ToolCallRequested):
for call in event.calls:
if not call.dispatched: # not auto-handled (unknown name, or auto_dispatch_tools=False)
result = run_it_yourself(call.name, call.args)
await session.send_tool_result(call.id, call.name, result)
Usage Tracking
from autourgos_live import UsageUpdate
async for event in session:
if isinstance(event, UsageUpdate):
print(event.input_tokens, event.output_tokens, event.total_tokens)
# 12 8 20
Live-tested caveat: UsageUpdate isn't guaranteed to arrive before TurnComplete — in testing it showed up after a second TurnComplete-mapped message in the same turn (Gemini's generationComplete and turnComplete server signals both map to this library's single TurnComplete event, and can both fire per turn). If you break on the first TurnComplete — the pattern every other example on this page uses — you may miss usage data. Keep reading a little past the first TurnComplete if usage tracking matters to you.
Long-Running Sessions (duration limits, goAway, resumption)
Real Live API sessions have hard duration limits: 15 minutes for audio-only, 2 minutes for audio+video, and the underlying connection lasts ~10 minutes regardless of modality — unless you enable context window compression. Before the server closes the connection, it sends a GoAway event with a warning window. This library never reconnects automatically (matches its "no silent reconnect" design elsewhere) — you decide whether and when to.
For sessions that might run long, enable both resumption (so you can pick a session back up on a new connection) and compression (so it doesn't hit the duration cap at all):
from autourgos_live import GeminiLiveSession, GoAway, SessionResumptionUpdate
session = GeminiLiveSession(
model="gemini-3.1-flash-live-preview",
enable_session_resumption=True,
enable_context_compression=True, # removes the duration cap
compression_trigger_tokens=16000, # optional tuning; server defaults apply if omitted
compression_target_tokens=8000,
)
async with session:
await session.send_text("Let's have a long conversation.")
async for event in session:
if isinstance(event, GoAway):
print(f"Disconnecting in {event.time_left_seconds}s — reconnect now")
break
# SessionResumptionUpdate events are tracked for you automatically;
# no need to handle them explicitly unless you want to persist the
# handle somewhere (a file, a DB) across process restarts.
# session.last_resumption_handle now holds the most recent handle - reconnect with it:
async with GeminiLiveSession(
model="gemini-3.1-flash-live-preview",
resumption_handle=session.last_resumption_handle,
) as new_session:
...
resumption_handle= on the constructor implies resumption is enabled — you don't need enable_session_resumption=True alongside it. A resumption handle is valid for 2 hours after the session that issued it ends.
Error Handling
from autourgos_live import GeminiLiveSession, LiveSessionError, LiveSessionImportError
try:
async with GeminiLiveSession(model="gemini-3.1-flash-live-preview") as session:
...
except LiveSessionImportError:
print("pip install websockets")
except LiveSessionError as exc:
print(f"Config/connection error: {exc}")
LiveError events (from within the async for event in session: loop) surface provider- or transport-reported errors that happen mid-session, as opposed to setup-time failures raised as exceptions above.
Constructor Reference
GeminiLiveSession
| Parameter | Type | Default | Description |
|---|---|---|---|
model |
str |
required | e.g. "gemini-3.1-flash-live-preview" (models/ prefix added automatically if omitted) |
api_key |
str |
GEMINI_API_KEY / GOOGLE_API_KEY env |
API key |
response_modalities |
list[str] |
["AUDIO"] |
["TEXT"] or ["AUDIO"] — exactly one; Gemini Live rejects combining both (raises LiveSessionError client-side rather than a confusing server-side socket close). Defaults to AUDIO because every currently available Gemini Live model is native-audio-only and rejects TEXT outright — see the note at the top of this README |
system_instruction |
str |
None |
System prompt |
thinking_level |
str |
None |
"minimal", "low", "medium", or "high" (Gemini 3.x live models only) |
tools |
list[dict] |
None |
{"name", "description", "parameters", "func"?} dicts (an autourgos_agent.Tool works directly) — declared to the model, and auto-dispatched if func is present (see Native Tool Calling) |
auto_dispatch_tools |
bool |
True |
If False, ToolCallRequested events are yielded but never auto-executed/replied to, even for tools with a func |
enable_input_transcription |
bool |
False |
Required for TranscriptDelta(role="user") events — the server never sends a transcript unless asked (see Transcripts) |
enable_output_transcription |
bool |
False |
Required for TranscriptDelta(role="model") events |
resumption_handle |
str |
None |
Resume a previous session (from .last_resumption_handle, valid 2 hours) instead of starting fresh — implies enable_session_resumption=True |
enable_session_resumption |
bool |
False |
Request periodic SessionResumptionUpdate events (auto-tracked on .last_resumption_handle) without resuming a prior session |
enable_context_compression |
bool |
False |
Removes the session duration cap (see Long-Running Sessions) |
compression_trigger_tokens |
int |
None |
Token count that triggers compression; server default if omitted |
compression_target_tokens |
int |
None |
Target size to compress the context window down to; server default if omitted |
disable_automatic_activity_detection |
bool |
False |
Turn off server-side VAD for manual turn control (see Manual Turn Control) — drive turns yourself with send_activity_start()/send_activity_end() |
activity_handling |
str |
None |
"START_OF_ACTIVITY_INTERRUPTS" or "NO_INTERRUPTION"; validated client-side against these two values |
open_timeout |
float |
10.0 |
Seconds to wait for the WebSocket handshake |
SpeakerPlayer (mic extra)
| Parameter | Type | Default | Description |
|---|---|---|---|
sample_rate |
int |
24000 |
PCM sample rate — 24kHz per the Live API's documented output format |
channels |
int |
1 |
Audio channel count |
device |
int |
None |
Output device index; None uses the system default |
MicrophoneStream's constructor reference (sample_rate, chunk_ms, channels, device, max_queue_chunks) lives in autourgos-micinput's README — this package no longer bundles it.
API Reference
BaseLiveSession methods
| Method | Description |
|---|---|
await session.connect() |
Open the WebSocket and complete the setup handshake (called automatically by async with) |
await session.send_text(text) |
Send a text input chunk |
await session.send_audio_chunk(data, mime_type=...) |
Send a chunk of raw audio input |
await session.send_video_frame(data, mime_type=...) |
Send a single video/image frame |
await session.send_activity_start() / send_activity_end() |
Manual turn boundaries — only meaningful with disable_automatic_activity_detection=True (see Manual Turn Control) |
await session.end_audio_stream() |
Signal no more audio for now (e.g. mic off) — distinct from ending a turn |
await session.send_tool_result(call_id, name, result) |
Reply to a ToolCallRequested |
async for event in session: |
Yield normalized LiveEvent objects as the server sends them |
await session.close() |
Close the connection (called automatically by async with) |
Event types (autourgos_live.events)
| Event | Fields | Meaning |
|---|---|---|
SessionOpened |
— | Setup handshake completed |
TextDelta |
.text |
A chunk of model text output |
AudioDelta |
.data, .mime_type |
A chunk of model audio output |
TranscriptDelta |
.text, .role ("user"/"model") |
Speech-to-text transcript chunk |
ToolCallRequested |
.calls: list[FunctionCallRequest] |
The model wants function(s) invoked |
FunctionCallRequest |
.id, .name, .args, .result, .error, .dispatched |
One requested tool call — .result/.error/.dispatched are filled in after auto-dispatch (see Native Tool Calling), otherwise stay None/None/False |
TurnComplete |
— | The model finished its turn |
Interrupted |
— | The turn was interrupted (e.g. barge-in) |
UsageUpdate |
.input_tokens, .output_tokens, .total_tokens |
Token usage so far, if reported |
SessionClosed |
.code, .reason |
The socket closed |
LiveError |
.message |
A provider or transport error |
GoAway |
.time_left, .time_left_seconds |
The server will close the connection soon (see Long-Running Sessions) — .time_left_seconds is a best-effort parse, .time_left keeps the raw value |
SessionResumptionUpdate |
.handle, .resumable |
A resumption handle — you don't need to handle this event yourself, session.last_resumption_handle/.last_resumption_resumable are updated automatically |
Every event also carries the original parsed server message on .raw, for anything not yet covered by the normalized shape.
Audio I/O helpers
| Name | Description |
|---|---|
SpeakerPlayer (mic extra) |
Sync context manager; await .awrite(data: bytes) plays a chunk (e.g. AudioDelta.data) to the output device without blocking the event loop — use this from async code (.write() is also available but blocks the whole event loop for its duration, so only use it from genuinely sync code) |
Microphone input isn't provided by this package — install autourgos-micinput separately and forward its chunks to session.send_audio_chunk() yourself (see Voice Conversation above for the three-line bridging loop).
All three raise ImportError with an install hint if the mic extra isn't installed; they're never required just to import autourgos_live.
License
Apache License 2.0, Copyright (c) 2026 Jitin Kumar Sengar
Metadata
Release files for autourgos-live 0.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| autourgos_live-0.4.0.tar.gz | 49.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| autourgos_live-0.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 84.4 kB
Release files / autourgos_live-0.4.0.tar.gz
| Download URL | autourgos_live-0.4.0.tar.gz |
|---|---|
| Size | 49.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8fbccfc33e516db1b6518a6eda68d95ed30eb856d79a2896113db0ede87469a7
|
|
BLAKE2b-256 checksum How to use checksums |
2f9784b31fcb3b3a3039621b8ae52a21ece9e40827e15fd13b22a04d8f35c03c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.9
|
Release files / autourgos_live-0.4.0-py3-none-any.whl
| Download URL | autourgos_live-0.4.0-py3-none-any.whl |
|---|---|
| Size | 34.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
64f030f6d5d17780fc4106f38235afa8dbe10d7f5a9374df16880d7e25f748f6
|
|
BLAKE2b-256 checksum How to use checksums |
be28a3743b2d8acd751e3f77917cee4e8828064e4b4b359430b67bd448a9f113
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.9
|