Skip to main content

autourgos-live

Framework: Autourgos Python License: Apache 2.0 Author Contributor Contributor

A self-contained, async wrapper for live/realtime bidirectional-streaming LLM APIs — persistent WebSocket sessions that stream text, audio, and video both ways, as opposed to the single call-and-reply request/response wrappers in autourgos-openaichat and autourgos-responses. Part of the Autourgos agentic-AI framework, but has zero dependency on it: pip install websockets and you're ready. Talks to each provider over a raw WebSocket connection using that provider's own documented wire protocol — no vendor SDK dependency.

import asyncio
from autourgos_live import GeminiLiveSession, TranscriptDelta, TurnComplete

async def main():
    async with GeminiLiveSession(
        model="gemini-3.1-flash-live-preview",  # reads GEMINI_API_KEY
        enable_output_transcription=True,  # required to get text back -- see note below
    ) as session:
        await session.send_text("What is the capital of France?")
        async for event in session:
            if isinstance(event, TranscriptDelta) and event.role == "model":
                print(event.text, end="", flush=True)
            elif isinstance(event, TurnComplete):
                break

asyncio.run(main())
# Paris

Note: every Gemini Live model currently available is native-audio-only — response_modalities therefore defaults to ["AUDIO"], and there is no way to get raw TextDelta output today. Pass enable_output_transcription=True and read TranscriptDelta(role="model") instead, as above. If Google ships a text-capable live model later, pass response_modalities=["TEXT"] to get TextDelta events directly.


Features

  • One normalized event stream, any live provider: text, audio, transcript, tool-call, turn-complete, interrupted, usage, goAway, and session-resumption events all come back as the same small set of dataclasses — calling code never branches on provider-specific event names
  • Text, audio, and video input; text and audio output, with opt-in speech-to-text transcripts for both sides of the conversation
  • Native tool/function calling: pass tools=[...] (accepts autourgos_agent's @tool-decorated functions directly) and matching calls are auto-executed and replied to for you, or handle ToolCallRequested yourself with send_tool_result()
  • Built-in speaker playback (pip install autourgos-live[mic]), async-safe (doesn't block the event loop); pair with the standalone autourgos-micinput package for microphone capture (buffer-bounded, drop-oldest backpressure) — a full duplex voice conversation in ~20 lines — or bring your own audio source (telephony pipeline, browser WebRTC, file) and skip both
  • Handles the real duration limits of a live session (15 min audio-only / 2 min audio+video / ~10 min connection lifetime by default): context window compression to remove the cap, session resumption to pick a session back up on a new connection, and a GoAway event surfaced ahead of a forced disconnect
  • Manual (non-VAD) turn control for push-to-talk or noisy environments, alongside the default server-side voice activity detection
  • Async-only by design: a live session is a concurrent send/receive loop over one socket, not a call-and-reply — no forced blocking sync facade
  • No silent reconnect: a dropped socket surfaces as a SessionClosed/LiveError event instead of automatically retrying mid-turn, which could duplicate or lose in-flight audio
  • Defensive parsing throughout — an unrecognized, malformed, or transport-level failure is logged and surfaced as a LiveError event rather than crashing the receive loop
  • Fully typed (py.typed)

Table of Contents


Install

pip install autourgos-live

Requires Python 3.10+ and websockets>=14.0. Microphone/speaker helpers additionally need pip install "autourgos-live[mic]" (sounddevice).


Supported Providers

Provider Class Model examples Get a key
Google Gemini Live GeminiLiveSession gemini-3.1-flash-live-preview, gemini-2.5-flash-native-audio-preview-09-2025 https://aistudio.google.com/apikey
OpenAI Realtime (planned) — —

Additional providers slot in as another adapter implementing the shared BaseLiveSession interface — the normalized event types and calling code above don't change.


Core Usage

Text Conversation

import asyncio
from autourgos_live import GeminiLiveSession, TranscriptDelta, TurnComplete

async def main():
    async with GeminiLiveSession(
        model="gemini-3.1-flash-live-preview",
        api_key="...",  # or set GEMINI_API_KEY / GOOGLE_API_KEY
        system_instruction="You are terse.",
        enable_output_transcription=True,  # see note near the top of this README -- no TEXT-capable live model exists today
    ) as session:
        await session.send_text("Name three prime numbers.")
        async for event in session:
            if isinstance(event, TranscriptDelta) and event.role == "model":
                print(event.text, end="", flush=True)
            elif isinstance(event, TurnComplete):
                break

asyncio.run(main())
# 2, 3, 5.

Voice Conversation (mic input + built-in speaker output)

Microphone capture isn't bundled here — it lives in the standalone autourgos-micinput package, so import autourgos_live never forces that dependency on callers who don't do local audio input at all (telephony, browser WebRTC, text-only). Speaker output stays built in (SpeakerPlayer, below) since it's genuinely specific to playing back what a live session sends you.

pip install "autourgos-live[mic]" "autourgos-micinput[mic]"
import asyncio
from autourgos_live import AudioDelta, GeminiLiveSession, SpeakerPlayer, TranscriptDelta
from autourgos_micinput import MicrophoneStream

async def pipe_microphone(session, **kwargs):
    async with MicrophoneStream(**kwargs) as mic:
        async for chunk in mic:
            await session.send_audio_chunk(chunk, mime_type=mic.mime_type)

async def main():
    async with GeminiLiveSession(
        model="gemini-3.1-flash-live-preview",
        response_modalities=["AUDIO"],
        enable_input_transcription=True,   # required for TranscriptDelta(role="user") - opt-in, has its own cost
        enable_output_transcription=True,  # required for TranscriptDelta(role="model")
    ) as session:
        mic_task = asyncio.create_task(pipe_microphone(session, sample_rate=16000))
        try:
            with SpeakerPlayer(sample_rate=24000) as speaker:
                async for event in session:
                    if isinstance(event, AudioDelta):
                        await speaker.awrite(event.data)  # not speaker.write() - that blocks the event loop
                    elif isinstance(event, TranscriptDelta):
                        print(f"[{event.role}] {event.text}")
        finally:
            mic_task.cancel()

asyncio.run(main())
# [user] What's the weather like on Mars?
# [model] Mars is cold and dusty, with an average temperature of about -63°C...

Custom Audio Source

No local mic/speaker? send_audio_chunk() / AudioDelta.data are plain bytes — feed audio from a telephony pipeline, browser WebRTC, or a file without depending on autourgos-micinput/SpeakerPlayer (or the mic extra) at all:

await session.send_audio_chunk(pcm_bytes, mime_type="audio/pcm;rate=16000")

Video Input

await session.send_video_frame(jpeg_bytes, mime_type="image/jpeg")

Manual Turn Control (push-to-talk)

By default the server detects speech activity itself (VAD). For push-to-talk or noisy environments, disable that and drive turns explicitly:

async with GeminiLiveSession(
    model="gemini-3.1-flash-live-preview",
    response_modalities=["AUDIO"],
    disable_automatic_activity_detection=True,
    activity_handling="START_OF_ACTIVITY_INTERRUPTS",  # or "NO_INTERRUPTION"
) as session:
    await session.send_activity_start()
    await session.send_audio_chunk(pcm_bytes, mime_type="audio/pcm;rate=16000")
    await asyncio.sleep(1.5)  # see the timing note below -- don't skip this for scripted/synthetic audio
    await session.send_activity_end()

await session.end_audio_stream() separately signals "no more audio for now" (e.g. the mic was turned off) — distinct from ending a turn.

⚠️ Live-tested timing requirement: send_activity_end() needs real wall-clock time to have passed since the preceding audio was sent — not audio duration, chunk count, or chunk size. Calling it back-to-back after send_audio_chunk() fails 100% of the time with "Precondition check failed" (code 1007), regardless of how much audio was sent; live trials found a ~20ms gap fails consistently, ~300-600ms is flaky, and >=1.5s passes reliably (5/5 sequential trials). This is almost certainly the server needing time to process/commit the buffered activity, not a schema or ordering bug — the wire messages here are confirmed correct. In practice this rarely bites real microphone input, since a genuine push-to-talk utterance naturally spans well over a second; it mainly matters for scripted/synthetic turns (tests, demos) that send audio and immediately call send_activity_end(). See gemini.py's module docstring for the full investigation.

Transcripts

Opt in with enable_input_transcription=True/enable_output_transcription=True on the constructor (both default False — transcription is a genuine extra feature on Google's side, with its own cost/latency, not something to turn on silently). Without at least one of these, no TranscriptDelta events will ever arrive. Once enabled, transcripts arrive for that side of the conversation:

async with GeminiLiveSession(
    model="gemini-3.1-flash-live-preview",
    response_modalities=["AUDIO"],
    enable_input_transcription=True,
    enable_output_transcription=True,
) as session:
    ...

async for event in session:
    if isinstance(event, TranscriptDelta):
        print(f"[{event.role}] {event.text}")
        # [user] What's the capital of France?
        # [model] The capital of France is Paris.

Native Tool Calling

Pass tools= a list of autourgos_agent's @tool-decorated functions (or plain {"name", "description", "parameters", "func"} dicts — Tool is a dict subclass, so both work identically, no dependency on autourgos-agent required) and matching calls are auto-executed and replied to for you:

from autourgos_agent import tool
from autourgos_live import GeminiLiveSession, TurnComplete

@tool
def get_weather(city: str) -> str:
    """Get the current weather for a city."""
    return f"22°C and sunny in {city}"

async def main():
    async with GeminiLiveSession(
        model="gemini-3.1-flash-live-preview",
        tools=[get_weather],
    ) as session:
        await session.send_text("What's the weather in Tokyo?")
        async for event in session:
            if isinstance(event, TurnComplete):
                break

The model calls get_weather, the result is sent back over the socket, and the conversation continues — no ToolCallRequested handling needed. You still see the event go by if you want to log it (event.calls[0].result / .error / .dispatched are filled in after auto-dispatch). Turn dispatch off with auto_dispatch_tools=False, or leave a tool's name unregistered, to handle a call yourself:

from autourgos_live import ToolCallRequested

async for event in session:
    if isinstance(event, ToolCallRequested):
        for call in event.calls:
            if not call.dispatched:  # not auto-handled (unknown name, or auto_dispatch_tools=False)
                result = run_it_yourself(call.name, call.args)
                await session.send_tool_result(call.id, call.name, result)

Usage Tracking

from autourgos_live import UsageUpdate

async for event in session:
    if isinstance(event, UsageUpdate):
        print(event.input_tokens, event.output_tokens, event.total_tokens)
        # 12 8 20

Live-tested caveat: UsageUpdate isn't guaranteed to arrive before TurnComplete — in testing it showed up after a second TurnComplete-mapped message in the same turn (Gemini's generationComplete and turnComplete server signals both map to this library's single TurnComplete event, and can both fire per turn). If you break on the first TurnComplete — the pattern every other example on this page uses — you may miss usage data. Keep reading a little past the first TurnComplete if usage tracking matters to you.

Long-Running Sessions (duration limits, goAway, resumption)

Real Live API sessions have hard duration limits: 15 minutes for audio-only, 2 minutes for audio+video, and the underlying connection lasts ~10 minutes regardless of modality — unless you enable context window compression. Before the server closes the connection, it sends a GoAway event with a warning window. This library never reconnects automatically (matches its "no silent reconnect" design elsewhere) — you decide whether and when to.

For sessions that might run long, enable both resumption (so you can pick a session back up on a new connection) and compression (so it doesn't hit the duration cap at all):

from autourgos_live import GeminiLiveSession, GoAway, SessionResumptionUpdate

session = GeminiLiveSession(
    model="gemini-3.1-flash-live-preview",
    enable_session_resumption=True,
    enable_context_compression=True,       # removes the duration cap
    compression_trigger_tokens=16000,      # optional tuning; server defaults apply if omitted
    compression_target_tokens=8000,
)

async with session:
    await session.send_text("Let's have a long conversation.")
    async for event in session:
        if isinstance(event, GoAway):
            print(f"Disconnecting in {event.time_left_seconds}s — reconnect now")
            break
        # SessionResumptionUpdate events are tracked for you automatically;
        # no need to handle them explicitly unless you want to persist the
        # handle somewhere (a file, a DB) across process restarts.

# session.last_resumption_handle now holds the most recent handle - reconnect with it:
async with GeminiLiveSession(
    model="gemini-3.1-flash-live-preview",
    resumption_handle=session.last_resumption_handle,
) as new_session:
    ...

resumption_handle= on the constructor implies resumption is enabled — you don't need enable_session_resumption=True alongside it. A resumption handle is valid for 2 hours after the session that issued it ends.

Error Handling

from autourgos_live import GeminiLiveSession, LiveSessionError, LiveSessionImportError

try:
    async with GeminiLiveSession(model="gemini-3.1-flash-live-preview") as session:
        ...
except LiveSessionImportError:
    print("pip install websockets")
except LiveSessionError as exc:
    print(f"Config/connection error: {exc}")

LiveError events (from within the async for event in session: loop) surface provider- or transport-reported errors that happen mid-session, as opposed to setup-time failures raised as exceptions above.


Constructor Reference

GeminiLiveSession

Parameter Type Default Description
model str required e.g. "gemini-3.1-flash-live-preview" (models/ prefix added automatically if omitted)
api_key str GEMINI_API_KEY / GOOGLE_API_KEY env API key
response_modalities list[str] ["AUDIO"] ["TEXT"] or ["AUDIO"] — exactly one; Gemini Live rejects combining both (raises LiveSessionError client-side rather than a confusing server-side socket close). Defaults to AUDIO because every currently available Gemini Live model is native-audio-only and rejects TEXT outright — see the note at the top of this README
system_instruction str None System prompt
thinking_level str None "minimal", "low", "medium", or "high" (Gemini 3.x live models only)
tools list[dict] None {"name", "description", "parameters", "func"?} dicts (an autourgos_agent.Tool works directly) — declared to the model, and auto-dispatched if func is present (see Native Tool Calling)
auto_dispatch_tools bool True If False, ToolCallRequested events are yielded but never auto-executed/replied to, even for tools with a func
enable_input_transcription bool False Required for TranscriptDelta(role="user") events — the server never sends a transcript unless asked (see Transcripts)
enable_output_transcription bool False Required for TranscriptDelta(role="model") events
resumption_handle str None Resume a previous session (from .last_resumption_handle, valid 2 hours) instead of starting fresh — implies enable_session_resumption=True
enable_session_resumption bool False Request periodic SessionResumptionUpdate events (auto-tracked on .last_resumption_handle) without resuming a prior session
enable_context_compression bool False Removes the session duration cap (see Long-Running Sessions)
compression_trigger_tokens int None Token count that triggers compression; server default if omitted
compression_target_tokens int None Target size to compress the context window down to; server default if omitted
disable_automatic_activity_detection bool False Turn off server-side VAD for manual turn control (see Manual Turn Control) — drive turns yourself with send_activity_start()/send_activity_end()
activity_handling str None "START_OF_ACTIVITY_INTERRUPTS" or "NO_INTERRUPTION"; validated client-side against these two values
open_timeout float 10.0 Seconds to wait for the WebSocket handshake

SpeakerPlayer (mic extra)

Parameter Type Default Description
sample_rate int 24000 PCM sample rate — 24kHz per the Live API's documented output format
channels int 1 Audio channel count
device int None Output device index; None uses the system default

MicrophoneStream's constructor reference (sample_rate, chunk_ms, channels, device, max_queue_chunks) lives in autourgos-micinput's README — this package no longer bundles it.


API Reference

BaseLiveSession methods

Method Description
await session.connect() Open the WebSocket and complete the setup handshake (called automatically by async with)
await session.send_text(text) Send a text input chunk
await session.send_audio_chunk(data, mime_type=...) Send a chunk of raw audio input
await session.send_video_frame(data, mime_type=...) Send a single video/image frame
await session.send_activity_start() / send_activity_end() Manual turn boundaries — only meaningful with disable_automatic_activity_detection=True (see Manual Turn Control)
await session.end_audio_stream() Signal no more audio for now (e.g. mic off) — distinct from ending a turn
await session.send_tool_result(call_id, name, result) Reply to a ToolCallRequested
async for event in session: Yield normalized LiveEvent objects as the server sends them
await session.close() Close the connection (called automatically by async with)

Event types (autourgos_live.events)

Event Fields Meaning
SessionOpened — Setup handshake completed
TextDelta .text A chunk of model text output
AudioDelta .data, .mime_type A chunk of model audio output
TranscriptDelta .text, .role ("user"/"model") Speech-to-text transcript chunk
ToolCallRequested .calls: list[FunctionCallRequest] The model wants function(s) invoked
FunctionCallRequest .id, .name, .args, .result, .error, .dispatched One requested tool call — .result/.error/.dispatched are filled in after auto-dispatch (see Native Tool Calling), otherwise stay None/None/False
TurnComplete — The model finished its turn
Interrupted — The turn was interrupted (e.g. barge-in)
UsageUpdate .input_tokens, .output_tokens, .total_tokens Token usage so far, if reported
SessionClosed .code, .reason The socket closed
LiveError .message A provider or transport error
GoAway .time_left, .time_left_seconds The server will close the connection soon (see Long-Running Sessions) — .time_left_seconds is a best-effort parse, .time_left keeps the raw value
SessionResumptionUpdate .handle, .resumable A resumption handle — you don't need to handle this event yourself, session.last_resumption_handle/.last_resumption_resumable are updated automatically

Every event also carries the original parsed server message on .raw, for anything not yet covered by the normalized shape.

Audio I/O helpers

Name Description
SpeakerPlayer (mic extra) Sync context manager; await .awrite(data: bytes) plays a chunk (e.g. AudioDelta.data) to the output device without blocking the event loop — use this from async code (.write() is also available but blocks the whole event loop for its duration, so only use it from genuinely sync code)

Microphone input isn't provided by this package — install autourgos-micinput separately and forward its chunks to session.send_audio_chunk() yourself (see Voice Conversation above for the three-line bridging loop).

All three raise ImportError with an install hint if the mic extra isn't installed; they're never required just to import autourgos_live.


License

Apache License 2.0, Copyright (c) 2026 Jitin Kumar Sengar

Metadata

Release files for autourgos-live 0.4.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for autourgos-live 0.4.1
File Size Uploaded
autourgos_live-0.4.1.tar.gz 52.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for autourgos-live 0.4.1
File Interpreter ABI Platform
autourgos_live-0.4.1-py3-none-any.whl Python 3 none any Details

Total release size: 89.2 kB

Release files / autourgos_live-0.4.1.tar.gz

Download URL autourgos_live-0.4.1.tar.gz
Size 52.5 kB
Tags Source
SHA-256 checksum
How to use checksums
8247737c08f7e392d563cb2fe5f45993a0eeda898a05b0244e2fddb8b3a8df7b
BLAKE2b-256 checksum
How to use checksums
9f2b9644886694b6c97743cf6bfcbbf67bb8dd4a0e26b1d19cdf62b65925a040
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release files / autourgos_live-0.4.1-py3-none-any.whl

Download URL autourgos_live-0.4.1-py3-none-any.whl
Size 36.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
274355e1fe91ee97e5a0836d90f5d23fe642a85d9b54185e74645f366f380b54
BLAKE2b-256 checksum
How to use checksums
740a2fb82f42d4ee033387b9665d3e12cb78bf030f17eedc0f7cb3039da1df5e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release history Release notifications | RSS feed

This release

0.4.1 This release

2 release files

0.4.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page