Python SDK for AttentionLabs real-time attention detection.

These details have not been verified by PyPI

Project links

Homepage

Project description

attenlabs-saa

Python SDK for Attention Labs real-time selective auditory attention.

Every voice pipeline has the same problem: the microphone hears everything, but your ASR should only process speech directed at the device. Wake words solve this with a rigid trigger phrase. SAA solves it without one — classifying every audio frame as silent, human-directed, or device-directed and routing only what matters.

attenlabs-saa streams mic and webcam data to the SAA inference server over WebSocket and emits typed events: attention predictions, voice activity, conversation state, and ready-to-forward speech audio. LLM routing is left to you.

Sign up

Get your API token at attentionlabs.ai/dashboard.

You need your API Key for this project to work

Install

pip install attenlabs-saa

Requires Python 3.10+. sounddevice and opencv-python are pulled in automatically for mic and camera access.

Quickstart

import time
from saa import AttentionClient

client = AttentionClient(token="your-token")

@client.on_prediction
def _(event):
    label = {0: "silent", 1: "human", 2: "device"}.get(event.cls, "?")
    print(f"{label}  {event.confidence:.0%}  faces={event.num_faces}  src={event.source}")

@client.on_speech_ready
def _(event):
    # event.audio_base64 — base64 PCM16 @ 16 kHz mono, ready for OpenAI Realtime / any LLM
    # event.audio_pcm16  — same audio as np.int16 array
    print(f"speech ready ({event.duration_sec:.2f}s)")

@client.on_error
def _(event):
    print(f"ERROR: {event.title}: {event.message}")

client.start()
try:
    while True:
        time.sleep(0.1)
except KeyboardInterrupt:
    client.stop()

A full CLI demo wiring SAA + OpenAI Realtime lives at saa-py-demo.

API

`AttentionClient`

from saa import AttentionClient, CameraConfig, MicConfig

client = AttentionClient(
    token="...",                    # Auth token — sent as WS subprotocol
    url=None,                      # Server URL (default: wss://server.attentionlabs.ai/ws)
    video=CameraConfig(),          # Webcam config
    audio=MicConfig(),             # Mic config
    initial_threshold=0.7,         # Device-class confidence threshold (0..1)
    enable_audio=True,             # Set False to skip mic capture
    enable_video=True,             # Set False to skip webcam capture
)

Configuration

`MicConfig`

field	type	default	notes
`device`	`int \| str \| None`	`None`	Device index, name, or `None` for system default
`channels`	`int`	`1`	Number of input channels

`CameraConfig`

field	type	default	notes
`device_index`	`int`	`0`	Webcam device index
`width`	`int`	`1920`	Capture width
`height`	`int`	`1080`	Capture height
`jpeg_quality`	`int`	`60`	JPEG compression quality 0–100

Methods

method	description
`start()`	Opens WebSocket, acquires mic + camera, starts capture threads. Non-blocking. Raises on handshake failure.
`stop()`	Tears down capture, joins threads, closes WebSocket.
`mute()`	Pauses upstream audio and signals server to stop VAD.
`unmute()`	Resumes upstream audio.
`mark_responding(bool)`	Tell the server an LLM response is in flight. Server stops emitting predictions while `True`.
`set_threshold(value: float)`	Update device-class confidence threshold (0..1). Server acks via `config` event.

Events

Register handlers with decorators. All callbacks fire on internal threads — keep them fast or hand work off to your own thread.

@client.on_prediction
def handle(event):
    ...

decorator	payload	fires when
`@on_connected`	—	WebSocket opens
`@on_started`	—	Server-side warmup complete
`@on_warmup_complete`	—	First non-zero-confidence prediction
`@on_prediction`	`PredictionEvent`	Each attention prediction
`@on_vad`	`VadEvent`	Voice activity update
`@on_state`	`StateEvent`	Conversation state transition
`@on_speech_ready`	`SpeechReadyEvent`	Complete speech segment ready to forward
`@on_config`	`ConfigEvent`	Server acks a threshold change
`@on_stats`	`StatsEvent`	Every ~10s with connection health
`@on_interrupt`	`InterruptEvent`	User is barging in mid-LLM-response
`@on_error`	`AttentionErrorEvent`	Connection, auth, or server error
`@on_disconnected`	`DisconnectedEvent`	WebSocket closes

Event types

`PredictionEvent`

cls: int            # 0 = silent, 1 = human-directed, 2 = device-directed
confidence: float   # 0..1
source: str         # "video" or "audio"
num_faces: int      # faces detected in frame

`VadEvent`

probability: float  # VAD probability 0..1
is_speech: bool     # whether speech was detected

`StateEvent`

state: ConversationState  # "listening" | "sending" | "cancelled" | "idle"

`SpeechReadyEvent`

audio_pcm16: np.ndarray   # int16 array @ 16 kHz mono
audio_base64: str          # same audio as base64 — ready for OpenAI Realtime, etc.
duration_sec: float        # duration in seconds

`ConfigEvent`

model_class2_threshold: float  # server-confirmed threshold

`StatsEvent`

rtt_ms: float | None  # round-trip latency in ms
sent_video: int        # total video frames sent
skipped_video: int     # total video frames skipped
sent_audio: int        # total audio chunks sent
uptime_s: float        # connection uptime in seconds

`InterruptEvent`

fade_ms: int        # suggested fade duration (ms) before stopping playback
confidence: float   # raw model confidence of the class-2 prediction that fired

Fires when the server detects the user trying to take the turn back while the LLM is mid-response. The server has already moved its state machine to listening and pre-rolled the user's recent audio into the next turn — the following turn_ready event will carry the actual barge-in question. The consumer's job is to (a) fade and stop its local LLM playback over fade_ms, (b) cancel any in-flight LLM response, and (c) re-open the mic immediately (do not wait for the fade to finish, or the user's continued speech is dropped for the duration of the fade).

`AttentionErrorEvent`

title: str                  # error category ("Auth Failed", "Connection Stalled", etc.)
message: str                # human-readable message
detail: str | None = None   # technical detail
code: int | None = None     # WebSocket close code, if applicable

`DisconnectedEvent`

code: int        # WebSocket close code
reason: str      # close reason
was_clean: bool  # True if code == 1000

LLM integration

LLM routing is intentionally not part of the SDK. The speech_ready event hands you PCM16 audio — both as a NumPy array and as base64 — forward it wherever you like.

When your LLM starts generating, call mute() + mark_responding(True) to suppress predictions during playback. When it finishes, unmute() + mark_responding(False).

from saa import AttentionClient

client = AttentionClient(token="...")

@client.on_speech_ready
def _(event):
    # Forward to your LLM of choice
    your_llm.send(event.audio_base64)

def on_llm_speaking():
    client.mute()
    client.mark_responding(True)

def on_llm_done():
    client.unmute()
    client.mark_responding(False)

Barge-in (interrupt) handling

When the server detects the user trying to take the turn back while the LLM is speaking, it fires interrupt. Wire it to a fade-and-cancel on your LLM playback layer, then re-open the mic immediately:

@client.on_interrupt
def _(event):
    # Fade your local LLM audio and cancel its in-flight response.
    your_llm.interrupt(event.fade_ms)
    # Re-open the mic immediately — do NOT wait for the fade to finish,
    # or the user's continued speech is dropped for the fade duration.
    client.unmute()
    client.mark_responding(False)

The server has already moved its state machine to listening and pre-rolled the user's recent audio into the chunk accumulator by the time this event arrives. The next turn_ready event will carry the user's actual barge-in question.

See saa-py-demo for a full working example with OpenAI Realtime.

Threading model

The SDK manages four threads internally:

thread	purpose
`saa-ws`	WebSocket send/receive
`saa-heartbeat`	JSON pings every 5s, stats every 10s
`saa-camera`	JPEG capture at 4 fps (250 ms)
(sounddevice)	Audio callback at native sample rate, resampled to 16 kHz

All event callbacks fire on saa-ws or saa-heartbeat. Don't block them — offload heavy work to your own thread.

License

MIT

Project details

These details have not been verified by PyPI

Project links

Homepage

Release history Release notifications | RSS feed

This version

0.3.1

May 19, 2026

0.3.0

May 14, 2026

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

attenlabs_saa-0.3.1.tar.gz (13.0 kB view details)

Uploaded May 19, 2026 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

attenlabs_saa-0.3.1-py3-none-any.whl (15.4 kB view details)

Uploaded May 19, 2026 Python 3

File details

Details for the file attenlabs_saa-0.3.1.tar.gz.

File metadata

Download URL: attenlabs_saa-0.3.1.tar.gz
Upload date: May 19, 2026
Size: 13.0 kB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/6.2.0 CPython/3.10.14

File hashes

Hashes for attenlabs_saa-0.3.1.tar.gz
Algorithm	Hash digest
SHA256	`cd8578660a192f2287bb1f792c106530a99ad6bc0118a08dc4e81ca272194bfc`
MD5	`c5c46e05ae745349723257e3401ced68`
BLAKE2b-256	`c613fb63dd5c749d5e785c8ffd05bb8d397484d5a7e0a11949b693915570ef69`

See more details on using hashes here.

File details

Details for the file attenlabs_saa-0.3.1-py3-none-any.whl.

File metadata

Download URL: attenlabs_saa-0.3.1-py3-none-any.whl
Upload date: May 19, 2026
Size: 15.4 kB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: twine/6.2.0 CPython/3.10.14

File hashes

Hashes for attenlabs_saa-0.3.1-py3-none-any.whl
Algorithm	Hash digest
SHA256	`c41f5df6c47a7f8c68dfb73ef1d1e3947c444c2a935b1e4e093de9093ee9c956`
MD5	`e452e2e68d0b53c0bef8617be0b9fa1c`
BLAKE2b-256	`5191bea849e34c0ad1bf6e6dd558169ff8e975009652fe4e114a6993e19f888f`

See more details on using hashes here.

attenlabs-saa 0.3.1

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

attenlabs-saa

Sign up

Install

Quickstart

API

AttentionClient

Configuration

MicConfig

CameraConfig

Methods

Events

Event types

PredictionEvent

VadEvent

StateEvent

SpeechReadyEvent

ConfigEvent

StatsEvent

InterruptEvent

AttentionErrorEvent

DisconnectedEvent

LLM integration

Barge-in (interrupt) handling

Threading model

License

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes

`AttentionClient`

`MicConfig`

`CameraConfig`

`PredictionEvent`

`VadEvent`

`StateEvent`

`SpeechReadyEvent`

`ConfigEvent`

`StatsEvent`

`InterruptEvent`

`AttentionErrorEvent`

`DisconnectedEvent`