Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

AudioPlane Python SDK

The source-aware, bidirectional async SDK for the local AudioPlane Runtime. It connects only to the local Unix-domain Runtime and keeps Core Audio details out of application code. The core package has no runtime dependencies and supports Python 3.9+.

Developer preview (1.0.0rc1). Native macOS Runtime required. The Python package does not install the Runtime or virtual audio driver. Start with the installation and first-audio guide before attempting capture or playback. Core audio I/O needs Python 3.9+; Gemini, OpenAI and MCP integrations need Python 3.10+.

Install in your Python application

The prepared PyPI release uses an explicit prerelease version:

python3 -m venv .venv
. .venv/bin/activate
python -m pip install audioplane==1.0.0rc1

Until that release is available on PyPI, install from a repository checkout: python -m pip install ./SDKs/python. For provider support after publication, use python -m pip install 'audioplane[gemini]==1.0.0rc1' or python -m pip install 'audioplane[openai]==1.0.0rc1'. Provider credentials are external; no cloud provider is required for basic I/O.

Documentation:

These links reference a fixed documentation revision so they also work on PyPI and outside a repository checkout. Installation, cloud model availability and hardware behavior are not guarantees of unattended production qualification.

Install the command-line client with pipx

Once the prepared prerelease is available, install the CLI separately:

pipx install 'audioplane==1.0.0rc1'

audioplane version
audioplane doctor
audioplane sources

Before publication, the source install is:

pipx install \
  "git+https://github.com/skanda-vyas-srinivasan/audioplane.git@python-v1.0.0rc1#subdirectory=SDKs/python"

This installs the audioplane client and audioplane-mcp control server in an isolated Python environment. It does not install or sign the native macOS Runtime; Process Tap permission belongs to that native executable. Build and start the signed Runtime using the repository quickstart before running commands that connect to it.

With the first-party AudioPlane Input device installed, the CLI can keep one virtual-microphone session open and speak each line entered at the prompt:

audioplane speak

Type /quit to stop. The command uses the local macOS say voice by default; audioplane speak --voice Samantha --rate 190 selects a voice and speech rate.

When an application uses AudioPlane Input instead of a physical microphone, keep the current macOS microphone in the mix with:

audioplane mic-through

This resolves the annotated default microphone and forwards 48 kHz mono PCM through the Runtime until Ctrl-C. audioplane speak and model response audio can run concurrently. Use --source ID_OR_NAME for a different physical input; the SDK rejects a source/destination pair that is the same Core Audio device.

For local development against a checkout:

pipx install --editable SDKs/python

New Python code can use the AudioPlane name:

from audioplane import AudioPlane

async with AudioPlane() as audio:
    for source in await audio.sources():
        print(source.name)

The existing sonexis import remains available for compatibility with v1.0 applications and examples.

Install for repository development

cd /path/to/audioplane
python3 -m venv .venv
. .venv/bin/activate
python -m pip install -e ./SDKs/python
python -c 'import audioplane; print(audioplane.__version__)'

The editable install is the normal repository-development path. Release engineering also builds a wheel and source distribution locally; install the wheel into a clean environment with python -m pip install PATH_TO_WHEEL. The package contains the SDK and CLI, not a separately installed native Runtime.

For completely offline testing, installation is unnecessary:

PYTHONPATH="$PWD/SDKs/python/src" /usr/bin/python3 -m unittest discover \
  -s SDKs/python/tests -p 'test_*.py' -v

Capture one application

Start with getting started for permissions and separate Runtime installation. Choose realtime or capture-then-analysis based on your application; neither requires implementing Core Audio.

import asyncio
from audioplane import AudioPlane

async def main():
    async with AudioPlane() as sx:
        async with await sx.capture("Discord") as stream:
            async for frame in stream:
                print(frame.source.name, frame.timestamp_ns, len(frame.data))

asyncio.run(main())

Async ownership

Use context managers as the canonical ownership boundary:

  • async with AudioPlane() as sx owns the control connection and closes every capture, output, and event subscription it created.
  • Capture and playback perform an asynchronous Runtime request first, so write async with await sx.capture(...) and async with await sx.playback(...).
  • sx.session() and sx.duplex(...) return convenience context managers directly, so do not add await after async with for those calls.
  • Breaking an async for loop does not replace explicit ownership. Keep the stream inside its context, especially when cancellation or exceptions can leave the loop early.
  • reconnect() creates only a new control connection. It never revives or silently replaces resources owned by an earlier Runtime instance.

Cancellation of a mutating request closes its owning control connection when the result is ambiguous, allowing Runtime cleanup to remain authoritative. Read-only request cancellation leaves an otherwise healthy connection usable. Control requests have a 10-second deadline covering socket backpressure and the response wait. Set AudioPlane(request_timeout=30.0) to change it. A timeout raises SonexisConnectionError with code request_timeout; timed-out mutating requests reconcile ownership by closing the control connection. Positive source/destination wait timeouts also cover an in-flight discovery lookup. For compatibility, a zero timeout performs one snapshot lookup under the normal control-request deadline and raises a wait-timeout error if no match exists.

capture accepts a Runtime source ID, bundle identifier, PID passed as a Python int, exact application name, or AudioSource. A numeric string remains a string selector. A name must resolve uniquely; the SDK raises AmbiguousSourceError rather than guessing. Use find_sources, get_source, or wait_for_source for discovery and applications that launch later. Physical inputs use stable microphone:<Core Audio UID> IDs, kind "microphone", and optional is_default metadata. default_microphone() resolves the current default without guessing when metadata is ambiguous.

Every delivered AudioFrame includes its immutable source snapshot, session and stream IDs, sequence, format, session-relative timestamp, discontinuity/drop state, and SDK receipt time. estimated_capture_at_ns is an approximate Runtime presentation coordinate—not a preserved Core Audio host timestamp.

Play realtime audio

The Runtime, rather than the Python process, owns the selected output device:

import asyncio
from audioplane import AudioFormat, AudioPlane

async def play(model_audio):
    async with AudioPlane() as sx:
        async with await sx.playback(
            destination="default",
            format=AudioFormat.openai_realtime_output(),
            target_buffer_milliseconds=60,
        ) as output:
            async for pcm_chunk in model_audio:
                await output.write(pcm_chunk)

# Call `await play(model_audio)` from your application's async entry point.

write() accepts bytes, bytearray, or memoryview containing whole, interleaved PCM sample frames. Each nonempty write must contain at least 1 ms of audio; coalesce smaller model fragments first. It serializes concurrent callers and splits large chunks into protocol packets no longer than 200 ms. Awaiting write() applies Unix-socket backpressure; the SDK does not add an unbounded queue. Supply timestamp_ns= when a producer has a stream-relative presentation timestamp. Otherwise the SDK derives continuous timestamps from frames sent.

Normal context-manager exit rejects new and waiting writes, finishes the active write, then sends EOS and waits for the Runtime to drain. The three-second drain deadline includes the active write; expiry discards remaining output. await output.cancel() can override graceful close immediately: it closes the stream without EOS and asks the Runtime to discard buffered playback. flush() discards queued audio while keeping the logical session, attaches the new stream epoch returned by the Runtime, and resets sequence/timestamp state.

Use await sx.output_destinations() to enumerate typed AudioOutputDestination values. await output.refresh() or await sx.output_status(output.info.id) returns OutputInfo and OutputMetrics, including input/rendered/dropped frames, underruns, overruns, queue depth, buffered duration, conversion work, route changes, and whether a producer is attached. The default destination follows the current macOS default output device. A Runtime without the v0.4 output_sessions capability fails these calls with unsupported_capability instead of sending unsupported commands.

Use find_output_destinations(), get_output_destination(), and wait_for_output_destination() to search snapshots, resolve an exact ID or name, or wait for a device to appear. "loopback" and "virtual_input" are accepted only when they identify one unambiguous installed loopback. The SDK raises typed not-found/ambiguity errors rather than choosing silently.

Installed loopback HAL devices are returned as virtual_input destinations when recognizable and can be selected by their coreaudio:<UID> ID or exact name. AudioPlane includes a source-built first-party driver, but installation remains explicit; destination enumeration is the authoritative source of availability. await output.refresh() updates the creation-time destination snapshot after a default-device change.

Capture and output can also be owned together without imposing agent policy:

async with AudioPlane() as sx:
    async with sx.duplex(
        "Discord",
        input_format=AudioFormat.gemini_live(),
        output_format=AudioFormat.gemini_live_output(),
    ) as session:
        async for frame in session.input:
            ...
            await session.output.write(response_pcm)

The application still decides turn-taking and feedback behavior. Call await session.output.flush() to discard a partial response while keeping the duplex session, or cancel() to tear output down immediately.

For a public-API-only loop with a fake passthrough model, negotiate the same format on both sides. Real model output must instead match output_format:

import asyncio
from audioplane import AudioFormat, AudioPlane

async def main():
    format = AudioFormat.speech_16k()
    async with AudioPlane() as sx:
        async with sx.duplex(
            "Discord", input_format=format, output_format=format
        ) as session:
            async for frame in session.input:
                # Replace this passthrough with a model producing `format`.
                await session.output.write(frame.data)

asyncio.run(main())

Use headphones for passthrough/duplex experiments; AudioPlane does not implement acoustic echo cancellation.

The output data plane uses the same 64-byte SXPC v2 PCM envelope as capture, but in the client-to-Runtime direction. Continuous audio never travels in JSON control requests.

Multiple labeled sources

async with AudioPlane() as sx:
    async with sx.session() as group:
        await group.add("conversation", "Discord")
        await group.add("media", "Spotify")

        async for item in group.frames():
            print(item.label, item.source.name, item.timestamp_ns)

Captures remain independent and are never mixed. Every label has a fair queue bounded to max_queue_packets complete AudioFrame packets. group.dropped_frames and dropped_frames_by_label count lost PCM sample frames; the next retained frame reports a local discontinuity. Independent Process Taps do not promise sample-accurate cross-application synchronization.

fail_fast=True is the default: one member failure raises from frames() and the context manager closes the group. Use fail_fast=False for independent long-running members and inspect errors_by_label when one source ends.

Format presets

from audioplane import AudioFormat

AudioFormat.speech_16k()
AudioFormat.openai_realtime()  # PCM16 mono, 24 kHz
AudioFormat.gemini_live()      # PCM16 mono, 16 kHz
AudioFormat.openai_realtime_output()  # PCM16 mono, 24 kHz
AudioFormat.gemini_live_output()      # PCM16 mono, 24 kHz

These are convenience values. The Runtime handshake remains authoritative for supported formats.

Optional realtime providers

Provider adapters are isolated above the SDK and require opt-in dependencies:

The dependency-free core supports Python 3.9. Provider and MCP extras require Python 3.10 or newer because their official upstream SDKs do. Create a separate environment with any installed 3.10+ interpreter before installing an extra:

python3.10 -m venv .venv-ai
. .venv-ai/bin/activate
python -m pip install -e 'SDKs/python[openai]'
python -m pip install -e 'SDKs/python[openai]'
python -m pip install -e 'SDKs/python[gemini]'
from audioplane import AudioFormat, AudioPlane
from audioplane.providers import GeminiLiveSink, GeminiTurnDetectionConfig

async with AudioPlane() as sx:
    turns = GeminiTurnDetectionConfig(silence_duration_ms=1200)
    async with await GeminiLiveSink.connect(model="gemini-3.8-live", turn_detection=turns) as model:
        async with await sx.capture(
            "Google Chrome", format=AudioFormat.gemini_live()
        ) as stream:
            async for frame in stream:
                await model.send_audio(frame)

This fragment shows the send path only. Consume model.events() concurrently to receive text/audio and surface provider failures; the runnable streaming example and packaged audioplane agent do that. A provider stall can backpressure direct forwarding; use the packaged bounded pipeline for conversational behavior.

For a complete recording review, run python Examples/capture-and-review.py --source "Google Chrome" --duration 205 from a checkout with the Gemini extra installed and the key exported. It uses captured audio directly, not a website transcript, and retains PCM in memory.

OpenAI reads OPENAI_API_KEY; Gemini reads GEMINI_API_KEY and optionally GEMINI_LIVE_MODEL. Each sink accepts one ordered AudioPlane stream; create one sink per source label. No adapter logs or persists audio or credentials. Network behavior must be validated with the developer's own provider account.

Gemini uses hybrid VAD by default: server automatic VAD remains enabled while the adapter buffers a short activity onset and sends one audio_stream_end after a debounced local pause. Silence after finalization is not transmitted, and new meaningful activity reopens the stream. Configure start/end RMS, minimum activity, and silence duration with GeminiTurnDetectionConfig, or inject the SDK's VoiceActivityDetector extension for speech-aware detection. WebRTCVoiceActivityDetector is a provided optional classifier, installed with the Gemini extra. The packaged Gemini agent uses it by default to avoid waiting for literal silence in a noisy room. It confirms 100 ms of speech, tolerates 100 ms onset gaps, and ends after 1,200 ms of non-speech. Configure it with --gemini-vad-mode, --gemini-min-activity-ms, --gemini-onset-gap-ms, and --gemini-silence-ms. Choose --gemini-vad energy for the legacy RMS detector; RMS start/end options affect only that mode. For direct sink usage:

from audioplane import WebRTCVoiceActivityDetector
from audioplane.providers import GeminiLiveSink

model = await GeminiLiveSink.connect(
    voice_activity_detector=WebRTCVoiceActivityDetector(mode=2),
)

--debug prints input levels and detector state once per second. The classifier retains less than one 10 ms block and runs outside the Runtime/audio callback.

For the complete packaged capture/provider/playback lifecycle, use the same entry point installed by the wheel:

audioplane agent \
  --provider gemini \
  --gemini-model gemini-3.8-live \
  --source "Google Chrome" \
  --response-output coreaudio:com.audioplane.input.device \
  --gemini-barge-in \
  --validate-live \
  --debug

The base audioplane commands remain usable without provider extras. The agent imports provider integrations lazily and reports missing optional dependencies or credentials as structured, actionable failures.

Replay, activity, and diagnostics

ReplayStream.from_wav(...) and ReplayStream.from_pcm(...) yield normal source-aware frames, with optional realtime pacing, for deterministic adapter tests. measure_activity(frame) reports RMS/peak and non-silence activity; it does not claim to detect speech. Applications may implement the VoiceActivityDetector protocol with their chosen VAD.

AudioActivityDetector(ActivityDetectionConfig(...)) provides debounced activity_started / activity_ended edges with source/session/stream context. It runs on the consuming task, resets on discontinuities, and accepts an optional VoiceActivityDetector for speech-aware classification. Use one detector per AudioPlane stream; reset() clears debounce state but deliberately retains stream affinity.

LatencyTracker keeps a bounded sample window and reports p50/p95/p99 estimates from Runtime session presentation time to SDK receipt. These estimates exclude provider response/network latency and are not live Process Tap capture latency.

MCP control plane

The optional MCP server requires Python 3.10+ and the mcp extra:

python -m pip install -e 'SDKs/python[mcp]'
python -m sonexis.mcp_server

Runtime/source/diagnostic/output-destination tools are local and low-bandwidth. Capture tools are not registered unless the server is launched with --allow-capture. MCP never carries PCM. Attach an audio process through the binary data plane:

async with AudioPlane() as sx:
    async with await sx.attach_capture(session_id) as stream:
        async for frame in stream:
            ...

The MCP process owns sessions it creates, cannot inspect another client's capture, and must remain running. Session results omit binary socket paths; closing its Runtime control connection stops those captures.

Failure and reconnect behavior

Errors preserve a stable code, message, retryable, and details mapping. Specialized exceptions cover missing/ambiguous/unavailable sources, permission, formats, session limits, slow consumers, protocol failures, connection failures, capture failures, and provider failures. reconnect() creates a fresh control connection but never silently recreates captures or binds to a relaunched app.

PermissionDeniedError maps an explicit permission_denied Runtime response. Core Audio does not consistently distinguish a denied Process Tap from other initialization failures, so missing permission may instead raise CaptureFailedError with code capture_initialization_failed and actionable permission guidance. In either case, enable sonexis-runtime in System Settings > Privacy & Security > Screen & System Audio Recording, restart the Runtime, and retry live capture; source enumeration alone does not verify it.

Metadata

Release files for audioplane 1.0.0rc1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for audioplane 1.0.0rc1
File Size Uploaded
audioplane-1.0.0rc1.tar.gz 134.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for audioplane 1.0.0rc1
File Interpreter ABI Platform
audioplane-1.0.0rc1-py3-none-any.whl Python 3 none any Details

Total release size: 225.4 kB

Release files / audioplane-1.0.0rc1.tar.gz

Download URL audioplane-1.0.0rc1.tar.gz
Size 134.3 kB
Tags Source
SHA-256 checksum
How to use checksums
ab606366f7077482b9fa0e6d94335eac64274758b2250641310b43b63c4fe9aa
BLAKE2b-256 checksum
How to use checksums
187ff97e1b1afc7766e5370bc376edfda1c49ad19edf71c723bc2018fca87731
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release files / audioplane-1.0.0rc1-py3-none-any.whl

Download URL audioplane-1.0.0rc1-py3-none-any.whl
Size 91.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
bf75acff2e6d36a74d4db18e1bcb06bb85b5c0c39acd35978a91edbb18b29629
BLAKE2b-256 checksum
How to use checksums
cebdd2bdf43e12c84cb627ef48a86a88ecdd39ade60edc9bb0141015b496255c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release history Release notifications | RSS feed

This release

1.0.0rc1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page