This release is a pre-release and may not be stable for production use.
AudioPlane Python SDK
The source-aware, bidirectional async SDK for the local AudioPlane Runtime. It connects only to the local Unix-domain Runtime and keeps Core Audio details out of application code. The core package has no runtime dependencies and supports Python 3.9+.
Developer preview (1.0.0rc1). Native macOS Runtime required. The Python
package does not install the Runtime or virtual audio driver. Start with the
installation and first-audio guide
before attempting capture or playback. Core audio I/O needs Python 3.9+;
Gemini, OpenAI and MCP integrations need Python 3.10+.
Install in your Python application
The prepared PyPI release uses an explicit prerelease version:
python3 -m venv .venv
. .venv/bin/activate
python -m pip install audioplane==1.0.0rc1
Until that release is available on PyPI, install from a repository checkout:
python -m pip install ./SDKs/python.
For provider support after publication, use
python -m pip install 'audioplane[gemini]==1.0.0rc1' or
python -m pip install 'audioplane[openai]==1.0.0rc1'.
Provider credentials are external; no cloud provider is required for basic I/O.
Documentation:
- Getting started and permissions
- Realtime streaming vs. capture-and-review
- AI integration
- Playback, loopback and duplex output
- Current qualification and outstanding manual checks
These links reference a fixed documentation revision so they also work on PyPI and outside a repository checkout. Installation, cloud model availability and hardware behavior are not guarantees of unattended production qualification.
Install the command-line client with pipx
Once the prepared prerelease is available, install the CLI separately:
pipx install 'audioplane==1.0.0rc1'
audioplane version
audioplane doctor
audioplane sources
Before publication, the source install is:
pipx install \
"git+https://github.com/skanda-vyas-srinivasan/audioplane.git@python-v1.0.0rc1#subdirectory=SDKs/python"
This installs the audioplane client and audioplane-mcp control server in an
isolated Python environment. It does not install or sign the native macOS
Runtime; Process Tap permission belongs to that native executable. Build and
start the signed Runtime using the repository quickstart before running commands
that connect to it.
With the first-party AudioPlane Input device installed, the CLI can keep one
virtual-microphone session open and speak each line entered at the prompt:
audioplane speak
Type /quit to stop. The command uses the local macOS say voice by default;
audioplane speak --voice Samantha --rate 190 selects a voice and speech rate.
When an application uses AudioPlane Input instead of a physical microphone, keep the current macOS microphone in the mix with:
audioplane mic-through
This resolves the annotated default microphone and forwards 48 kHz mono PCM
through the Runtime until Ctrl-C. audioplane speak and model response audio
can run concurrently. Use --source ID_OR_NAME for a different physical input;
the SDK rejects a source/destination pair that is the same Core Audio device.
For local development against a checkout:
pipx install --editable SDKs/python
New Python code can use the AudioPlane name:
from audioplane import AudioPlane
async with AudioPlane() as audio:
for source in await audio.sources():
print(source.name)
The existing sonexis import remains available for compatibility with v1.0
applications and examples.
Install for repository development
cd /path/to/audioplane
python3 -m venv .venv
. .venv/bin/activate
python -m pip install -e ./SDKs/python
python -c 'import audioplane; print(audioplane.__version__)'
The editable install is the normal repository-development path. Release
engineering also builds a wheel and source distribution locally; install the
wheel into a clean environment with python -m pip install PATH_TO_WHEEL.
The package contains the SDK and CLI, not a separately installed native Runtime.
For completely offline testing, installation is unnecessary:
PYTHONPATH="$PWD/SDKs/python/src" /usr/bin/python3 -m unittest discover \
-s SDKs/python/tests -p 'test_*.py' -v
Capture one application
Start with getting started for permissions and separate Runtime installation. Choose realtime or capture-then-analysis based on your application; neither requires implementing Core Audio.
import asyncio
from audioplane import AudioPlane
async def main():
async with AudioPlane() as sx:
async with await sx.capture("Discord") as stream:
async for frame in stream:
print(frame.source.name, frame.timestamp_ns, len(frame.data))
asyncio.run(main())
Async ownership
Use context managers as the canonical ownership boundary:
async with AudioPlane() as sxowns the control connection and closes every capture, output, and event subscription it created.- Capture and playback perform an asynchronous Runtime request first, so write
async with await sx.capture(...)andasync with await sx.playback(...). sx.session()andsx.duplex(...)return convenience context managers directly, so do not addawaitafterasync withfor those calls.- Breaking an
async forloop does not replace explicit ownership. Keep the stream inside its context, especially when cancellation or exceptions can leave the loop early. reconnect()creates only a new control connection. It never revives or silently replaces resources owned by an earlier Runtime instance.
Cancellation of a mutating request closes its owning control connection when
the result is ambiguous, allowing Runtime cleanup to remain authoritative.
Read-only request cancellation leaves an otherwise healthy connection usable.
Control requests have a 10-second deadline covering socket backpressure and the
response wait. Set AudioPlane(request_timeout=30.0) to change it. A timeout
raises SonexisConnectionError with code request_timeout; timed-out mutating
requests reconcile ownership by closing the control connection.
Positive source/destination wait timeouts also cover an in-flight discovery
lookup. For compatibility, a zero timeout performs one snapshot lookup under
the normal control-request deadline and raises a wait-timeout error if no match exists.
capture accepts a Runtime source ID, bundle identifier, PID passed as a Python
int, exact application name, or AudioSource. A numeric string remains a
string selector. A name must resolve uniquely; the SDK raises
AmbiguousSourceError rather than guessing. Use find_sources, get_source,
or wait_for_source for discovery and applications that launch later.
Physical inputs use stable microphone:<Core Audio UID> IDs, kind
"microphone", and optional is_default metadata. default_microphone()
resolves the current default without guessing when metadata is ambiguous.
Every delivered AudioFrame includes its immutable source snapshot, session and
stream IDs, sequence, format, session-relative timestamp, discontinuity/drop
state, and SDK receipt time. estimated_capture_at_ns is an approximate Runtime
presentation coordinate—not a preserved Core Audio host timestamp.
Play realtime audio
The Runtime, rather than the Python process, owns the selected output device:
import asyncio
from audioplane import AudioFormat, AudioPlane
async def play(model_audio):
async with AudioPlane() as sx:
async with await sx.playback(
destination="default",
format=AudioFormat.openai_realtime_output(),
target_buffer_milliseconds=60,
) as output:
async for pcm_chunk in model_audio:
await output.write(pcm_chunk)
# Call `await play(model_audio)` from your application's async entry point.
write() accepts bytes, bytearray, or memoryview containing whole,
interleaved PCM sample frames. Each nonempty write must contain at least 1 ms
of audio; coalesce smaller model fragments first. It serializes concurrent callers and splits
large chunks into protocol packets no longer than 200 ms. Awaiting write()
applies Unix-socket backpressure; the SDK does not add an unbounded queue.
Supply timestamp_ns= when a producer has a stream-relative presentation
timestamp. Otherwise the SDK derives continuous timestamps from frames sent.
Normal context-manager exit rejects new and waiting writes, finishes the active
write, then sends EOS and waits for the Runtime to drain. The three-second drain
deadline includes the active write; expiry discards remaining output.
await output.cancel() can override graceful close immediately: it closes the stream
without EOS and asks the Runtime to discard buffered playback. flush()
discards queued audio while keeping the logical session, attaches the new
stream epoch returned by the Runtime, and resets sequence/timestamp state.
Use await sx.output_destinations() to enumerate typed
AudioOutputDestination values. await output.refresh() or
await sx.output_status(output.info.id) returns OutputInfo and
OutputMetrics, including input/rendered/dropped frames, underruns, overruns,
queue depth, buffered duration, conversion work, route changes, and whether a
producer is attached. The default destination follows the current macOS
default output device. A Runtime without the v0.4 output_sessions capability
fails these calls with unsupported_capability instead of sending unsupported
commands.
Use find_output_destinations(), get_output_destination(), and
wait_for_output_destination() to search snapshots, resolve an exact ID or
name, or wait for a device to appear. "loopback" and "virtual_input" are
accepted only when they identify one unambiguous installed loopback. The SDK
raises typed not-found/ambiguity errors rather than choosing silently.
Installed loopback HAL devices are returned as virtual_input destinations
when recognizable and can be selected by their coreaudio:<UID> ID or exact
name. AudioPlane includes a source-built first-party driver, but installation
remains explicit; destination enumeration is the authoritative source of
availability. await output.refresh() updates the
creation-time destination snapshot after a default-device change.
Capture and output can also be owned together without imposing agent policy:
async with AudioPlane() as sx:
async with sx.duplex(
"Discord",
input_format=AudioFormat.gemini_live(),
output_format=AudioFormat.gemini_live_output(),
) as session:
async for frame in session.input:
...
await session.output.write(response_pcm)
The application still decides turn-taking and feedback behavior. Call
await session.output.flush() to discard a partial response while keeping the
duplex session, or cancel() to tear output down immediately.
For a public-API-only loop with a fake passthrough model, negotiate the same
format on both sides. Real model output must instead match output_format:
import asyncio
from audioplane import AudioFormat, AudioPlane
async def main():
format = AudioFormat.speech_16k()
async with AudioPlane() as sx:
async with sx.duplex(
"Discord", input_format=format, output_format=format
) as session:
async for frame in session.input:
# Replace this passthrough with a model producing `format`.
await session.output.write(frame.data)
asyncio.run(main())
Use headphones for passthrough/duplex experiments; AudioPlane does not implement acoustic echo cancellation.
The output data plane uses the same 64-byte SXPC v2 PCM envelope as capture, but in the client-to-Runtime direction. Continuous audio never travels in JSON control requests.
Multiple labeled sources
async with AudioPlane() as sx:
async with sx.session() as group:
await group.add("conversation", "Discord")
await group.add("media", "Spotify")
async for item in group.frames():
print(item.label, item.source.name, item.timestamp_ns)
Captures remain independent and are never mixed. Every label has a fair queue
bounded to max_queue_packets complete AudioFrame packets. group.dropped_frames and
dropped_frames_by_label count lost PCM sample frames; the next retained frame
reports a local discontinuity. Independent Process Taps do not promise sample-accurate
cross-application synchronization.
fail_fast=True is the default: one member failure raises from frames() and
the context manager closes the group. Use fail_fast=False for independent
long-running members and inspect errors_by_label when one source ends.
Format presets
from audioplane import AudioFormat
AudioFormat.speech_16k()
AudioFormat.openai_realtime() # PCM16 mono, 24 kHz
AudioFormat.gemini_live() # PCM16 mono, 16 kHz
AudioFormat.openai_realtime_output() # PCM16 mono, 24 kHz
AudioFormat.gemini_live_output() # PCM16 mono, 24 kHz
These are convenience values. The Runtime handshake remains authoritative for supported formats.
Optional realtime providers
Provider adapters are isolated above the SDK and require opt-in dependencies:
The dependency-free core supports Python 3.9. Provider and MCP extras require Python 3.10 or newer because their official upstream SDKs do. Create a separate environment with any installed 3.10+ interpreter before installing an extra:
python3.10 -m venv .venv-ai
. .venv-ai/bin/activate
python -m pip install -e 'SDKs/python[openai]'
python -m pip install -e 'SDKs/python[openai]'
python -m pip install -e 'SDKs/python[gemini]'
from audioplane import AudioFormat, AudioPlane
from audioplane.providers import GeminiLiveSink, GeminiTurnDetectionConfig
async with AudioPlane() as sx:
turns = GeminiTurnDetectionConfig(silence_duration_ms=1200)
async with await GeminiLiveSink.connect(model="gemini-3.8-live", turn_detection=turns) as model:
async with await sx.capture(
"Google Chrome", format=AudioFormat.gemini_live()
) as stream:
async for frame in stream:
await model.send_audio(frame)
This fragment shows the send path only. Consume model.events() concurrently
to receive text/audio and surface provider failures; the runnable streaming
example
and packaged audioplane agent do that. A provider stall can backpressure direct
forwarding; use the packaged bounded pipeline for conversational behavior.
For a complete recording review, run
python Examples/capture-and-review.py --source "Google Chrome" --duration 205
from a checkout with the Gemini extra installed and the key exported. It uses
captured audio directly, not a website transcript, and retains PCM in memory.
OpenAI reads OPENAI_API_KEY; Gemini reads GEMINI_API_KEY and optionally
GEMINI_LIVE_MODEL. Each sink accepts one ordered AudioPlane stream; create one
sink per source label. No adapter logs or persists audio or credentials. Network
behavior must be validated with the developer's own provider account.
Gemini uses hybrid VAD by default: server automatic VAD remains enabled while
the adapter buffers a short activity onset and sends one audio_stream_end
after a debounced local pause. Silence after finalization is not transmitted,
and new meaningful activity reopens the stream. Configure start/end RMS,
minimum activity, and silence duration with GeminiTurnDetectionConfig, or
inject the SDK's VoiceActivityDetector extension for speech-aware detection.
WebRTCVoiceActivityDetector is a provided optional classifier, installed with
the Gemini extra. The packaged Gemini agent uses it by default to avoid waiting
for literal silence in a noisy room. It confirms 100 ms of speech, tolerates
100 ms onset gaps, and ends after 1,200 ms of non-speech. Configure it with
--gemini-vad-mode, --gemini-min-activity-ms, --gemini-onset-gap-ms, and
--gemini-silence-ms. Choose --gemini-vad energy for the legacy RMS detector;
RMS start/end options affect only that mode. For direct sink usage:
from audioplane import WebRTCVoiceActivityDetector
from audioplane.providers import GeminiLiveSink
model = await GeminiLiveSink.connect(
voice_activity_detector=WebRTCVoiceActivityDetector(mode=2),
)
--debug prints input levels and detector state once per second. The classifier
retains less than one 10 ms block and runs outside the Runtime/audio callback.
For the complete packaged capture/provider/playback lifecycle, use the same entry point installed by the wheel:
audioplane agent \
--provider gemini \
--gemini-model gemini-3.8-live \
--source "Google Chrome" \
--response-output coreaudio:com.audioplane.input.device \
--gemini-barge-in \
--validate-live \
--debug
The base audioplane commands remain usable without provider extras. The
agent imports provider integrations lazily and reports missing optional
dependencies or credentials as structured, actionable failures.
Replay, activity, and diagnostics
ReplayStream.from_wav(...) and ReplayStream.from_pcm(...) yield normal
source-aware frames, with optional realtime pacing, for deterministic adapter
tests. measure_activity(frame) reports RMS/peak and non-silence activity; it
does not claim to detect speech. Applications may implement the
VoiceActivityDetector protocol with their chosen VAD.
AudioActivityDetector(ActivityDetectionConfig(...)) provides debounced
activity_started / activity_ended edges with source/session/stream context.
It runs on the consuming task, resets on discontinuities, and accepts an
optional VoiceActivityDetector for speech-aware classification. Use one
detector per AudioPlane stream; reset() clears debounce state but deliberately
retains stream affinity.
LatencyTracker keeps a bounded sample window and reports p50/p95/p99 estimates
from Runtime session presentation time to SDK receipt. These estimates exclude
provider response/network latency and are not live Process Tap capture latency.
MCP control plane
The optional MCP server requires Python 3.10+ and the mcp extra:
python -m pip install -e 'SDKs/python[mcp]'
python -m sonexis.mcp_server
Runtime/source/diagnostic/output-destination tools are local and low-bandwidth.
Capture tools are not registered unless the server is launched with
--allow-capture. MCP
never carries PCM. Attach an audio process through the binary data plane:
async with AudioPlane() as sx:
async with await sx.attach_capture(session_id) as stream:
async for frame in stream:
...
The MCP process owns sessions it creates, cannot inspect another client's capture, and must remain running. Session results omit binary socket paths; closing its Runtime control connection stops those captures.
Failure and reconnect behavior
Errors preserve a stable code, message, retryable, and details mapping.
Specialized exceptions cover missing/ambiguous/unavailable sources, permission,
formats, session limits, slow consumers, protocol failures, connection failures,
capture failures, and provider failures. reconnect() creates a fresh control
connection but never silently recreates captures or binds to a relaunched app.
PermissionDeniedError maps an explicit permission_denied Runtime response.
Core Audio does not consistently distinguish a denied Process Tap from other
initialization failures, so missing permission may instead raise
CaptureFailedError with code capture_initialization_failed and actionable
permission guidance. In either case, enable sonexis-runtime in System
Settings > Privacy & Security > Screen & System Audio Recording, restart the
Runtime, and retry live capture; source enumeration alone does not verify it.
Metadata
Release files for audioplane 1.0.0rc1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| audioplane-1.0.0rc1.tar.gz | 134.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| audioplane-1.0.0rc1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 225.4 kB
Release files / audioplane-1.0.0rc1.tar.gz
| Download URL | audioplane-1.0.0rc1.tar.gz |
|---|---|
| Size | 134.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ab606366f7077482b9fa0e6d94335eac64274758b2250641310b43b63c4fe9aa
|
|
BLAKE2b-256 checksum How to use checksums |
187ff97e1b1afc7766e5370bc376edfda1c49ad19edf71c723bc2018fca87731
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|
Release files / audioplane-1.0.0rc1-py3-none-any.whl
| Download URL | audioplane-1.0.0rc1-py3-none-any.whl |
|---|---|
| Size | 91.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
bf75acff2e6d36a74d4db18e1bcb06bb85b5c0c39acd35978a91edbb18b29629
|
|
BLAKE2b-256 checksum How to use checksums |
cebdd2bdf43e12c84cb627ef48a86a88ecdd39ade60edc9bb0141015b496255c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|