olab_audio
Audio device I/O (Mic, Speaker, recording, device enumeration,
PulseAudio port control) — the core install needs only pyaudio,
pulsectl, and numpy — plus two optional extras: resample (lightweight
cross-rate PCM conversion) and analysis (a DSP/teaching/research toolkit:
Wave, Spectrogram, Spectrum, tone/chirp/pitch synthesis, trim,
normalize, matplotlib plotting).
Extracted from ~/Projects/ofm/ofm/sensor/ub_audio.py per
docs/plans/olab_packages_reorg_plan.md's
"olab_audio v1 scope" section and Migration sequence step 5.
Installing
python3 -m venv venv
source venv/bin/activate
pip install olab-audio
Add resample, analysis, and/or mp3 as needed:
pip install "olab-audio[analysis,mp3]".
Local development, against an olab_code checkout:
pip install -e "packages/olab_audio[analysis,mp3]"
Extras
| Extra | Adds | Needed for |
|---|---|---|
| (core, default) | pyaudio, pulsectl, numpy |
Mic/Speaker/recording at the device's own native rate, device enumeration, PulseAudio port control. |
resample |
soxr |
Cross-rate recording (Mic.recordStart(samplerateRec=...) at a rate other than the mic's native one) and olab_audio.resample's resample()/StreamResampler. Not yet validated on target Raspberry Pi hardware — see the plan doc's acceptance checklist. |
analysis |
resample (soxr) + librosa, soundfile, matplotlib |
olab_audio.analysis's Wave/Spectrogram/Spectrum, tone/chirp/pitch synthesis, trim, read_wave_librosa, plotting, Recording.make_wave(), and Recording_np's explicit .resample() method. |
mp3 |
lameenc (bundled LAME bindings) |
Save a capture as MP3 or convert a 16-bit PCM WAV with wav_to_mp3(). No system ffmpeg binary is required. |
Recording at the microphone's own native sample rate — the default, and
almost always what you want — needs only the core install. Cross-rate
recording fails fast and clearly at recordStart() time (not from inside
the audio callback thread) if resample isn't installed.
API parity with the original ub_audio module: every analysis-extra
symbol (Wave, Spectrum, Spectrogram, createTone, trim,
read_wave, pitch_map, etc.) is available directly at olab_audio.<name>
— not just olab_audio.analysis.<name> — via lazy module __getattr__.
olab_audio.analysis is only actually imported the first time one of
those names is accessed, so a core-only install never pays for it, but
[analysis] installed gives you the same flat namespace the original
module had. resample() (the function) is always at olab_audio.resample
directly, matching the original API exactly — the backend module itself is
named olab_audio._resample (private) specifically to avoid that name
colliding with the function.
Quick start
import olab_audio
mics = olab_audio.get_input_devices() # ALSA pseudo-device plugins (e.g. 'vdownmix') filtered out
mic = olab_audio.Mic(deviceID=mics[0]['deviceID'])
mic.start() # queries the device's own default sample rate if none is given
mic.recordStart(filename="test.wav")
# ... let it capture some audio ...
mic.recordStop()
mic.stop()
PipeWire capture identity
On PipeWire's ALSA compatibility device, each Mic.start() open is also given
a fresh node.name/application.name identity for that open. This prevents a
WirePlumber stream-restore rule made for one start_loopback_capture() stream
from being replayed onto an unrelated process's microphone capture. Any valid
caller-provided PIPEWIRE_PROPS dictionary is retained for the open and the
environment is restored immediately afterward.
MP3 output
Install the optional encoder first:
pip install "olab-audio[mp3]"
Use an .mp3 filename to encode when saving a recording; WAV remains the
default behavior for every other filename. bitrate is optional: it defaults
to 128 kbps for rates of 16 kHz and above, and 64 kbps for 8/11.025/12 kHz.
mic.recordStart(filename="capture.mp3")
# ... let it capture some audio ...
mic.recordStop() # saves a 128 kbps MP3 for a normal 44.1/48 kHz capture
# Or select a valid constant bitrate while saving a recording manually:
mic.recording.save(filename="capture.mp3", bitrate=192)
Convert an already-saved uncompressed 16-bit PCM WAV without invoking an external command:
olab_audio.wav_to_mp3("system_audio.wav")
olab_audio.wav_to_mp3("system_audio.wav", out_filepath="share.mp3", bitrate=192)
MP3 supports only mono or stereo 16-bit PCM input and these sample rates:
8, 11.025, 12, 16, 22.05, 24, 32, 44.1, and 48 kHz. olab_audio rejects
other rates and invalid rate/bitrate combinations rather than allowing the
encoder to silently change them. If a device's native capture rate is not
MP3-compatible, first capture/resample with Recording_np.resample() (the
analysis extra) to a supported rate, then save the recording as MP3.
Loopback (system-output) capture
Record "whatever is playing on this speaker/output" via PulseAudio/PipeWire's monitor sources -- Linux with a PulseAudio or PipeWire's PulseAudio-compatible server only, ALSA host API only.
import olab_audio
loopbacks = olab_audio.get_loopback_input_devices()
loopback = loopbacks[0]
mic = olab_audio.Mic(deviceID=loopback['deviceID'])
olab_audio.start_loopback_capture(mic, loopback) # starts mic AND routes only this stream
mic.recordStart(filename="system_audio.wav")
# ... let it capture some audio ...
mic.recordStop()
mic.stop() # also removes this stream's PulseAudio source-output -- nothing to restore
Notes:
- All entries returned by
get_loopback_input_devices()share onedeviceID(the single PulseAudio/PipeWire-routed PortAudio device) --start_loopback_capture()is what actually determines which sink's audio is captured, not thedeviceID. micmust not be started before callingstart_loopback_capture()-- it callsmic.start()itself (forwarding any keyword arguments you pass aftersource, e.g.reachbackFunc=, exactly like callingmic.start()directly), then moves only that one resulting PulseAudio source-output to the selected monitor. The system default source, and every other application's capture, are never touched.- If the new capture stream can't be identified (timeout, or an ambiguous
new stream), or the move itself is rejected by PulseAudio/PipeWire,
start_loopback_capture()stopsmicand raisesRuntimeErrorrather than leaving it silently capturing the wrong source. A move call that doesn't raise is trusted as successful -- somepipewire-pulseversions don't reliably report a source-output's routing state afterward, so this doesn't attempt to re-confirm it. - Capture only receives audio actually routed to the selected sink (silent otherwise -- expected, not a bug), and uses the sink's native device rate, which may not be 44.1kHz.
Normalizing recordings
Loopback (and mic) recordings capture whatever level PulseAudio/PipeWire happened to be playing at -- e.g. a quiet per-app stream volume (PipeWire gives each playback app its own independent volume, restored per media role, separate from the sink/master volume) that's unrelated to what the system volume slider shows. Two ways to fix a too-quiet recording, depending on when you catch it:
Already saved to a WAV file -- normalize_wav() peak-normalizes the
file on disk:
import olab_audio
olab_audio.normalize_wav("system_audio.wav") # overwrites in place, peak -> 0dBFS
# or write to a separate file instead of overwriting:
olab_audio.normalize_wav("system_audio.wav", out_filepath="system_audio_normalized.wav")
Still in memory, before the first save -- skip recordStart(filename=...)
so recordStop()'s automatic save is a no-op, call Recording.normalize()
on the buffer, then save it yourself:
import olab_audio
loopback = olab_audio.get_loopback_input_devices()[0]
mic = olab_audio.Mic(deviceID=loopback['deviceID'])
olab_audio.start_loopback_capture(mic, loopback)
mic.recordStart() # no filename -- recordStop() below won't auto-save
# ... let it capture some audio ...
mic.recordStop() # writes nothing yet, since no filename was given
mic.recording.normalize() # peak -> 0dBFS, in place, before saving
mic.recording.save(filename="system_audio.wav")
mic.stop()
Both use the same peak-scaling math (loudest sample -> +-amp, default
1.0 == 0dBFS) and print a NOTE instead of raising if the audio is
silent (nothing to normalize). Neither touches PulseAudio/PipeWire device
or mixer state -- they only rescale the samples you already captured.
amp must be a finite value in (0, 1.0] -- 1.0 (0dBFS) is the loudest
peak 16-bit PCM can represent at all, so an amp above 1.0 is rejected
with ValueError rather than silently hard-clipping at 1.0 and falling
short of the peak it promised.
Recording.normalize() on a Recording_bytes only supports paInt16
(the default frmt) -- it also raises ValueError for any other PyAudio
format, since the int16 decode/re-encode it uses internally would
otherwise corrupt other bit depths/formats. There's currently no working
alternative for non-int16 capture: Mic's NumPy callback path
(Mic._callback_np()) also hardcodes an int16 decode regardless of
frmt, so Recording_np can't correctly normalize (or otherwise
process) a non-int16 capture either -- that's a separate, pre-existing
limitation of Mic itself (tracked as
issue #29), not
something normalize() works around.
TLS/security note
Unlike olab_camera, olab_audio has no network-facing streaming server
in v1 — see the plan doc's "olab_audio v1 scope" item 5 (resolved: no
Camera-style network streaming in v1, deferred to v2). No TLS/cert
concerns apply here.
Known bugs fixed during this migration
Found in the original ub_audio.py (which had zero automated tests) and
fixed here, not just carried forward:
- ALSA pseudo-device segfault risk:
get_input_devices()/get_output_devices()/get_connected_devices()now filter to real hardware (hw:-named) devices plus the safedefault/pipewire/pulsealiases — never offering resampling/mixing plugins (vdownmix,sysdefault,lavrate, etc.) as selectable inputs, since opening one as a capture stream is a C-level segfaulttry/exceptcannot catch. - Hardcoded 44.1kHz default sample rate:
Mic.start()now queries the device's own reported default rate when none is given, instead of assuming 44100Hz universally (some hardware, e.g. certain USB mics, only supports other rates). Mic.start()failure left a half-open object:self.streamis now initialized toNoneand guarded everywhere it's used, so.stop()is always safe to call — including after a failed.start()— and is idempotent.get_connected_devices()'smaxInputChannelsbug: it was populated with themaxOutputChannelsvalue (a copy-paste bug), not the actual input channel count. Fixed.ftt_freq()'sNameError: its body referencednfft, but the parameter is namedn_fft— any call would crash. Fixed.Wave.zero_pad()'sNameError: called a module-levelzero_pad()function that was never defined anywhere in the file. Implemented.Wave.__add__()'sNameError: calledwarnings.warn(...)butwarningswas never imported. Fixed.- Lazy PyAudio initialization: the module-level
audiosingleton no longer constructspyaudio.PyAudio()(which opens the whole PortAudio subsystem) unconditionally at import time —import olab_audioalone no longer touches audio hardware or fails on a machine with no audio drivers. - Embedded Whisper transcription hooks removed (
Mic.transcribeStart/transcribeStop/_thread_transcribe/_transcribePrep) — not migrated, per the plan's explicit decision. Transcription isolab_voice's territory now. - Core/analysis dependency split:
Recording_np.append()'s automatic cross-rate resampling previously calledlibrosa.resample()unconditionally on every captured chunk — heavyweight, and run even on the (default, same-rate) common case. It now skips conversion entirely when rates match, and uses a persistent, statefulStreamResampler(soxr-backed, notlibrosa) when they don't — a fresh one-shot conversion per chunk would introduce boundary artifacts and drift at every chunk edge, which the persistent converter avoids; it's flushed exactly once at save time.saveAudio()'s numpy-array save path now uses the stdlibwavemodule instead of requiringsoundfile, so basic recording never needs the DSP/teaching dependency stack. Recording.duration's frame-count bug: it waslen(self.ys) / samplerateRec— wrong forRecording_bytes(self.ysis a list of raw byte chunks, not samples) and wrong for multi-channelRecording_np(interleaved sample count overcounts by a factor ofchannels), which broketimeLimitSeccutoff behavior in both cases. Fixed with an explicit per-frame counter maintained by each subclass'sappend().Recording_bytescross-rate silent mislabeling:Mic.recordStart( samplerateRec=...)could construct aRecording_bytesat a rate different from the mic's capture rate;Recording_bytesnever resamples, so it would save the original-rate bytes into a WAV file labeled with the wrong rate — wrong playback speed/pitch, no error. Now rejected explicitly at construction.Mic.recordStart()'s failure was silently swallowed: it caught every exception (including a cross-rateRecording_npfailing becauseresampleisn't installed) and only reported it viaexcFunc, with no way for a caller to detect failure except inferring it frommic.isRecordingafterward. It now also returnsTrue/False.- Stereo/multi-channel WAV files got a mono header:
Recording.save()calledsaveAudio()without passingself.channels/self.frmt, so it silently defaulted to mono — a stereo recording's interleaved samples were written with a one-channel WAV header, doubling apparent duration and corrupting playback. Fixed by passing them through explicitly. - Cross-rate stereo recording misinterpreted as mono:
Mic._callback_np()produces a flat interleaved buffer, but that flat buffer was passed straight tosoxr.ResampleStream, which requires 2D(frames, channels)for anything but mono — silently treating a stereo buffer as a mono one twice as long.StreamResampler.process()/.flush()now reshape to(frames, channels)immediately before the soxr call and flatten the result immediately after, so every other caller still only ever sees flat/interleaved 1D data. Recording_np.resample()double-converted already-cross-rate-captured audio: after a 32kHz→16kHz capture,self.ysis already at 16kHz, but.resample()defaultedframerateOrigtoself.samplerateMic(32kHz) — treating the already-converted 16kHz data as if it were still 32kHz and converting it a second time.framerateOrignow defaults toself.samplerateRec(the rate the data is actually at), andself.samplerateRecitself is updated after an explicit resample.
Not yet done
- The future
olab_voicestreaming integration adapter (consuming nativeolab_audio.Micframes and resampling them asynchronously to a streaming STT engine's target rate, e.g. 16kHz) is out of scope for this migration — see the plan doc. - Migrating OFM's
sensor_node.pyandCoG/realtime_transcription'sAudioCapture/audio_processing.pyontoolab_audio.Mic— deferred to a follow-up commit, same pattern as the otherolab_*migrations. This closes the segfault risk currently live inrealtime_transcriptionspecifically (see the plan doc's "Consumer migration candidates"). Mic.start()'s failure is still reported only viaexcFunc(unlikerecordStart(), which also returnsTrue/False). Safe to call.stop()afterward regardless, but a documentedTrue/Falseresult (or a dedicated exception) would make consumer migration cleaner. Flagged by review as a follow-up-quality item, not a blocker.recordStop()can do disk I/O (the WAV write insave()) from inside the PortAudio callback thread when atimeLimitSeccutoff triggers it automatically from_callback_np/_callback_record_bytes. Fine for short recordings; for long ones, finalization should eventually move to a non-callback worker thread. Flagged by review as a follow-up-quality item, not a blocker.- Real Raspberry Pi hardware validation of device filtering,
sample-rate selection, and the
soxrbackend's performance — the non-hardware test suite is thorough, but real-hardware confirmation (especially before deploying to vehicles) hasn't happened yet.
Metadata
Release files for olab-audio 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| olab_audio-0.1.0.tar.gz | 63.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| olab_audio-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 109.4 kB
Release files / olab_audio-0.1.0.tar.gz
| Download URL | olab_audio-0.1.0.tar.gz |
|---|---|
| Size | 63.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f387dfc91d7badbde204d51f4172b111f5c678065a74496e194ff4cade71fd6b
|
|
BLAKE2b-256 checksum How to use checksums |
d458355e8033bde437d8f05d1d27e6663cbefb05f0a69ebb4a34f09f446a738c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.
Transparency logRelease files / olab_audio-0.1.0-py3-none-any.whl
| Download URL | olab_audio-0.1.0-py3-none-any.whl |
|---|---|
| Size | 45.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
d4747c0e36e18f925fdd092b5024495fb15148dc477ddb991a8f9f5a7dc21f68
|
|
BLAKE2b-256 checksum How to use checksums |
2b489b74ba702e80dda662857ce1fe038a48cbd9500947fc6f5cdd19426ca189
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.
Transparency log