sonarwise
Pluggable audio perception engine. Hear. Search. Retrieve.
sonarwise indexes any audio — meetings, calls, podcasts, factory floors, lectures — and makes it searchable by text, speaker, sound events, or audio similarity. Every component is pluggable: swap transcription, embedding, diarization, or storage without changing your code.
What's New in v0.2.0
- Audio Preprocessing Pipeline — DC removal, bandpass filtering, spectral-gating noise reduction, normalization, dynamic range compression
- Language Detection — Whisper, faster-whisper, or script-based heuristic detection for 99 languages
- CLAP Event Classifier — Zero-shot audio event classification using CLAP text-audio similarity
- Conversation Analytics — Turn-taking stats, talk-time ratios, interruption detection, speaker overlap analysis
- Audio Clip Extraction — Extract segments to WAV files by time range, speaker, or search results
- Webhook Notifications — HTTP callbacks on indexing completion, keyword detection, event triggers
- Health Dashboard — Pipeline health monitoring, component status, memory/latency tracking
- Batch Operations — Batch segment insertion for faster indexing of large files
- Improved Robustness — Thread-safe SQLite store, batched search (prevents OOM on large datasets), input validation, error recovery
Features
- Transcription — Whisper, Faster Whisper, or bring your own ASR
- Audio Embeddings — CLAP joint text-audio space for semantic search
- Speaker Diarization — know who said what (pyannote)
- Speaker Registry — track speakers across files by voiceprint
- Audio Event Detection — detect alarms, machinery, glass breaks, and 50+ sound types
- Audio Preprocessing — noise reduction, normalization, bandpass filtering, compression
- Language Detection — automatic spoken language identification (99 languages)
- Live Streaming — real-time transcription with keyword and event callbacks
- Conversation Analytics — turn-taking, talk-time, interruptions, speaker overlap
- Clip Extraction — extract audio segments to WAV by time, speaker, or search
- Export — SRT, VTT, JSON, CSV, TXT, meeting notes
- Pluggable Architecture — every component swappable via base classes
- CLI — full command-line interface
Install
# Core (no ML dependencies)
pip install sonarwise
# With all features
pip install sonarwise[all]
# Pick what you need
pip install sonarwise[whisper] # Whisper transcription
pip install sonarwise[faster-whisper] # Faster Whisper (CTranslate2)
pip install sonarwise[clap] # CLAP audio embeddings
pip install sonarwise[diarization] # Speaker diarization
pip install sonarwise[speaker] # Speaker identification
pip install sonarwise[events] # Audio event detection (PANNs)
pip install sonarwise[preprocess] # Audio preprocessing (scipy)
pip install sonarwise[live] # Live microphone capture
pip install sonarwise[vad] # Silero VAD chunking
Requires: ffmpeg (sudo apt install ffmpeg or brew install ffmpeg)
Quick Start
from sonarwise import SonarWise
sw = SonarWise(diarization=True, events=True)
# Index audio
sw.index("meeting.wav")
sw.index_folder("./recordings/")
# Search by text
results = sw.query("budget discussion", top_k=5)
for r in results:
print(f"[{r.speaker_name}] {r.transcript} (score: {r.score})")
# Search by audio similarity
results = sw.query_audio("alarm_clip.wav", top_k=5)
# Search by speaker
results = sw.query("budget", speaker="Ant", top_k=5)
# Search by event
results = sw.query_events(event="machine_fault", top_k=5)
Audio Preprocessing (New in v0.2.0)
Clean and enhance audio before processing:
from sonarwise.core.preprocessor import AudioPreprocessor, NoiseReducer, Normalizer
from sonarwise import AudioData
# Full pipeline
preprocessor = AudioPreprocessor(
remove_dc=True, # Remove DC offset
bandpass=True, # Bandpass filter (80Hz-7500Hz)
noise_reduce=True, # Spectral gating noise reduction
normalize=True, # Peak normalization
compress=True, # Dynamic range compression
)
# Or use convenience wrappers
reducer = NoiseReducer(strength=1.5)
normalizer = Normalizer(target_peak=0.95)
Language Detection (New in v0.2.0)
Identify spoken language before or after transcription:
from sonarwise.core.language_detector import (
SimpleLanguageDetector,
WhisperLanguageDetector,
FasterWhisperLanguageDetector,
)
# Script-based detection (zero dependencies)
detector = SimpleLanguageDetector()
result = detector.detect_from_text("Bonjour le monde")
print(f"{result.language_name}: {result.confidence}") # French: 0.5
# Whisper-based detection (highly accurate)
detector = WhisperLanguageDetector(model_size="base")
result = detector.detect(audio_segment)
print(f"{result.language_name}: {result.confidence}") # French: 0.97
Conversation Analytics (New in v0.2.0)
# After indexing a meeting
analytics = sw.conversation_analytics("meeting.wav")
print(f"Total speakers: {analytics['total_speakers']}")
print(f"Total turns: {analytics['total_turns']}")
for speaker, stats in analytics['speaker_stats'].items():
print(f" {speaker}: {stats['talk_time_pct']:.1f}% talk time, "
f"{stats['turn_count']} turns")
Clip Extraction (New in v0.2.0)
# Extract a specific time range
sw.extract_clip("meeting.wav", start_ms=30000, end_ms=60000,
output="clip_30s_60s.wav")
# Extract all segments from a speaker
sw.extract_clip("meeting.wav", speaker="Ant",
output="ant_segments.wav")
Speaker Intelligence
# Register a speaker
sw.register_speaker("Ant", reference_audio="ant_voice.wav")
# Query by speaker
results = sw.query("budget", speaker="Ant")
# Speaker timeline
timeline = sw.speaker_timeline("meeting.wav")
# Speaker stats
stats = sw.speaker_stats("meeting.wav")
# Find speaker across files
presence = sw.find_speaker_across(speaker="Ant", folders=["./meetings/"])
Live Mode
sw = SonarWise(mode="live", diarization=True, events=True)
@sw.on("transcript")
def on_speech(segment):
print(f"[{segment.speaker_name}] {segment.transcript}")
@sw.on("keyword", words=["budget", "deadline", "risk"])
def on_keyword(segment):
send_alert(segment.transcript)
@sw.on("sound_event", events=["alarm", "glass_break"])
def on_danger(event):
trigger_alert(event)
sw.listen(source="microphone")
Plug Any Model
Every component is swappable:
from sonarwise import SonarWise
from sonarwise.core.transcriber import BaseTranscriber
class MyTranscriber(BaseTranscriber):
def transcribe(self, audio):
return my_model.process(audio)
sw = SonarWise(transcriber=MyTranscriber())
Pluggable slots:
| Component | Base Class | Default |
|---|---|---|
| Transcriber | BaseTranscriber |
Whisper |
| Audio Embedder | BaseAudioEmbedder |
CLAP |
| Vector Store | BaseVectorStore |
SQLite |
| Chunker | BaseChunker |
Silero VAD |
| Diarizer | BaseDiarizer |
pyannote |
| Speaker Embedder | BaseSpeakerEmbedder |
ECAPA-TDNN |
| Event Classifier | BaseEventClassifier |
PANNs / CLAP |
| Preprocessor | BasePreprocessor |
AudioPreprocessor |
| Language Detector | BaseLanguageDetector |
Whisper |
| Stream Listener | BaseStreamListener |
sounddevice |
CLI
sonarwise index meeting.wav
sonarwise query "budget discussion"
sonarwise speakers meeting.wav
sonarwise timeline meeting.wav
sonarwise listen --source microphone --on-keyword "budget,risk"
sonarwise export meeting.wav --format srt
sonarwise stats
Export
sw.export("meeting.wav", format="srt", output="subtitles.srt")
sw.export("meeting.wav", format="json", output="segments.json")
sw.export("meeting.wav", format="csv", output="data.csv")
sw.export("meeting.wav", format="txt", output="transcript.txt")
sw.export("meeting.wav", format="notes", output="meeting_notes.md")
Ecosystem
sonarwise is part of the Ant Intelligence Ecosystem:
| Library | Purpose | Tagline |
|---|---|---|
| SightRAG | Visual perception | See. Search. Retrieve. |
| sonarwise | Audio perception | Hear. Search. Retrieve. |
| adaptive-intelligence | Reasoning & memory | Learn. Remember. Adapt. |
| llmevalkit | Evaluation | Evaluate. Score. Improve. |
| docqwise | Document intelligence | Read. Query. Understand. |
| wavqwise | Audio Q&A | Ask. Listen. Answer. |
| AntGuard | AI safety & guardrails | Guard. Filter. Protect. |
License
Apache 2.0
Author
Built by Venkatkumar Rajan (VK-Ant)
- GitHub: https://github.com/VK-Ant
- Portfolio: https://vk-ant.github.io/Venkatkumar
Metadata
Release files for sonarwise 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| sonarwise-0.2.0.tar.gz | 66.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| sonarwise-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 127.9 kB
Release files / sonarwise-0.2.0.tar.gz
| Download URL | sonarwise-0.2.0.tar.gz |
|---|---|
| Size | 66.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f08577a91c1ee963bb5ab1bdad37bf1c1c11b39769436adb96ecf7525464c164
|
|
BLAKE2b-256 checksum How to use checksums |
8da91bbfc65f3b3801a45997092849950b9d11e39bd23b48ce423132d1986ab5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.0
|
Release files / sonarwise-0.2.0-py3-none-any.whl
| Download URL | sonarwise-0.2.0-py3-none-any.whl |
|---|---|
| Size | 61.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2d21f1065364a196674415391cfc85d446596f02136a448a4efcb77b5a00656b
|
|
BLAKE2b-256 checksum How to use checksums |
273e8e35e597999ae50634301f2a703c90efb5e111de9bf75e87df01de29a512
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.0
|