Voice Activity Detection (VAD) and Transcription package for audio/video processing
Project description
Status: Active
Last Updated: 2025-12-27
Version: Current
VAD Service Documentation
This document provides comprehensive documentation for the Voice Activity Detection (VAD) service, which runs as a serverless HTTP API on Modal.com.
Overview
The VAD service is a serverless HTTP API that analyzes audio/video files to detect speech segments. It uses the lansonai-vadtools Python package (published on PyPI) and runs on Modal.com for automatic scaling and isolation.
Key Features:
- Serverless architecture (Modal.com)
- Automatic scaling (up to 10 concurrent requests)
- Supports audio and video formats
- Exports detected speech segments
- Returns detailed analysis results
Architecture
The service consists of:
- Modal API:
modal_api.py- Serverless HTTP API endpoint - Python Package:
lansonai-vadtools- Core VAD processing logic (published on PyPI) - Integration: Main API calls Modal endpoint via HTTP
Quick Start
For API Users
The VAD service is already deployed and accessible via Modal. The main API automatically uses it when processing audio tasks.
Production Endpoint: https://deth--analyze.modal.run
For Developers
Prerequisites
- Python 3.12+
- Modal CLI installed globally:
pip install modal - Modal account (free tier available)
Local Development
-
Install Modal CLI (if not already installed):
pip install modal
-
Login to Modal:
modal token new
-
Run locally for testing:
cd scripts/python/vad modal serve modal_api.py
-
Test the local API:
curl -X POST http://localhost:8000/analyze \ -H "Content-Type: application/json" \ -d '{ "file_url": "https://example.com/audio.wav", "threshold": 0.3 }'
-
Deploy to production (after testing):
modal deploy modal_api.py
API Documentation
POST /analyze
Analyzes an audio/video file for voice activity.
Request Body (JSON):
{
"file_url": "https://example.com/audio.wav", // Required: Public URL to audio/video file
"threshold": 0.3, // Optional: VAD threshold (0.0-1.0), default 0.3
"min_segment_duration": 0.5, // Optional: Minimum segment duration (seconds), default 0.5
"max_merge_gap": 0.2, // Optional: Maximum merge gap (seconds), default 0.2
"export_segments": false, // Optional: Export audio segments, default false
"output_format": "wav", // Optional: Output format ("wav" or "flac"), default "wav"
"request_id": "custom-id" // Optional: Custom request ID
}
Response (Success):
{
"request_id": "abc123...",
"total_segments": 42,
"total_duration": 120.5,
"overall_speech_ratio": 0.85,
"segments": [
{
"id": 1,
"start_time": 0.0,
"end_time": 2.5,
"duration": 2.5,
"speech_confidence": 0.95
}
],
"summary": {
"total_duration": 120.5,
"total_speech_duration": 102.4,
"overall_speech_ratio": 0.85,
"num_segments": 42
},
"performance": {
"total_processing_time": 15.2,
"speed_ratio": 7.9
}
}
GET /health
Health check endpoint.
Response:
{
"status": "ok",
"service": "vad-api",
"version": "0.2.0"
}
Using the Python Package
The VAD service is built on top of the lansonai-vadtools Python package, which can also be used directly:
Installation
pip install lansonai-vadtools
Basic Usage
from lansonai.vadtools import analyze
result = analyze(
input_path="audio.wav",
output_dir="./output",
threshold=0.3,
min_segment_duration=0.5,
max_merge_gap=0.2,
export_segments=True,
output_format="wav"
)
print(f"Detected {result['total_segments']} speech segments")
print(f"Speech ratio: {result['overall_speech_ratio'] * 100:.1f}%")
Return Value Structure
{
"request_id": str,
"input_file": str,
"output_dir": str,
"json_path": str, # Path to timestamps.json
"segments_dir": str | None, # Path to segments directory (if exported)
"segments": List[Dict], # VAD segment list
"summary": Dict, # Statistics
"performance": Dict, # Performance metrics
"metadata": Dict, # Metadata
"total_segments": int,
"total_duration": float,
"overall_speech_ratio": float
}
For detailed package usage examples, see USAGE.md.
Deployment
Modal Configuration
The service uses Modal secrets for environment variables:
# Create secret with Supabase credentials (for segment uploads)
modal secret create vad-secrets \
SUPABASE_URL=https://your-project.supabase.co \
SUPABASE_ANON_KEY=your-anon-key
Resource Allocation
Current configuration (in modal_api.py):
- CPU: 2.0 cores
- Memory: 4096 MB (4 GB)
- Timeout: 300 seconds (5 minutes)
- Concurrency: 10 requests
Cost Estimation
Based on current configuration:
- Per request: ~$0.002-0.004 (1-2 minutes processing)
- Free tier: $30/month covers ~7,500-15,000 requests
For detailed deployment instructions, see MODAL_DEPLOY.md.
Testing
Test the Deployed Service
# Health check
curl https://deth--health.modal.run
# Analyze audio
curl -X POST https://deth--analyze.modal.run \
-H "Content-Type: application/json" \
-d '{
"file_url": "https://r2.deth.us/audio/example.mp3",
"threshold": 0.3
}'
Local Testing
# Start local server
cd scripts/python/vad
modal serve modal_api.py
# Test with local URL
curl -X POST http://localhost:8000/analyze \
-H "Content-Type: application/json" \
-d '{"file_url": "https://example.com/audio.wav"}'
For comprehensive testing guide, see TESTING.md.
Package Publishing
The lansonai-vadtools package is published on PyPI. To publish updates:
Prerequisites
- Get PyPI API token from https://pypi.org/manage/account/token/
- Set environment variable:
export UV_PUBLISH_TOKEN="pypi-your-token-here"
Publish Process
cd scripts/python/vad
# Build package
uv build
# Publish to PyPI
uv publish
For detailed publishing instructions, see README_PACKAGE.md.
Supported Formats
Input Formats
- Audio: WAV, MP3, M4A, FLAC, OGG
- Video: MP4, AVI, MOV, MKV, FLV, WMV, WEBM, M4V (requires ffmpeg)
Output Formats
- WAV
- FLAC
Environment Setup
For Python environment setup (pyenv, modal CLI, etc.), see SETUP_PYTHON_ENV.md.
Troubleshooting
Common Issues
- Download failures: Ensure file URL is publicly accessible
- Timeout errors: Increase timeout for large files or optimize parameters
- Format not supported: Check file format compatibility
- Modal deployment errors: Check Modal logs with
modal app logs vad-api
Viewing Logs
# View Modal app logs
modal app logs vad-api
Related Documentation
- MODAL_DEPLOY.md - Detailed deployment guide
- USAGE.md - Python package usage examples
- TESTING.md - Testing guide
- SETUP_PYTHON_ENV.md - Environment setup
- README_PACKAGE.md - Package publishing guide
- CHANGELOG.md - Package version history
Integration with Main API
The main API integrates with the VAD service via src/services/vadCliService.ts, which:
- Calls the Modal endpoint with audio URL
- Adapts the response format to match internal types
- Handles errors and retries
The service is automatically used when creating audio tasks via the main API.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file lansonai_vadtools-0.3.3.tar.gz.
File metadata
- Download URL: lansonai_vadtools-0.3.3.tar.gz
- Upload date:
- Size: 48.6 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.9.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b2e3f8d8767cb7b2141b8db82340539ba98518d5c3498fa818a61e1bc7cff597
|
|
| MD5 |
0e68f184348fc4d2b6e8cbcb1ab2703a
|
|
| BLAKE2b-256 |
bb301192f861ad04fe344abd785eb778f4fabab77db2b3dcd85a5da2ec974052
|
File details
Details for the file lansonai_vadtools-0.3.3-py3-none-any.whl.
File metadata
- Download URL: lansonai_vadtools-0.3.3-py3-none-any.whl
- Upload date:
- Size: 24.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.9.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2b2adba6b6b1e1dbd2e88a691acaaee9268706351c2d38f8964ec5ec48395763
|
|
| MD5 |
035d6c216cc5feb1efe7982e00b42773
|
|
| BLAKE2b-256 |
a3ac5e468a95e2547a2f06d9eeee4518108c74349a22f03639a29c9c57301b0f
|