Skip to main content

Multimodal Voice API Proxy

A privacy-first API proxy that intelligently routes voice and vision requests to local models with automatic cloud fallback.

What is this?

This proxy lets you run AI transcription (Whisper) and vision analysis (LLaVA) locally for privacy and cost savings, while automatically falling back to cloud APIs (OpenAI, Anthropic) when local resources are unavailable. It includes smart caching, usage metering, cost tracking, and multi-tenant API key management—perfect for developers building voice-enabled applications who want local-first privacy with cloud reliability as a safety net.

Features

  • Intelligent routing: Local-first processing with automatic cloud fallback
  • Privacy-focused: Audio and images processed on your hardware by default
  • Smart caching: Redis-backed caching prevents reprocessing identical inputs
  • Cost tracking: Real-time dashboard showing savings vs. cloud-only approach
  • Multi-tenant: API key management with per-key usage limits and metrics
  • OpenAPI-compatible: Drop-in replacement for OpenAI/Anthropic endpoints
  • Production-ready: Rate limiting, monitoring middleware, and async processing
  • Easy deployment: Docker Compose for local, one-click configs for Railway/Fly.io

Quick Start

Prerequisites

  • Docker and Docker Compose
  • 8GB+ RAM (for local models)
  • GPU recommended but optional

Installation

  1. Clone and configure
git clone <repository-url>
cd multimodal-voice-api-proxy
cp .env.example .env
  1. Edit .env with your settings
# Required for cloud fallback
OPENAI_API_KEY=sk-...
ANTHROPIC_API_KEY=sk-ant-...

# Database
DATABASE_URL=postgresql://user:pass@db:5432/proxy

# Redis cache
REDIS_URL=redis://redis:6379/0
  1. Launch with Docker Compose
docker-compose up -d
  1. Run migrations
docker-compose exec api alembic upgrade head
  1. Create your first API key
curl -X POST http://localhost:8000/keys \
  -H "Content-Type: application/json" \
  -d '{"name": "My App", "rate_limit": 100}'

The API will be available at http://localhost:8000

Usage

Audio Transcription

curl -X POST http://localhost:8000/v1/audio/transcriptions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -F file=@audio.mp3 \
  -F model=whisper-1

Vision Analysis

curl -X POST http://localhost:8000/v1/vision/analyze \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "image_url": "https://example.com/image.jpg",
    "prompt": "Describe this image"
  }'

Check Usage Stats

curl http://localhost:8000/keys/YOUR_API_KEY/stats \
  -H "Authorization: Bearer YOUR_API_KEY"

Response includes:

  • Total requests (local vs. cloud)
  • Cost savings
  • Cache hit rate
  • Rate limit status

Deployment

Railway

railway up

Fly.io

fly deploy

Self-Hosted

Use the included Dockerfile and docker-compose.yml for custom deployments.

Tech Stack

  • Framework: FastAPI (Python 3.11+)
  • Local Models: faster-whisper, llama-cpp-python
  • Cloud APIs: OpenAI, Anthropic
  • Database: PostgreSQL + SQLAlchemy + Alembic
  • Cache: Redis
  • Deployment: Docker, Railway, Fly.io

Configuration

Key environment variables:

Variable Description Default
LOCAL_MODELS_ENABLED Enable local model processing true
MAX_LOCAL_REQUESTS Concurrent local requests before fallback 5
CACHE_TTL Cache expiration in seconds 3600
RATE_LIMIT_WINDOW Rate limit window in seconds 60

See .env.example for complete configuration options.

License

MIT License - see LICENSE file for details.


Built for developers who value privacy without sacrificing reliability.

Release files for multimodal-voice-api-proxy 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for multimodal-voice-api-proxy 0.1.0
File Size Uploaded
multimodal_voice_api_proxy-0.1.0.tar.gz 15.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for multimodal-voice-api-proxy 0.1.0
File Interpreter ABI Platform
multimodal_voice_api_proxy-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 33.8 kB

Release files / multimodal_voice_api_proxy-0.1.0.tar.gz

Download URL multimodal_voice_api_proxy-0.1.0.tar.gz
Size 15.7 kB
Tags Source
SHA-256 checksum
How to use checksums
c885f31d35c493c1921ff6bb13cd5addeac4179b44e5ad714e98b27523f3b144
BLAKE2b-256 checksum
How to use checksums
e53c98bc903b8558759d40a1163636ea9353fbdd2fc9dfbf8eea51dbdf59473f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.9.25

Release files / multimodal_voice_api_proxy-0.1.0-py3-none-any.whl

Download URL multimodal_voice_api_proxy-0.1.0-py3-none-any.whl
Size 18.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9e41e1382487ae3be7ad72da5b5dc7dee1be35266fe202e0db0f318cfb72a2c7
BLAKE2b-256 checksum
How to use checksums
e612e000a1c281e39e27876d6064dc6634a160ce24870091d384618e7c223b26
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.9.25

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page