Skip to main content

Python SDK for RedenLab ML inference service

Project description

RedenLab Extract SDK

Python SDK for RedenLab's ML inference service. Run machine learning models on audio files using AWS SageMaker async inference endpoints.

Features

  • Simple, Pythonic API for ML inference
  • Automatic file upload to S3
  • Async job management with polling
  • Built-in retry logic and error handling
  • Support for multiple ML models (intelligibility, speaker diarization, etc.)

Installation

pip install redenlab-extract

Quick Start

from redenlab_ml import InferenceClient

# Initialize client with API key
client = InferenceClient(api_key="sk_live_your_api_key_here")

# Run inference on an audio file
result = client.predict(file_path="audio.wav")

print(result)
# {
#   'job_id': 'uuid',
#   'status': 'completed',
#   'result': {'intelligibility_score': 0.85},
#   'created_at': '2025-01-15T10:30:00Z',
#   'completed_at': '2025-01-15T10:35:00Z'
# }

Authentication

The SDK looks for your API key in the following order:

  1. Constructor parameter: InferenceClient(api_key="sk_live_...")
  2. Environment variable: REDENLAB_ML_API_KEY
  3. Config file: ~/.redenlab-ml/config.yaml

Using Environment Variables

export REDENLAB_ML_API_KEY=sk_live_your_api_key_here
from redenlab_ml import InferenceClient

# API key is loaded from environment
client = InferenceClient()
result = client.predict(file_path="audio.wav")

Using Config File

Create ~/.redenlab-ml/config.yaml:

api_key: sk_live_your_api_key_here
base_url: https://your-api-gateway-url.amazonaws.com/prod  # optional
model_name: intelligibility  # optional
from redenlab_ml import InferenceClient

# API key is loaded from config file
client = InferenceClient()
result = client.predict(file_path="audio.wav")

Advanced Usage

Specify Model

client = InferenceClient(
    api_key="sk_live_...",
    model_name="speaker_diarisation_workflow"
)

Available models:

  • intelligibility (default)
  • speaker_diarisation_workflow
  • ataxia-naturalness
  • ataxia-intelligibility

Custom Timeout

client = InferenceClient(
    api_key="sk_live_...",
    timeout=7200  # 2 hours
)

Batch Processing (Submit + Poll Pattern)

For processing multiple files efficiently, use submit() and poll() separately:

# Submit all jobs first (fast - no waiting)
job_ids = []
for audio_file in audio_files:
    job_id = client.submit(file_path=audio_file)
    job_ids.append(job_id)
    print(f"Submitted: {job_id}")

# Poll for results (efficient - single loop checking all jobs)
results = {}
pending = set(job_ids)

while pending:
    for job_id in list(pending):
        status = client.get_status(job_id)

        if status['status'] == 'completed':
            results[job_id] = status['result']
            pending.remove(job_id)
        elif status['status'] == 'failed':
            results[job_id] = {'error': status.get('error')}
            pending.remove(job_id)

    if pending:
        time.sleep(10)  # Check all jobs every 10 seconds

print(f"Processed {len(results)} files!")

Submit and Poll Separately

# Submit job and get job_id immediately
job_id = client.submit(file_path="audio.wav")
print(f"Job submitted: {job_id}")

# ... do other work, or even exit and resume later ...

# Poll when ready
result = client.poll(job_id)
print(result['result'])

Check Job Status (Non-blocking)

# Get status of a running job without blocking
status = client.get_status(job_id="existing-job-uuid")
print(status['status'])  # 'upload_pending', 'processing', 'completed', 'failed'

Error Handling

from redenlab_ml import InferenceClient, InferenceError, TimeoutError, AuthenticationError

try:
    client = InferenceClient(api_key="sk_live_...")
    result = client.predict(file_path="audio.wav")
except AuthenticationError as e:
    print(f"Authentication failed: {e}")
except TimeoutError as e:
    print(f"Inference timed out: {e}")
except InferenceError as e:
    print(f"Inference failed: {e}")

Configuration Options

Parameter Environment Variable Config File Default
api_key REDENLAB_ML_API_KEY api_key Required
base_url REDENLAB_ML_BASE_URL base_url Production endpoint
model_name REDENLAB_ML_MODEL model_name intelligibility
timeout REDENLAB_ML_TIMEOUT timeout 3600 (1 hour)

Supported File Types

  • WAV (.wav)
  • MP3 (.mp3)
  • M4A (.m4a)
  • FLAC (.flac)

Requirements

  • Python 3.8+
  • requests
  • tenacity
  • pyyaml

Development

Install Development Dependencies

# Clone from CodeCommit (requires AWS credentials)
git clone ssh://git-codecommit.us-west-2.amazonaws.com/v1/repos/redenlab-extract
cd redenlab-extract
pip install -e ".[dev]"

Run Tests

pytest tests/

Code Formatting

black src/ tests/
ruff check src/ tests/

Type Checking

mypy src/

Examples

See the examples/ directory for more usage examples:

Documentation

Support

License

MIT License - see LICENSE file for details

Changelog

See CHANGELOG.md for version history and release notes.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

redenlab_extract-0.2.0.tar.gz (45.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

redenlab_extract-0.2.0-py3-none-any.whl (20.7 kB view details)

Uploaded Python 3

File details

Details for the file redenlab_extract-0.2.0.tar.gz.

File metadata

  • Download URL: redenlab_extract-0.2.0.tar.gz
  • Upload date:
  • Size: 45.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.25

File hashes

Hashes for redenlab_extract-0.2.0.tar.gz
Algorithm Hash digest
SHA256 445c5c95d910ddb392c1c17faf566901f0ada54c1227942a08b550b3671baf47
MD5 124801928b42ee3de579b2d221491e0a
BLAKE2b-256 e9bb508db9800a0e90cb38786651307a87da97aae3363977c174cc69a0870160

See more details on using hashes here.

File details

Details for the file redenlab_extract-0.2.0-py3-none-any.whl.

File metadata

File hashes

Hashes for redenlab_extract-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 9b3013b2d2751911d32fad032bb899fdba4bb3ff594ec9d43d1ee2d89787206b
MD5 d4f62fa0cd0dad1f97186ffece862a38
BLAKE2b-256 7a66c59c0e7724075eba41b2298d1b7135deb5d699f7e668c7020f556b5e5b58

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page