Skip to main content

Speechall Python SDK

Python SDK for the Speechall API - A powerful speech-to-text transcription service supporting multiple AI models and providers.

PyPI version Python 3.8+

Features

  • Multiple AI Models: Access various speech-to-text models from different providers (OpenAI Whisper, and more)
  • Flexible Input: Transcribe local audio files or remote URLs
  • Rich Output Formats: Get results in text, JSON, SRT, or VTT formats
  • Speaker Diarization: Identify and separate different speakers in audio
  • Custom Vocabulary: Improve accuracy with domain-specific terms
  • Replacement Rules: Apply custom text transformations to transcriptions
  • Language Support: Auto-detect languages or specify from a wide range of supported languages
  • Async Support: Built with async/await support using httpx

Installation

pip install speechall

Quick Start

Basic Transcription

import os
from speechall import SpeechallApi

# Initialize the client
client = SpeechallApi(token=os.getenv("SPEECHALL_API_TOKEN"))

# Transcribe a local audio file
with open("audio.mp3", "rb") as audio_file:
    audio_data = audio_file.read()

response = client.speech_to_text.transcribe(
    model="openai.whisper-1",
    request=audio_data,
    language="en",
    output_format="json",
    punctuation=True
)

print(response.text)

Transcribe Remote Audio

from speechall import SpeechallApi

client = SpeechallApi(token=os.getenv("SPEECHALL_API_TOKEN"))

response = client.speech_to_text.transcribe_remote(
    file_url="https://example.com/audio.mp3",
    model="openai.whisper-1",
    language="auto",  # Auto-detect language
    output_format="json"
)

print(response.text)

Advanced Features

Speaker Diarization

Identify different speakers in your audio:

response = client.speech_to_text.transcribe(
    model="openai.whisper-1",
    request=audio_data,
    language="en",
    output_format="json",
    diarization=True,
    speakers_expected=2
)

for segment in response.segments:
    print(f"[Speaker {segment.speaker}] {segment.text}")

Custom Vocabulary

Improve accuracy for specific terms:

response = client.speech_to_text.transcribe(
    model="openai.whisper-1",
    request=audio_data,
    language="en",
    output_format="json",
    custom_vocabulary=["Kubernetes", "API", "Docker", "microservices"]
)

Replacement Rules

Apply custom text transformations:

from speechall import ReplacementRule, ExactRule

replacement_rules = [
    ReplacementRule(
        rule=ExactRule(find="API", replace="Application Programming Interface")
    )
]

response = client.speech_to_text.transcribe_remote(
    file_url="https://example.com/audio.mp3",
    model="openai.whisper-1",
    language="en",
    output_format="json",
    replacement_ruleset=replacement_rules
)

List Available Models

models = client.speech_to_text.list_speech_to_text_models()

for model in models:
    print(f"{model.model_identifier}: {model.display_name}")
    print(f"  Provider: {model.provider}")

Configuration

Authentication

Get your API token from speechall.com and set it as an environment variable:

export SPEECHALL_API_TOKEN="your-token-here"

Or pass it directly when initializing the client:

from speechall import SpeechallApi

client = SpeechallApi(token="your-token-here")

Output Formats

  • text: Plain text transcription
  • json: JSON with detailed information (segments, timestamps, metadata)
  • json_text: JSON with simplified text output
  • srt: SubRip subtitle format
  • vtt: WebVTT subtitle format

Language Codes

Use ISO 639-1 language codes (e.g., en, es, fr, de) or auto for automatic detection.

API Reference

Client Classes

  • SpeechallApi: Main client for the Speechall API
  • AsyncSpeechallApi: Async client for the Speechall API

Main Methods

speech_to_text.transcribe()

Transcribe a local audio file.

Parameters:

  • model (str): Model identifier (e.g., "openai.whisper-1")
  • request (bytes): Audio file content
  • language (str): Language code or "auto"
  • output_format (str): Output format (text, json, srt, vtt)
  • punctuation (bool): Enable automatic punctuation
  • diarization (bool): Enable speaker identification
  • speakers_expected (int, optional): Expected number of speakers
  • custom_vocabulary (list, optional): List of custom terms
  • initial_prompt (str, optional): Context prompt for the model
  • temperature (float, optional): Model temperature (0.0-1.0)

speech_to_text.transcribe_remote()

Transcribe audio from a URL.

Parameters: Same as transcribe() but with file_url instead of request

speech_to_text.list_speech_to_text_models()

List all available models.

Examples

Check out the examples directory for more detailed usage examples:

Requirements

  • Python 3.8+
  • httpx >= 0.27.0
  • pydantic >= 2.0.0
  • typing-extensions >= 4.0.0

Development

# Install with dev dependencies
pip install -e ".[dev]"

# Run tests
pytest

# Type checking
mypy .

Support

License

MIT License - see LICENSE file for details

Release files for speechall 0.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for speechall 0.6.0
File Size Uploaded
speechall-0.6.0.tar.gz 32.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for speechall 0.6.0
File Interpreter ABI Platform
speechall-0.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 89.1 kB

Release files / speechall-0.6.0.tar.gz

Download URL speechall-0.6.0.tar.gz
Size 32.3 kB
Tags Source
SHA-256 checksum
How to use checksums
8dd03c84b75915e77bcdf02549d080be5cf3c14bbb53d4c570474e68fa00f645
BLAKE2b-256 checksum
How to use checksums
345e2fc3846262409f493f257c78eddaacf1fcc01518c82f38cd4b36b73a2a61
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / speechall-0.6.0-py3-none-any.whl

Download URL speechall-0.6.0-py3-none-any.whl
Size 56.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c84cf6664e6847d963a3d76bbe54e9b700390adf834c626fa7aabdfae07b0650
BLAKE2b-256 checksum
How to use checksums
186413f24be6728b22f795b5506770b73d1ef8febb7a77b3a790c0558cfe4b8f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

0.6.0 This release

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page