Skip to main content

Modules required for conversation, released as a Python SDK.

Project description

respeak

Provides the modules required for conversation and is released as a Python SDK.

Install

pip install respeak-ai

# local development
pip install -e .

Optional extras:

pip install "respeak-ai[vad]"
pip install "respeak-ai[tts]"
pip install "respeak-ai[llm]"
pip install "respeak-ai[a2f]"
pip install "respeak-ai[dev]"

Supported models

Module Class Recommended weights Source
ASR StreamingParaformerAsr paraformer-zh-streaming (iic/speech_paraformer-large_asr_nat-zh-cn-16k-common-vocab8404-online) ModelScope · Hugging Face · FunASR
ASR (punc) used via punc_model= ct-punc (iic/punc_ct-transformer_cn-en-common-vocab471067-large) ModelScope · Hugging Face
TTS CosyVoice3Tts Fun-CosyVoice3-0.5B-2512 ModelScope · Hugging Face · CosyVoice
VAD SileroVad Silero VAD (auto-download via silero-vad) GitHub · PyPI
LLM Qwen3LLM Qwen3-8B (or other Qwen3 / Qwen2.5 Instruct checkpoints) ModelScope · Hugging Face · Qwen
Audio2Face NvidiaAudio2Face3D Audio2Face-3D-v2.3.1-Claire Hugging Face · Audio2Face-3D

Download examples:

# ASR + punctuation (ModelScope)
modelscope download --model iic/speech_paraformer-large_asr_nat-zh-cn-16k-common-vocab8404-online \
  --local_dir ./models/paraformer-zh-streaming
modelscope download --model iic/punc_ct-transformer_cn-en-common-vocab471067-large \
  --local_dir ./models/punc_ct-transformer_cn-en-common-vocab471067-large

# TTS (ModelScope or Hugging Face)
modelscope download --model FunAudioLLM/Fun-CosyVoice3-0.5B-2512 \
  --local_dir ./models/Fun-CosyVoice3-0.5B-2512

# LLM
modelscope download --model Qwen/Qwen3-8B --local_dir ./models/Qwen3-8B

# Audio2Face 3D (Hugging Face)
huggingface-cli download nvidia/Audio2Face-3D-v2.3.1-Claire \
  --local-dir ./models/audio2face-3d-v2.3.1-claire

Silero VAD needs no manual download: pip install "respeak-ai[vad]" then SileroVad.from_pretrained().

Usage

from respeak import StreamingParaformerAsr, CosyVoice3Tts, SileroVad, Qwen3LLM, NvidiaAudio2Face3D

# ASR
asr = StreamingParaformerAsr.from_pretrained(asr="path/or/model_id", punc_model="path/or/model_id")
text = asr.generate(input=audio_chunk, is_final=False, cache={})

# CosyVoice3 TTS (vLLM)
tts = CosyVoice3Tts.from_pretrained(
    model_dir="path/to/TTS_Model",
    prompt_text="...",
    prompt_wav="path/to/prompt.wav",
    load_vllm=True,
)
for wav in tts.generate("你好,欢迎使用本系统。", stream=True):
    ...

# VAD
vad = SileroVad.from_pretrained()
speech_prob = vad.generate(audio_chunk)

# LLM sentence-level streaming
llm = Qwen3LLM.from_pretrained(
    "path/to/LLM_Model_ft",
    base_prompt_file="path/to/base_prompt.md",
    persona_setting="你叫小睿,是一个温和的中文助手。",
    strategy_prompt="回答要简洁自然。",
)
for sentence in llm.generate("你好,介绍一下你自己。", stream=True):
    ...

# Audio2Face 3D (requires .[a2f])
a2f = NvidiaAudio2Face3D.from_pretrained("path/to/audio2face-3d-model", use_cuda=True)
weights = a2f.generate(audio_window)          # one 520ms window -> [51]
frames = a2f.generate(audio, stream=True, is_final=True)  # sliding window -> list[[51]]

New model types should live under src/respeak/models/<name>/ (subclass BaseModel, expose from_pretrained / generate). Add matching tests under tests/models/<name>/.

Examples

# ASR
python examples/paraformer_asr_streaming.py \
  --asr path/to/ASR_Model \
  --punc path/to/PUNC_Model \
  --audio path/to/test.wav

# TTS (requires .[tts])
python examples/cosyvoice3_tts_streaming.py \
  --model-dir path/to/Fun-CosyVoice3 \
  --prompt-wav path/to/prompt.wav \
  --prompt-text "..." \
  --text "你好,欢迎使用本系统。" \
  --output output.wav

# VAD (requires .[vad])
python examples/silerovad_real_inference.py \
  --audio path/to/test.wav \
  --threshold 0.5

# LLM (requires .[llm])
python examples/qwen3_llm_streaming.py \
  --model-path path/to/LLM_Model_ft \
  --base-prompt path/to/base_prompt.md \
  --text "你好,介绍一下你自己。"

# Audio2Face 3D (requires .[a2f])
python examples/nvidia_audio2face_3d_streaming.py \
  --model-dir path/to/audio2face-3d-v2.3.1-claire \
  --audio path/to/test.wav \
  --output arkit_weights.npy

Tests

pip install -e ".[dev]"
pytest tests/models -q

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

respeak_ai-0.1.6.tar.gz (241.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

respeak_ai-0.1.6-py3-none-any.whl (261.6 kB view details)

Uploaded Python 3

File details

Details for the file respeak_ai-0.1.6.tar.gz.

File metadata

  • Download URL: respeak_ai-0.1.6.tar.gz
  • Upload date:
  • Size: 241.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for respeak_ai-0.1.6.tar.gz
Algorithm Hash digest
SHA256 6ed8777f1e249ea9ee956400e4b363ba4abe855eb5955b9a471b9087dbaec1bb
MD5 435cd1ecc8266f7e193d40e766aa7712
BLAKE2b-256 9a0dba1237e1530014e7953dae954bc6746f906176ab408f781824d708e97dc6

See more details on using hashes here.

Provenance

The following attestation bundles were made for respeak_ai-0.1.6.tar.gz:

Publisher: publish-pypi.yml on shanguanma/respeak

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file respeak_ai-0.1.6-py3-none-any.whl.

File metadata

  • Download URL: respeak_ai-0.1.6-py3-none-any.whl
  • Upload date:
  • Size: 261.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for respeak_ai-0.1.6-py3-none-any.whl
Algorithm Hash digest
SHA256 29b4878c0167181629c96c7d2786cc94727568487d0061bf5ef68d052bee0da3
MD5 26764bf94f3017cd13a77d6739d4474e
BLAKE2b-256 32f3dc8a4000c07c592d15f3e215ac1f033fe8bcaaf846814562468b04c95ac5

See more details on using hashes here.

Provenance

The following attestation bundles were made for respeak_ai-0.1.6-py3-none-any.whl:

Publisher: publish-pypi.yml on shanguanma/respeak

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page