Skip to main content

A unified, extensible, and modern Python toolkit for LLM-based Automatic Speech Recognition (ASR).

Project description

Modern ASR

A unified, extensible, and future-proof Python toolkit for locally running state-of-the-art LLM-based Automatic Speech Recognition (ASR) models.

Python License


✨ Features

  • 🧩 23 Models — Whisper, SenseVoice, Qwen, MiMo, FireRedASR, GLM-ASR, and more
  • 🔌 Plugin Architecture — Add new models with @register_model decorator
  • 🚀 Hot-Swap — Switch models at runtime without restarting
  • 🌍 Multi-Language — 52 languages, 22 Chinese dialects
  • 🎯 Multi-Task — Transcription, translation, diarization, emotion, events
  • 💻 Local-First — All inference on-device. No APIs. No data leaves your machine.
  • 🍎 Apple Silicon — MPS (Metal Performance Shaders) support on macOS
  • 🐍 Modern Python — uv-native packaging, Pydantic configs, rich CLI

📦 Installation

From PyPI (Recommended)

pip install modern-asr

模型依赖和权重会在第一次使用时自动安装 — 你只需要输入模型名字,其余全自动化:

from modern_asr import ASRPipeline

# SenseVoice — 自动安装 funasr、modelscope,自动下载权重
pipe = ASRPipeline("sensevoice-small")

# MiMo-ASR — 自动 clone 官方仓库,自动下载 HF 权重
pipe = ASRPipeline("mimo-asr-v2.5")

# Whisper — 自动安装 openai-whisper,自动下载权重
pipe = ASRPipeline("whisper-small")

如果需要预装所有依赖(离线环境):

pip install modern-asr[all-models]

Available extras: transformers, vllm, onnx, firered-asr, sensevoice, fun-asr, qwen-asr, mimo-asr, canary-qwen, glm-asr, whisper, moonshine, all-models, all-backends, all.

From Source

# Clone the repository
git clone https://github.com/vra/modern-asr.git
cd modern-asr

# Sync dependencies (recommended)
uv sync --all-extras

# Or install specific extras only
uv sync --extra transformers --extra whisper

# Or just core dependencies
uv sync

Python 3.10+ recommended. Some models (Qwen3-ASR, MiMo) require Python ≥ 3.10.


🚀 Quick Start

from modern_asr import ASRPipeline

# Transcribe with SenseVoice (Alibaba)
pipe = ASRPipeline("sensevoice-small")
result = pipe("audio.wav", language="zh")
print(result.text)

# Switch to Qwen3-ASR for dialect support
pipe.switch_model("qwen3-asr-0.6b")
result = pipe("audio.wav", language="zh")
print(result.text)

# English with Whisper
pipe.switch_model("whisper-small")
result = pipe("audio.wav", language="en")

📚 Documentation

Full documentation with Material for MkDocs:

mkdocs serve

🏗️ Architecture

Modern ASR is built on three layers:

  1. ASRPipeline — Unified user API. Handles input normalization, task dispatch, model lifecycle.
  2. ASRModel / AudioLLMModel — Adapter layer. New models often need only 8 lines of config via AudioLLMModel.
  3. Backends — Transformers, vLLM, ONNX Runtime.

Adding a New Model

from modern_asr.core.audio_llm import AudioLLMModel
from modern_asr.core.registry import register_model

@register_model("my-model-1b")
class MyModel1B(AudioLLMModel):
    HF_PATH = "org/MyModel-1B"
    SUPPORTED_LANGUAGES = {"zh", "en"}
    CHUNK_DURATION = 30.0

    @property
    def model_id(self) -> str:
        return "my-model-1b"

That's it. The registry auto-discovers it at runtime.


🤝 Contributing

See Contributing Guide for development setup, code style, and PR checklist.


📄 License

Apache-2.0

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

modern_asr-0.2.2.tar.gz (374.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

modern_asr-0.2.2-py3-none-any.whl (54.0 kB view details)

Uploaded Python 3

File details

Details for the file modern_asr-0.2.2.tar.gz.

File metadata

  • Download URL: modern_asr-0.2.2.tar.gz
  • Upload date:
  • Size: 374.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.8 {"installer":{"name":"uv","version":"0.11.8","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for modern_asr-0.2.2.tar.gz
Algorithm Hash digest
SHA256 38c83c268832b3949e3506e709b93a84445bc955e97c12f7d3d1b9c89c3fd863
MD5 c0d308b540c4313a6613ec2f990ceb3e
BLAKE2b-256 9205e4bd7af2fa2a4069356a08f0140f8e270d9231ef1118dbbb021e2866c180

See more details on using hashes here.

File details

Details for the file modern_asr-0.2.2-py3-none-any.whl.

File metadata

  • Download URL: modern_asr-0.2.2-py3-none-any.whl
  • Upload date:
  • Size: 54.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.8 {"installer":{"name":"uv","version":"0.11.8","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for modern_asr-0.2.2-py3-none-any.whl
Algorithm Hash digest
SHA256 78dfe8c1e13ceb50ab200a52f4f5d5947667ba0781faf0a55f3e91c53a5d30ea
MD5 a84dfc037f1f775fc1ab385ae1453679
BLAKE2b-256 59379163326c2ebdc741bee2194cc348e3a432acde228be2bef9bf9ab22776f1

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page