Skip to main content

DataoceanAI Open-source Large Speech Model

Project description

Dolphin

Paper Github Huggingface Modelscope Openi Wisemodel

Dolphin is a multilingual, multitask ASR model developed through a collaboration between Dataocean AI and Tsinghua University. It supports 40 Eastern languages across East Asia, South Asia, Southeast Asia, and the Middle East, while also supporting 22 Chinese dialects. It is trained on over 210,000 hours of data, which includes both DataoceanAI's proprietary datasets and open-source datasets. The model can perform speech recognition, voice activity detection (VAD), segmentation, and language identification (LID).

Approach

Mulitask data format Dolphin largely follows the innovative design approach of Whisper and OWSM. A joint CTC-Attention architecture is adopted, with encoder based on E-Branchformer and decoder based on standard Transformer. Several key modifications are introduced for its specific focus on ASR. Dolphin does not support translation tasks, and eliminates the use of previous text and its related tokens.

A significant enhancement in Dolphin is the introduction of a two-level language token system to better handle linguistic and regional diversity, especially in Dataocean AI dataset. The first token specifies the language (e.g., <zh>, <ja>), while the second token indicates the region (e.g., <CN>, <JP>). See details in paper.

Setup

Dolphin requires FFmpeg to convert audio file to WAV format. If FFmpeg is not installed on your system, please install it first:

# Ubuntu or Debian
sudo apt update && sudo apt install ffmpeg

# MacOS
brew install ffmpeg

# Windows
choco install ffmpeg

You can install the latest version of Dolphin using the following command:

pip install -U dataoceanai-dolphin

Alternatively, it can also be installed from the source:

pip install git+https://github.com/SpeechOceanTech/Dolphin.git 

Available Models and Languages

Models

There are 4 models in Dolphin, and 2 of them are available now. See details in paper.

Model Parameters Average WER Publicly Available
base 140 M 33.3
small 372 M 25.2
medium 910 M 23.1
large 1679 M 21.6

Languages

Dolphin supports 40 Eastern languages and 22 Chinese dialects. For a complete list of supported languages, see languages.md.

Usage

Command-line usage

dolphin audio.wav

# Download model and specify the model path
dolphin audio.wav --model small --model_dir /data/models/dolphin/

# Specify language and region
dolphin audio.wav --model small --model_dir /data/models/dolphin/ --lang_sym "zh" --region_sym "CN"

# padding speech to 30 seconds
dolphin audio.wav --model small --model_dir /data/models/dolphin/ --lang_sym "zh" --region_sym "CN" --padding_speech true

Python usage

import dolphin

waveform = dolphin.load_audio("audio.wav")
model = dolphin.load_model("small", "/data/models/dolphin", "cuda")
result = model(waveform)

# Specify language
result = model(waveform, lang_sym="zh")

# Specify language and region
result = model(waveform, lang_sym="zh", region_sym="CN")
print(result.text)

License

Dolphin's code and model weights are released under the Apache 2.0 License.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

dataoceanai-dolphin-20250716.tar.gz (608.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

dataoceanai_dolphin-20250716-py3-none-any.whl (614.8 kB view details)

Uploaded Python 3

File details

Details for the file dataoceanai-dolphin-20250716.tar.gz.

File metadata

  • Download URL: dataoceanai-dolphin-20250716.tar.gz
  • Upload date:
  • Size: 608.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.10.14

File hashes

Hashes for dataoceanai-dolphin-20250716.tar.gz
Algorithm Hash digest
SHA256 bfd3c995f13392984683d6bffeb6e77e395786c8a0364b93bc6c985b6487574c
MD5 0f0f37e2b56fd903f7373f0c31fa05e7
BLAKE2b-256 d0f36d32704c59bc6423dad49f3b57f57c2c2fff6d692e1f6d58b44ed9ddcdde

See more details on using hashes here.

File details

Details for the file dataoceanai_dolphin-20250716-py3-none-any.whl.

File metadata

File hashes

Hashes for dataoceanai_dolphin-20250716-py3-none-any.whl
Algorithm Hash digest
SHA256 7a058ddd7d0c3aebcf6348787b004eae4a49cb8d3b4835df4716747ea4f65633
MD5 5346f223b1ed488aa8eded69446ff1e6
BLAKE2b-256 b9011a5a05d42c9e02e756323313d4abba917ad10e464f50aa86c7119e8a6a6d

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page