Skip to main content

Kaldi alignment methods wrapped into Python

Project description

kaldialign

A small package that exposes edit distance computation functions from Kaldi. It uses the original Kaldi code and wraps it using pybind11.

Installation

conda install -c kaldialign kaldialign

or

pip install --verbose kaldialign

or

pip install --verbose -U git+https://github.com/pzelasko/kaldialign.git

or

git clone https://github.com/pzelasko/kaldialign.git
cd kaldialign
python3 -m pip install --verbose .

Examples

  • align(seq1, seq2, epsilon) - used to obtain the alignment between two string sequences. epsilon should be a null symbol (indicating deletion/insertion) that doesn't exist in either sequence.
from kaldialign import align

EPS = '*'
a = ['a', 'b', 'c']
b = ['a', 's', 'x', 'c']
ali = align(a, b, EPS)
assert ali == [('a', 'a'), ('b', 's'), (EPS, 'x'), ('c', 'c')]
  • edit_distance(seq1, seq2) - used to obtain the total edit distance, as well as the number of insertions, deletions and substitutions.
from kaldialign import edit_distance

a = ['a', 'b', 'c']
b = ['a', 's', 'x', 'c']
results = edit_distance(a, b)
assert results == {
    'ins': 1,
    'del': 0,
    'sub': 1,
    'total': 2
}
  • For both of the above examples, you can pass sclite_mode=True to compute WER or alignments based on SCLITE style weights, i.e., insertion/deletion cost 3 and substitution cost 4.

Motivation

The need for this arised from the fact that practically all implementations of the Levenshtein distance have slight differences, making it impossible to use a different scoring tool than Kaldi and get the same error rate results. This package copies code from Kaldi directly and wraps it using Cython, avoiding the issue altogether.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

kaldialign-0.7.2.tar.gz (23.4 kB view hashes)

Uploaded Source

Built Distributions

kaldialign-0.7.2-cp311-cp311-win_amd64.whl (62.7 kB view hashes)

Uploaded CPython 3.11 Windows x86-64

kaldialign-0.7.2-cp311-cp311-win32.whl (55.3 kB view hashes)

Uploaded CPython 3.11 Windows x86

kaldialign-0.7.2-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (82.6 kB view hashes)

Uploaded CPython 3.11 manylinux: glibc 2.17+ x86-64

kaldialign-0.7.2-cp311-cp311-manylinux_2_17_i686.manylinux2014_i686.whl (88.6 kB view hashes)

Uploaded CPython 3.11 manylinux: glibc 2.17+ i686

kaldialign-0.7.2-cp311-cp311-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (78.3 kB view hashes)

Uploaded CPython 3.11 manylinux: glibc 2.17+ ARM64

kaldialign-0.7.2-cp311-cp311-macosx_10_9_universal2.whl (100.0 kB view hashes)

Uploaded CPython 3.11 macOS 10.9+ universal2 (ARM64, x86-64)

kaldialign-0.7.2-cp310-cp310-win_amd64.whl (62.9 kB view hashes)

Uploaded CPython 3.10 Windows x86-64

kaldialign-0.7.2-cp310-cp310-win32.whl (55.1 kB view hashes)

Uploaded CPython 3.10 Windows x86

kaldialign-0.7.2-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (82.5 kB view hashes)

Uploaded CPython 3.10 manylinux: glibc 2.17+ x86-64

kaldialign-0.7.2-cp310-cp310-manylinux_2_17_i686.manylinux2014_i686.whl (88.6 kB view hashes)

Uploaded CPython 3.10 manylinux: glibc 2.17+ i686

kaldialign-0.7.2-cp310-cp310-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (78.3 kB view hashes)

Uploaded CPython 3.10 manylinux: glibc 2.17+ ARM64

kaldialign-0.7.2-cp310-cp310-macosx_10_9_universal2.whl (100.0 kB view hashes)

Uploaded CPython 3.10 macOS 10.9+ universal2 (ARM64, x86-64)

kaldialign-0.7.2-cp39-cp39-win_amd64.whl (62.9 kB view hashes)

Uploaded CPython 3.9 Windows x86-64

kaldialign-0.7.2-cp39-cp39-win32.whl (55.4 kB view hashes)

Uploaded CPython 3.9 Windows x86

kaldialign-0.7.2-cp39-cp39-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (82.7 kB view hashes)

Uploaded CPython 3.9 manylinux: glibc 2.17+ x86-64

kaldialign-0.7.2-cp39-cp39-manylinux_2_17_i686.manylinux2014_i686.whl (88.5 kB view hashes)

Uploaded CPython 3.9 manylinux: glibc 2.17+ i686

kaldialign-0.7.2-cp39-cp39-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (78.4 kB view hashes)

Uploaded CPython 3.9 manylinux: glibc 2.17+ ARM64

kaldialign-0.7.2-cp39-cp39-macosx_10_9_universal2.whl (100.2 kB view hashes)

Uploaded CPython 3.9 macOS 10.9+ universal2 (ARM64, x86-64)

kaldialign-0.7.2-cp38-cp38-win_amd64.whl (62.6 kB view hashes)

Uploaded CPython 3.8 Windows x86-64

kaldialign-0.7.2-cp38-cp38-win32.whl (55.2 kB view hashes)

Uploaded CPython 3.8 Windows x86

kaldialign-0.7.2-cp38-cp38-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (82.5 kB view hashes)

Uploaded CPython 3.8 manylinux: glibc 2.17+ x86-64

kaldialign-0.7.2-cp38-cp38-manylinux_2_17_i686.manylinux2014_i686.whl (88.5 kB view hashes)

Uploaded CPython 3.8 manylinux: glibc 2.17+ i686

kaldialign-0.7.2-cp38-cp38-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (78.3 kB view hashes)

Uploaded CPython 3.8 manylinux: glibc 2.17+ ARM64

kaldialign-0.7.2-cp38-cp38-macosx_10_9_universal2.whl (99.9 kB view hashes)

Uploaded CPython 3.8 macOS 10.9+ universal2 (ARM64, x86-64)

kaldialign-0.7.2-cp37-cp37m-win_amd64.whl (63.2 kB view hashes)

Uploaded CPython 3.7m Windows x86-64

kaldialign-0.7.2-cp37-cp37m-win32.whl (55.9 kB view hashes)

Uploaded CPython 3.7m Windows x86

kaldialign-0.7.2-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (82.9 kB view hashes)

Uploaded CPython 3.7m manylinux: glibc 2.17+ x86-64

kaldialign-0.7.2-cp37-cp37m-manylinux_2_17_i686.manylinux2014_i686.whl (89.2 kB view hashes)

Uploaded CPython 3.7m manylinux: glibc 2.17+ i686

kaldialign-0.7.2-cp37-cp37m-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (79.5 kB view hashes)

Uploaded CPython 3.7m manylinux: glibc 2.17+ ARM64

kaldialign-0.7.2-cp36-cp36m-win_amd64.whl (63.2 kB view hashes)

Uploaded CPython 3.6m Windows x86-64

kaldialign-0.7.2-cp36-cp36m-win32.whl (55.9 kB view hashes)

Uploaded CPython 3.6m Windows x86

kaldialign-0.7.2-cp36-cp36m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (82.8 kB view hashes)

Uploaded CPython 3.6m manylinux: glibc 2.17+ x86-64

kaldialign-0.7.2-cp36-cp36m-manylinux_2_17_i686.manylinux2014_i686.whl (89.1 kB view hashes)

Uploaded CPython 3.6m manylinux: glibc 2.17+ i686

Supported by

AWS AWS Cloud computing and Security Sponsor Datadog Datadog Monitoring Fastly Fastly CDN Google Google Download Analytics Microsoft Microsoft PSF Sponsor Pingdom Pingdom Monitoring Sentry Sentry Error logging StatusPage StatusPage Status page