Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

Documentation Status

k2

The vision of k2 is to be able to seamlessly integrate Finite State Automaton (FSA) and Finite State Transducer (FST) algorithms into autograd-based machine learning toolkits like PyTorch and TensorFlow. For speech recognition applications, this should make it easy to interpolate and combine various training objectives such as cross-entropy, CTC and MMI and to jointly optimize a speech recognition system with multiple decoding passes including lattice rescoring and confidence estimation. We hope k2 will have many other applications as well.

One of the key algorithms that we want to make efficient in the short term is pruned composition of a generic FSA with a "dense" FSA (i.e. one that corresponds to log-probs of symbols at the output of a neural network). This can be used as a fast implementation of decoding for ASR, and for CTC and LF-MMI training. This won't give a direct advantage in terms of Word Error Rate when compared with existing technology; but the point is to do this in a much more general and extensible framework to allow further development of ASR technology.

Implementation

A few key points on our implementation strategy.

Most of the code is in C++ and CUDA. We implement a templated class Ragged, which is quite like TensorFlow's RaggedTensor (actually we came up with the design independently, and were later told that TensorFlow was using the same ideas). Despite a close similarity at the level of data structures, the design is quite different from TensorFlow and PyTorch. Most of the time we don't use composition of simple operations, but rely on C++11 lambdas defined directly in the C++ implementations of algorithms. The code in these lambdas operate directly on data pointers and, if the backend is CUDA, they can run in parallel for each element of a tensor. (The C++ and CUDA code is mixed together and the CUDA kernels get instantiated via templates).

It is difficult to adequately describe what we are doing with these Ragged objects without going in detail through the code. The algorithms look very different from the way you would code them on CPU because of the need to avoid sequential processing. We are using coding patterns that make the most expensive parts of the computations "embarrassingly parallelizable"; the only somewhat nontrivial CUDA operations are generally reduction-type operations such as exclusive-prefix-sum, for which we use NVidia's cub library. Our design is not too specific to the NVidia hardware and the bulk of the code we write is fairly normal-looking C++; the nontrivial CUDA programming is mostly done via the cub library, parts of which we wrap with our own convenient interface.

The Finite State Automaton object is then implemented as a Ragged tensor templated on a specific data type (a struct representing an arc in the automaton).

Autograd

If you look at the code as it exists now, you won't find any references to autograd. The design is quite different to TensorFlow and PyTorch (which is why we didn't simply extend one of those toolkits). Instead of making autograd come from the bottom up (by making individual operations differentiable) we are implementing it from the top down, which is much more efficient in this case (and will tend to have better roundoff properties).

An example: suppose we are finding the best path of an FSA, and we need derivatives. We implement this by keeping track of, for each arc in the output best-path, which input arc it corresponds to. (For more complex algorithms an arc in the output might correspond to a sum of probabilities of a list of input arcs). We can make this compatible with PyTorch/TensorFlow autograd at the Python level, by, for example, defining a Function class in PyTorch that remembers this relationship between the arcs and does the appropriate (sparse) operations to propagate back the derivatives w.r.t. the weights.

Current state of the code

A lot of the code is still unfinished (Sep 11, 2020). We finished the CPU versions of many algorithms and this code is in k2/csrc/host/; however, after that we figured out how to implement things on the GPU and decided to change the interfaces so the CPU and GPU code had a more unified interface. Currently in k2/csrc/ we have more GPU-oriented implementations (although these algorithms will also work on CPU). We had almost finished the Python wrapping for the older code, in the k2/python/ subdirectory, but we decided not to release code with that wrapping because it would have had to be reworked to be compatible with our GPU algorithms. Instead we will use the interfaces drafted in k2/csrc/ e.g. the Context object (which encapsulates things like memory managers from external toolkits) and the Tensor object which can be used to wrap tensors from external toolkits; and wrap those in Python (using pybind11). The code in host/ will eventually be either deprecated, rewritten or wrapped with newer-style interfaces.

Plans for initial release

We hope to get the first version working in early October. The current short-term aim is to finish the GPU implementation of pruned composition of a normal FSA with a dense FSA, which is the same as decoder search in speech recognition and can be used to implement CTC training and lattice-free MMI (LF-MMI) training. The proof-of-concept that we will release initially is something that's like CTC but allowing more general supervisions (general FSAs rather than linear sequences). This will work on GPU. The same underlying code will support LF-MMI so that would be easy to implement soon after. We plan to put example code in a separate repository.

Plans after initial release

We will then gradually implement more algorithms in a way that's compatible with the interfaces in k2/csrc/. Some of them will be CPU-only to start with. The idea is to eventually have very rich capabilities for operating on collections of sequences, including methods to convert from a lattice to a collection of linear sequences and back again (for purposes of neural language model rescoring, neural confidence estimation and the like).

Quick start

Want to try it out without installing anything? We have setup a Google Colab.

Caution: k2 is not nearly ready for actual use! We are still coding the core algorithms, and hope to have an early version working by early October.

Metadata

Release files for k2 0.3.5.dev20210608

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distributions (wheels)

Table of built distributions (wheels) for k2 0.3.5.dev20210608
File
k2-0.3.5.dev20210608-py38-none-any.whl Python 3.8 none any Details
k2-0.3.5.dev20210608-py37-none-any.whl Python 3.7 none any Details
k2-0.3.5.dev20210608-py36-none-any.whl Python 3.6 none any Details
k2-0.3.5.dev20210608-cp38-cp38-macosx_10_15_x86_64.whl CPython 3.8 CPython 3.8 macOS 10.15+ x86-64 Details
k2-0.3.5.dev20210608-cp37-cp37m-macosx_10_15_x86_64.whl CPython 3.7 CPython 3.7 pymalloc macOS 10.15+ x86-64 Details
k2-0.3.5.dev20210608-cp36-cp36m-macosx_10_15_x86_64.whl CPython 3.6 CPython 3.6 pymalloc macOS 10.15+ x86-64 Details

Total release size: 170.3 MB

Release files / k2-0.3.5.dev20210608-py38-none-any.whl

Download URL k2-0.3.5.dev20210608-py38-none-any.whl
Size 55.4 MB
Tags Python 3.8
SHA-256 checksum
How to use checksums
22c000543b863f2d0fa5093e628c52d9baa3599ff58296a4b4ec41f9d45d0d24
BLAKE2b-256 checksum
How to use checksums
95dcc7b6000d58e65b335d342954512bcc8d495845ae019c8a9165d35443e4e6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.4.1 importlib_metadata/4.5.0 pkginfo/1.7.0 requests/2.25.1 requests-toolbelt/0.9.1 tqdm/4.61.0 CPython/3.8.10

Release files / k2-0.3.5.dev20210608-py37-none-any.whl

Download URL k2-0.3.5.dev20210608-py37-none-any.whl
Size 55.4 MB
Tags Python 3.7
SHA-256 checksum
How to use checksums
c2bb583f2678fe8f769f6cc3578736d3341a9ccb95598f73e306482e19eddf06
BLAKE2b-256 checksum
How to use checksums
123e98a06b8e13eb7af2515c145edc64ca81d4e73b5bbbad397745fd6b6544b8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.4.1 importlib_metadata/4.5.0 pkginfo/1.7.0 requests/2.25.1 requests-toolbelt/0.9.1 tqdm/4.61.0 CPython/3.7.10

Release files / k2-0.3.5.dev20210608-py36-none-any.whl

Download URL k2-0.3.5.dev20210608-py36-none-any.whl
Size 55.4 MB
Tags Python 3.6
SHA-256 checksum
How to use checksums
cab06e00fe460be2de9be11cea27c449238280674f7f7df6d36cf958ad63c2ec
BLAKE2b-256 checksum
How to use checksums
9c39661590258fd41f4366c7b9f55e223a37332edafcfddc3c7247d60d8441da
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.4.1 importlib_metadata/4.5.0 pkginfo/1.7.0 requests/2.25.1 requests-toolbelt/0.9.1 tqdm/4.61.0 CPython/3.6.13

Release files / k2-0.3.5.dev20210608-cp38-cp38-macosx_10_15_x86_64.whl

Download URL k2-0.3.5.dev20210608-cp38-cp38-macosx_10_15_x86_64.whl
Size 1.4 MB
Tags CPython 3.8 macOS 10.15+ x86-64
SHA-256 checksum
How to use checksums
dbb36d7d792e31e56db2b776ae5ff4a35a152796efb9b33b6debe7af7142387f
BLAKE2b-256 checksum
How to use checksums
050678a924f005da6688b4c42b0ea31f7868e6d932ac1f4eac9d6fee226ed171
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.4.1 importlib_metadata/4.5.0 pkginfo/1.7.0 requests/2.25.1 requests-toolbelt/0.9.1 tqdm/4.61.0 CPython/3.8.6

Release files / k2-0.3.5.dev20210608-cp37-cp37m-macosx_10_15_x86_64.whl

Download URL k2-0.3.5.dev20210608-cp37-cp37m-macosx_10_15_x86_64.whl
Size 1.4 MB
Tags CPython 3.7 CPython 3.7 pymalloc macOS 10.15+ x86-64
SHA-256 checksum
How to use checksums
d98a0b3811e79a60f9524387d0d709cf061e12a3672e12fe5f2f7eb8a540e082
BLAKE2b-256 checksum
How to use checksums
ab3b65de1b8351d41649440d7a393da30cb498aa80c0307940a6d190fd962a4e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.4.1 importlib_metadata/4.5.0 pkginfo/1.7.0 requests/2.25.1 requests-toolbelt/0.9.1 tqdm/4.61.0 CPython/3.7.10

Release files / k2-0.3.5.dev20210608-cp36-cp36m-macosx_10_15_x86_64.whl

Download URL k2-0.3.5.dev20210608-cp36-cp36m-macosx_10_15_x86_64.whl
Size 1.4 MB
Tags CPython 3.6 CPython 3.6 pymalloc macOS 10.15+ x86-64
SHA-256 checksum
How to use checksums
b9c489e68e59bee0c20ae70bcd30d0aa21eba22a05dad036f79b8e127ad23568
BLAKE2b-256 checksum
How to use checksums
95b767c26af10b70d176c655192f0be1648923c4ee88829856d27274d1b6dc9e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.4.1 importlib_metadata/4.5.0 pkginfo/1.7.0 requests/2.25.1 requests-toolbelt/0.9.1 tqdm/4.61.0 CPython/3.6.13

Release history Release notifications | RSS feed

1.24.1

4 release files

1.24.0

4 release files

1.23.4

9 release files

1.23.2

8 release files

1.23.1

8 release files

1.22

8 release files

1.21

4 release files

1.20

4 release files

1.19

6 release files

1.18

6 release files

1.17

6 release files

1.16

6 release files

1.15.1

6 release files

1.15

6 release files

1.14

6 release files

1.13

6 release files

1.12

6 release files

1.11

6 release files

1.10

6 release files

1.9

6 release files

1.8

6 release files

1.7

6 release files

1.6

6 release files

1.5

6 release files

1.4

6 release files

1.3

6 release files

1.2

6 release files

1.1

6 release files

1.0

6 release files

0.3.5

6 release files

This release

0.3.0

3 release files

0.1.2

3 release files

0.1

3 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page