Skip to main content

Convenience functions for generating ML features from audio data

Project description

Audio ML Spec Tools

Convenience functions for generating ML features from audio data. Breaks audio ML dependencies on torchaudio. Unlike pytorch features, these functions can be exported to ExecuTorch and ONNX with no issues.

Motivation

Except in specific circumstances like wav2vec, raw audio has proven to be a much worse input for ML models than spectrogram-based features across a wide variety of problem domains, including environmental sound classificarion (Guzhov et al. (2021)), singing technique classification (Yamamoto et al. (2021)), and ship classification (Xie, Ren, and Xu (2024)).

There is no scientific consensus on the relative benefits of mel-scale spectrograms, linear spectrograms, and MFCCs. Different researchers have shown good results with each type of spectrogram; see respectively Raponi, Oligeri, and Ali (2021), Jung at al. (2021), and Razani et al (2017).

With this library, you can easily try as many feature extraction methods as you want to see what works for your use case.

Prerequisites

  • Python 3.12 runtime
  • pip for package installation
  • Note that torchcodec depends on a system installation of FFmpeg

Installation

  • Install using pip:
pip install AudioMlSpecTools

Local Installation

Install the dependencies and library with pip:

pip install .

Usage

See examples/features.py.

Testing

# If needed, install test dependencies
# pip install .[test]

python3 -m coverage run -m unittest discover -s test -p "*_test.py" && python -m coverage report --skip-covered
python -m coverage html

Versioning

We use SemVer for versioning. For the versions available, see the tags on this repository.

Authors

  • Ryan Quinn - Initial work

License

MIT.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

audiomlspectools-0.8.0.tar.gz (16.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

audiomlspectools-0.8.0-py3-none-any.whl (23.1 kB view details)

Uploaded Python 3

File details

Details for the file audiomlspectools-0.8.0.tar.gz.

File metadata

  • Download URL: audiomlspectools-0.8.0.tar.gz
  • Upload date:
  • Size: 16.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.12

File hashes

Hashes for audiomlspectools-0.8.0.tar.gz
Algorithm Hash digest
SHA256 d2e5d4d50f36763e966f005c710ca3d4ef5c173d2393746f3475ec8fc262bc30
MD5 a636fba66c2ce11a00f9d1c5d8bed0c0
BLAKE2b-256 643bb2f4a5472d178178e4d392e6d5d8f9f0a945e3137e61c32be9707270251e

See more details on using hashes here.

File details

Details for the file audiomlspectools-0.8.0-py3-none-any.whl.

File metadata

File hashes

Hashes for audiomlspectools-0.8.0-py3-none-any.whl
Algorithm Hash digest
SHA256 3ebf5261a33c954bb12bf2333250f08aed69aa7716a6023589131b7ab0be3312
MD5 e25c0b5998dd08c7c2db7fa12a1778a5
BLAKE2b-256 4a2103b683e60c20da512491c9025615f472405cc796dfaca0262586949a5f20

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page