Skip to main content

Audio Generation Evaluation

This toolbox aims to unify audio generation model evaluation for easier future comparison.

Quick Start

First, prepare the environment

git clone https://github.com/haoheliu/audioldm_eval.git
cd audioldm_eval
pip install -e .

Second, generate test dataset by

python3 gen_test_file.py

Finally, perform a test run. A result for reference is attached here.

python3 test.py # Evaluate and save the json file to disk (example/paired.json)

Evaluation metrics

We have the following metrics in this toolbox:

  • FD: Frechet distance, realized by PANNs, a state-of-the-art audio classification model
  • FAD: Frechet audio distance
  • ISc: Inception score
  • KID: Kernel inception score
  • KL: KL divergence (softmax over logits)
  • KL_Sigmoid: KL divergence (sigmoid over logits)
  • PSNR: Peak signal noise ratio
  • SSIM: Structural similarity index measure
  • LSD: Log-spectral distance

The evaluation function will accept the paths of two folders as main parameters.

  1. If two folder have files with same name and same numbers of files, the evaluation will run in paired mode.
  2. If two folder have different numbers of files or files with different name, the evaluation will run in unpaired mode.

These metrics will only be calculated in paried mode: KL, KL_Sigmoid, PSNR, SSIM, LSD. In the unpaired mode, these metrics will return minus one.

Evaluation on AudioCaps and AudioSet

The AudioCaps test set consists of audio files with multiple text annotations. To evaluate the performance of AudioLDM, we randomly selected one annotation per audio file, which can be found in the accompanying json file.

Given the size of the AudioSet evaluation set with approximately 20,000 audio files, it may be impractical for audio generative models to perform evaluation on the entire set. As a result, we randomly selected 2,000 audio files for evaluation, with the corresponding annotations available in a json file.

For more information on our evaluation process, please refer to our paper.

Example

import torch
from audioldm_eval import EvaluationHelper

# GPU acceleration is preferred
device = torch.device(f"cuda:{0}")

generation_result_path = "example/paired"
target_audio_path = "example/reference"

# Initialize a helper instance
evaluator = EvaluationHelper(16000, device)

# Perform evaluation, result will be print out and saved as json
metrics = evaluator.main(
    generation_result_path,
    target_audio_path,
    limit_num=None # If you only intend to evaluate X (int) pairs of data, set limit_num=X
)

TODO

  • Add pretrained AudioLDM model.
  • Add CLAP score

Cite this repo

If you found this tool useful, please consider citing

@article{liu2023audioldm,
  title={AudioLDM: Text-to-Audio Generation with Latent Diffusion Models},
  author={Liu, Haohe and Chen, Zehua and Yuan, Yi and Mei, Xinhao and Liu, Xubo and Mandic, Danilo and Wang, Wenwu and Plumbley, Mark D},
  journal={arXiv preprint arXiv:2301.12503},
  year={2023}
}

Reference

https://github.com/toshas/torch-fidelity

https://github.com/v-iashin/SpecVQGAN

Metadata

Release files for audioldm-eval 0.0.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for audioldm-eval 0.0.3
File Size Uploaded
audioldm_eval-0.0.3.tar.gz 53.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for audioldm-eval 0.0.3
File Interpreter ABI Platform
audioldm_eval-0.0.3-py3-none-any.whl Python 3 none any Details

Total release size: 115.0 kB

Release files / audioldm_eval-0.0.3.tar.gz

Download URL audioldm_eval-0.0.3.tar.gz
Size 53.1 kB
Tags Source
SHA-256 checksum
How to use checksums
409c913420cf3add2e064fd4ce166abac3fea663232fa25f359241d0fbae952d
BLAKE2b-256 checksum
How to use checksums
2b6c38ef37a6793e5207ae59acee125d9ee631395c2168c9f5ffb776a128d92d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.9.2

Release files / audioldm_eval-0.0.3-py3-none-any.whl

Download URL audioldm_eval-0.0.3-py3-none-any.whl
Size 61.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
002b6b9e493918d18d31ed33aeb72176edfc0384618136e8e2e0f77990fa0d4d
BLAKE2b-256 checksum
How to use checksums
be493000bc8b4e2f4adbf11749f6475ded6a47ad4e58e2c92807ae459ad63787
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.9.2

Release history Release notifications | RSS feed

This release

0.0.3 This release

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page