Skip to main content
RVCBench logo

RVCBench

Comprehensive Voice Cloning Evaluation

NeurIPS 2026 arXiv PyPI Dataset Website CI License: CC0-1.0

Documentation · Paper · Website · Dataset · Demo · Evaluate your model · Reproduce the paper

News

  • 2026-10 · pip install "rvcbench[eval]": score your own model from ZipVoice- or Seed-TTS-style batch lists, compare models in one table, or use the metrics in your own code.
  • 2026-09 · RVCBench is accepted to NeurIPS 2026.

RVCBench is a general-purpose package for evaluating voice cloning. It brings speaker similarity, speech quality, intelligibility, pronunciation accuracy and emotion consistency into one scoring API, and provides ready-to-use datasets for evaluating a model across languages, speakers and recording conditions.

What you want to do Start here
Score your own audio with automatic speech metrics Python metrics API — use your own files and data
Evaluate your model with our data and get a complete report Dataset evaluation — export prompts, generate, score

Both workflows are included in the pip package. Your model can run in its own environment or through an API; RVCBench scores the audio it produces.

First time here? Follow the complete getting started guide. It walks through installation, required files, both workflows and reading the results.

Install

python -m pip install "rvcbench[eval]"
# Linux: install FFmpeg if it is not already available (e.g. sudo apt-get install ffmpeg).
rvcbench setup-scorers

Python 3.10+ and Linux are supported. rvcbench[eval] installs the package with all seven public metrics; plain rvcbench installs the runner and datasets without the optional scoring dependencies. setup-scorers downloads and verifies the metric models once. They are reused from your local cache. A GPU is optional; choose device="cpu" or --device cpu to score on CPU.

For a smaller initial download, select only the metrics you need: rvcbench setup-scorers --metrics sim wer speechmos. See the installation guide for CPU/GPU installation, caches and troubleshooting.

Score your own audio

from rvcbench import metrics

with metrics.Evaluator(["sim", "wer", "speechmos"], device="cpu") as evaluator:
    scores = evaluator.score(
        "generated.wav",
        reference="speaker_reference.wav",  # a recording of the intended speaker
        text="Hello there.",                 # what the generated audio should say
        language="en",
    )
    print(scores)  # dictionary with sim, wer and speechmos

Reuse the evaluator for a whole dataset: call score() for each file, and each metric model loads once. For every available metric, use metrics.Evaluator("all") and also pass target="same_text_recording.wav" for MCD and STOI. The target is a recording of the same text as the synthesized audio; the speaker reference may contain different words.

Metric Evaluates Required input besides generated audio
sim, sva Speaker similarity and speaker verification Speaker reference recording
wer Pronunciation/content accuracy via ASR word error rate Expected text; optional language
speechmos Predicted perceptual quality (UTMOS) None
mcd, stoi Acoustic distortion and intelligibility Same-text target recording
emotion Emotion consistency Reference recording

The metrics guide includes one-line functions, scoring all metrics, a batch example and metric definitions. No RVCBench dataset or model adapter is needed for this workflow.

Evaluate your model

RVCBench downloads the selected data, prepares reference clips and texts, and scores the audio your model generates. To run only selected scenarios, use rvcbench tasks --suite core-v1 to list them, then pass --tasks chinese (or several task IDs) to both prompts and score. See the scenario selection guide.

Start with the 52-utterance onboarding suite:

rvcbench prompts --suite onboarding-v1 --output prompts/
# Generate every prompt with your model. Save outputs as <id>.wav in outputs/my-model/.
rvcbench score --suite onboarding-v1 --generated outputs/my-model --output results/my-model --device cpu

prompts/ contains prompts.jsonl, prompts.tsv (ZipVoice) and prompts.lst (Seed-TTS-style scripts such as F5-TTS and CosyVoice). For example, a model environment with ZipVoice installed can run:

python3 -m zipvoice.bin.infer_zipvoice --model-name zipvoice \
  --test-list prompts/prompts.tsv --res-dir outputs/my-model

The model evaluation guide shows the exact file format and an example loop for your own inference function. Model inference is the step supplied by you; data preparation, metric scoring and report generation are automatic.

Suite Utterances to generate Use it for
onboarding-v1 52 Check the complete workflow
core-v1 480 Broad evaluation across languages, speaker groups, long speech, noise, compression and protection
full-v1 12,724 The larger evaluation datasets; protection tasks are currently covered by core-v1

Use the same suite name for prompts and score. Every input is checked against a fixed dataset revision; target recordings stay in the scoring data and are not exported with the prompts. submission.json contains per-task scores, coverage and failures. Task directories include per-sample scores and bootstrap confidence intervals. These suites are currently previews; published paper results use a separate protocol.

Resume an interrupted or partial evaluation with the original command plus --resume:

rvcbench score --suite onboarding-v1 --generated outputs/my-model --output results/my-model --device cpu --resume

Compare several models in one call, or compare reports you already have:

rvcbench score --suite core-v1 --generated outputs/model-a outputs/model-b --output results/comparison --device cuda
rvcbench compare results/model-a results/model-b --output results/compare

Comparison checks both the suite and scoring fingerprints. Different scoring environments must be rescored together, or explicitly inspected with --allow-incompatible, which disables ranking. The suite guide lists tasks, coverage and approximate scoring costs.

Evaluation coverage

The package covers clean voice cloning, multilingual and cross-lingual speech, speaker demographics, long-form speech, unusual text, background noise, overlapping speakers, compression, and protected references. Its seven public metrics cover identity, content, quality, intelligibility and emotion. The paper additionally studies deepfake detectability and an audio-LLM expression judge; these two components are not yet available through the packaged suites.

The paper defines 18 evaluations over 14,370 utterances from 204 speakers. The package also includes adapters for running existing models and tools for protection and denoising research.

Results

From the paper (arXiv v3). SIM, MOS, WER and MCD are on clean LibriTTS prompts; the last three columns are speaker similarity on other datasets. Bold marks the best value per column, and † marks models reported in the paper's appendix.

Model SIM ↑ MOS ↑ WER ↓ MCD ↓ VCTK SIM ↑ Chinese SIM ↑ Cross-lingual SIM ↑
Qwen3-TTS 0.61 4.39 0.05 5.79 0.62 0.72 0.67
IndexTTS 0.61 4.06 0.05 6.61 0.57 0.72 0.67
dots.tts † 0.60 4.17 0.06 6.11 0.57 0.68 0.61
CosyVoice 2 0.58 4.37 0.05 6.02 0.58 0.72 0.65
ZipVoice 0.58 4.13 0.05 7.09 0.55 0.71 0.63
MOSS-TTS v1.5 † 0.57 4.32 0.06 6.66 0.51 0.68 0.65
GLM-TTS 0.57 4.08 0.09 6.41 0.57 0.69 0.66
MaskGCT 0.57 3.93 0.09 6.91 0.56 0.67 0.63
Higgs TTS 3 † 0.56 4.23 0.05 6.32 0.45 0.63 0.30
F5-TTS 0.56 3.99 0.12 6.96 0.54 0.70 0.65
Higgs Audio 0.56 4.30 0.25 6.06 0.52 0.58 0.54
Fish Audio S2 † 0.54 4.37 0.04 6.16 0.51 0.66 0.62
MGM-Omni 0.54 4.28 0.09 5.82 0.45 0.71 0.63
PlayDiffusion 0.51 4.15 0.05 8.06 0.43 0.44 0.46
MOSS-TTSD 0.49 4.10 0.38 7.09 0.44 0.44 0.44
VibeVoice 0.48 3.83 0.23 6.76 0.44 0.56 0.53
FishSpeech 0.47 4.37 0.17 6.47 0.43 0.61 0.57
XTTS-v2 0.45 3.81 0.07 8.62 0.45 0.57 0.51
Spark-TTS 0.41 4.06 0.33 5.83 0.53 0.57 0.48
OZSpeech 0.39 3.21 0.06 6.87 0.25 0.00 0.17
OpenVoice V2 0.24 4.30 0.07 7.06 0.39 0.43 0.30
StyleTTS 2 0.23 4.30 0.05 6.81 0.24 0.11 0.21
Speaker similarity under anti-cloning protection (LibriTTS)

A lower SIM under protection means the protection hides the speaker's voice better.

Model Clean SafeSpeech SPEC Enkidu Gaussian POP
Qwen3-TTS 0.61 0.38 0.36 0.50 0.41 0.58
IndexTTS 0.61 0.35 0.32 0.47 0.39 0.57
dots.tts † 0.60 0.41 0.39 0.49 0.44 0.57
CosyVoice 2 0.58 0.32 0.30 0.45 0.38 0.55
ZipVoice 0.58 0.29 0.26 0.44 0.26 0.54
MOSS-TTS v1.5 † 0.57 0.33 0.31 0.43 0.33 0.53
GLM-TTS 0.57 0.33 0.31 0.44 0.39 0.53
MaskGCT 0.57 0.30 0.28 0.41 0.31 0.53
Higgs TTS 3 † 0.56 0.48 0.48 0.49 0.34 0.53
F5-TTS 0.56 0.21 0.18 0.43 0.14 0.52
Higgs Audio 0.56 0.26 0.24 0.43 0.27 0.52
Fish Audio S2 † 0.54 0.32 0.30 0.43 0.34 0.52
MGM-Omni 0.54 0.18 0.17 0.32 0.23 0.49
PlayDiffusion 0.51 0.17 0.15 0.34 0.16 0.47
MOSS-TTSD 0.49 0.24 0.22 0.34 0.25 0.08
VibeVoice 0.48 0.27 0.25 0.37 0.28 0.45
FishSpeech 0.47 0.24 0.21 0.33 0.23 0.01
XTTS-v2 0.45 0.26 0.24 0.31 0.24 0.41
Spark-TTS 0.41 0.13 0.11 0.14 0.06 0.36
OZSpeech 0.39 0.16 0.15 0.19 0.15 0.34
OpenVoice V2 0.24 0.18 0.18 0.19 0.18 0.24
StyleTTS 2 0.23 0.09 0.08 0.12 0.03 0.21

Results for compression, deepfake detectability, long-form and expressive speech are in the paper and on the website.

Reproduce the paper

To reproduce the paper, use the code released with it, kept on the v1 branch (tag v1.0). This branch, main, is v2: the same benchmark rebuilt as an installable package for evaluating new models.

git clone --branch v1 https://github.com/Nanboy-Ronan/RVCBench.git RVCBench-v1
v1 (branch v1) v2 (main)
Use it to Reproduce the paper Evaluate a new model
Status Frozen at the paper release Under active development
Install Clone, then pip install a list of packages pip install the rvcbench package
Run a built-in model python run_vc.py --config-name ... rvcbench run --config-name ...
Evaluate your own model Add an adapter to the codebase Score audio generated anywhere, or a one-file adapter
Evaluation data Full datasets Full datasets, plus the core-v1 and full-v1 suites

Run the built-in models

RVCBench includes adapters for 32 voice cloning models, 5 protection methods and a denoising stage. Each model runs in its own environment from envs/.

git clone https://github.com/Nanboy-Ronan/RVCBench.git && cd RVCBench
python -m pip install -e '.[eval]'
rvcbench run --config-name ots_vc/clean/libritts/qwen3_tts_ots dataset.speaker_id=1089 adversary.max_samples=5

Every run writes a per-sample run_manifest.json with input and output hashes, seeds, failures and metric coverage. See Running the built-in models and the run guide.

Documentation

Guide Covers
Getting started Your first evaluation, from installation to results
Evaluate your own model Prompt and output formats, batch lists, several models, adapters
Core and full suites Tasks, data, metrics and scoring time
Install the package pip installation, CPU/GPU, downloads and troubleshooting
Metrics API Speaker similarity, WER, MOS, MCD, STOI and emotion in your own code
Running the built-in models Installation options, supported models, protection and denoising
Datasets Hub folders, manifest format, preprocessing
Run guide Run records, resuming, scoring saved audio, timing
Codebase versions What changed between v1 and v2
Contributing Development setup, checks and repository layout

Citation

@inproceedings{jin2026rvcbench,
  title     = {RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models},
  author    = {Jin, Ruinan and Liao, Xinting and Yu, Hanlin and Pandya, Deval and Li, Xiaoxiao},
  booktitle = {Advances in Neural Information Processing Systems},
  url       = {https://arxiv.org/abs/2602.00443},
  year      = {2026}
}

License

CC0-1.0. Model checkpoints, upstream code and source corpora keep their own licenses. Questions and contributions are welcome through issues and pull requests, or at ruinanjin@alumni.ubc.ca.

Metadata

Release files for rvcbench 2.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for rvcbench 2.2.0
File Size Uploaded
rvcbench-2.2.0.tar.gz 6.1 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for rvcbench 2.2.0
File Interpreter ABI Platform
rvcbench-2.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 12.6 MB

Release files / rvcbench-2.2.0.tar.gz

Download URL rvcbench-2.2.0.tar.gz
Size 6.1 MB
Tags Source
SHA-256 checksum
How to use checksums
d6e9d6b25db762ae778674f7b0db1bb3f99af57c3f61b7111605f184a82fcf53
BLAKE2b-256 checksum
How to use checksums
0e40bd45e6ca50ab3eb7547111a36c200689da59ee1046e3c350596785a7985c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release files / rvcbench-2.2.0-py3-none-any.whl

Download URL rvcbench-2.2.0-py3-none-any.whl
Size 6.5 MB
Tags Python 3
SHA-256 checksum
How to use checksums
a796c6131ea8cfef3db8d78f8ef8d27371f3c5f05fcd888f2bb3ac5aaa6b3c0a
BLAKE2b-256 checksum
How to use checksums
7d8d42325d81d17c5c74f81e4f7f58845cf778707ff59c1126ae3b4e83da3adb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release history Release notifications | RSS feed

2.2.1

2 release files

This release

2.2.0 This release

2 release files

2.1.1

2 release files

2.1.0

2 release files

2.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page