MultiMetric-Eval
English | 中文
MultiMetric-Eval is an evaluation toolkit centered on translation and speech translation. It provides a unified way to score text translation quality, speech output quality, preservation-related properties, and streaming latency.
What It Can Be Used For
This project is best suited for these directions:
- MT or S2TT text-side evaluation with
BLEU,chrF++,COMET, andBLEURT - S2ST evaluation by combining text quality, speech quality, speaker similarity, and latency
- Streaming or simultaneous speech translation latency evaluation with a custom agent
- Preservation analysis for speech translation outputs, including speaker similarity, emotion, and paralinguistic similarity
Core Modules
| Module | Main Use | Typical Metrics |
|---|---|---|
TranslationEvaluator |
Text-side translation quality | sacreBLEU, chrF++, COMET, BLEURT |
SpeechQualityEvaluator |
Naturalness and text-speech consistency | UTMOS, WER_Consistency, CER_Consistency |
SpeakerSimilarityEvaluator |
Speaker preservation | wavlm_similarity, resemblyzer_similarity |
EmotionEvaluator |
Emotion preservation or classification accuracy | Emotion2Vec_Cosine_Similarity, Audio_Emotion_Accuracy |
ParalinguisticEvaluator |
Non-verbal and paralinguistic preservation | Paralinguistic_Fidelity_Cosine, Acoustic_Event_Preservation_Rate, Acoustic_Event_Preservation_Macro_F1, Acoustic_Event_Preservation_Macro_Recall, Event_Aligned_Preservation_Rate, Conditional_Relative_Onset_Error |
LatencyEvaluator |
Streaming / simultaneous translation latency | StartOffset, ATD, CustomATD, RTF, Model_Generate_RTF |
Installation
Basic install:
pip install multimetriceval
Optional extras:
pip install "multimetriceval[comet]"
pip install "multimetriceval[whisper]"
pip install "multimetriceval[speech_quality]"
pip install "multimetriceval[emotion]"
pip install "multimetriceval[paralinguistics]"
pip install "multimetriceval[all]"
If you need BLEURT:
pip install git+https://github.com/lucadiliello/bleurt-pytorch.git
Import
PyPI package name:
multimetriceval
Python import name:
multimetric_eval
Example:
from multimetric_eval import TranslationEvaluator, SpeechQualityEvaluator
Quick Start
Quick-start scripts live under examples/.
Python examples:
examples/python/translation_eval.pyexamples/python/speech_quality_eval.pyexamples/python/speaker_similarity_eval.pyexamples/python/emotion_eval.pyexamples/python/paralinguistic_eval.pyexamples/python/paralinguistic_identity_baseline.pyexamples/python/latency_eval.py
Shell examples:
examples/bash/install_extras.shexamples/bash/run_latency_cli.sh
Latency output now distinguishes two RTF variants:
Real_Time_Factor_(RTF): system-level RTF. This includes agent policy overhead, pre/post-processing, and other runtime costs around model inference.Model_Generate_RTF: model-level RTF. This is reported only when the agent explicitly records model inference time viarecord_model_inference_time(...)or returns it inSegment.config["model_inference_time"].
Examples
Examples have been moved into the examples/ directory.
Input Conventions
Common text inputs support:
- Python
List[str] .txtfiles with one sample per line.jsonfiles
Common audio inputs support:
- folder path
- Python
List[str] .txtfiles.jsonfiles
Notes
- For
zh/ja/ko, the toolkit uses CJK-aware handling for text-side evaluation. SpeechQualityEvaluatorreturnsCER_Consistencyforzh/ja/ko, andWER_Consistencyfor most other languages.ParalinguisticEvaluatoralways supportsParalinguistic_Fidelity_Cosine, a continuous CLAP-based audio similarity score between source and target speech.- The discrete preservation branch is an utterance-level single-label task. With source-side gold labels, it reports
Acoustic_Event_Preservation_Rate,Acoustic_Event_Preservation_Macro_F1, andAcoustic_Event_Preservation_Macro_Recall. - If
source_onsets_msare available, the evaluator can also report alignment-aware metrics:Event_Aligned_Preservation_RateandConditional_Relative_Onset_Error. - Alignment is computed on relative onset position, not absolute wall-clock time. This makes it suitable for cross-lingual S2ST where source and target utterance durations naturally differ.
- If target-side onset timestamps are not provided, the default localizer estimates them with CLAP sliding-window scoring conditioned on the target event label.
- These alignment metrics should be interpreted as weak, coarse-grained alignment signals rather than timestamp-accurate event localization benchmarks.
- If source-side gold labels are not available, the evaluator can still run in prediction-only mode and reports
Predicted_Event_Consistency_Rate,Predicted_Event_Consistency_Macro_F1, andPredicted_Event_Consistency_Macro_Recall. - The default discrete predictor is a closed-set CLAP classifier over
candidate_labels. Users may replace it with any custom predictor object that implementspredict(audio_paths, candidate_labels). - The default event localizer is also replaceable. Custom localizers only need to implement
localize(audio_paths, labels, candidate_labels). - Dataset-specific label mapping is intentionally outside the core package. Pass
candidate_labelsandlabel_normalizerat call time so the same evaluator works across datasets without changing core code. - For offline environments,
clap_model_pathaccepts either a Hugging Face repo id or a local model directory or snapshot. - In S2S latency evaluation, alignment prefers the model's native transcript when available. If the model is audio-only, the evaluator can optionally use ASR fallback to prepare alignment text.
- For S2S forced alignment, pass language-appropriate MFA models through
alignment_acoustic_modelandalignment_dictionary_model. The defaults are English. - Some modules rely on optional dependencies or local model paths in offline environments.
License
MIT License
Metadata
Release files for multimetriceval 0.8.4
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| multimetriceval-0.8.4.tar.gz | 41.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| multimetriceval-0.8.4-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 85.6 kB
Release files / multimetriceval-0.8.4.tar.gz
| Download URL | multimetriceval-0.8.4.tar.gz |
|---|---|
| Size | 41.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
fa4325687356ee46198f224c513923382223c1ec69e6b42c63c4b0fd9e5f4edb
|
|
BLAKE2b-256 checksum How to use checksums |
30c83298929c26e1ce994891ef555e91822ec48aa8b0c40dc7c77619be5fa999
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.25
|
Release files / multimetriceval-0.8.4-py3-none-any.whl
| Download URL | multimetriceval-0.8.4-py3-none-any.whl |
|---|---|
| Size | 43.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c0e667bc0e14236001c63aa205553bbdc27502d5978159588a9575aadc11496b
|
|
BLAKE2b-256 checksum How to use checksums |
34e212d833fe92d94012fff7420f104da1fa4f529e19702d4300e54f79e16f15
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.25
|