Skip to main content

compute-wer

PyPI License

A Python package for computing Word Error Rate (WER) and Sentence Error Rate (SER) for evaluating speech recognition systems.

Features

  • Compute WER and SER for speech recognition evaluation
  • Support for both word-level and character-level WER calculation
  • Detailed alignment visualization between reference and hypothesis texts
  • Support for case-sensitive and case-insensitive matching
  • Cluster-based error analysis (Chinese, English, Numbers, etc.)
  • Support for filtering results based on maximum WER threshold
  • Handle tagged text with option to remove tags
  • Support for ignoring punctuation in WER calculation

Installation

pip install compute-wer

Usage

Command Line Interface

Basic Usage

Compute WER between reference and hypothesis texts:

# Compare two texts directly
compute-wer "你好世界" "你好"

# Compare texts from files
compute-wer ref.txt hyp.txt wer.txt

File Format

The input files should contain lines in the format utterance_id text. For example:

ref.txt:

utt1 你好世界
utt2 欢迎使用 compute-wer

hyp.txt:

utt1 你好
utt2 欢迎使用 computer-wer

Advanced Options

# Character-level WER
compute-wer --char ref.txt hyp.txt

# Case-sensitive matching
compute-wer --case-sensitive ref.txt hyp.txt

# Sort results by utterance-id or WER
compute-wer --sort utt ref.txt hyp.txt
compute-wer --sort wer ref.txt hyp.txt

# Remove tags from text
compute-wer --remove-tag ref.txt hyp.txt

# Filter results with WER <= 50%
compute-wer --max-wer 0.5 ref.txt hyp.txt

# Ignore specific words from a file
compute-wer --ignore-file ignore_words.txt ref.txt hyp.txt

# Ignore punctuation (except single quotes)
compute-wer --ignore-punctuation ref.txt hyp.txt

Python API

from compute_wer import Calculator

# Initialize calculator
calculator = Calculator(
    to_char=False,          # Character-level WER
    case_sensitive=False,   # Case-sensitive matching
    remove_tag=True,        # Remove tags from text
    ignore_punctuation=True,# Ignore punctuation (except single quotes)
    max_wer=float('inf')    # Maximum WER threshold
)

# Calculate WER
wer = calculator.calculate("你好世界", "你好")
print(f"WER: {wer}")
print(f"Reference : {' '.join(wer.reference)}")
print(f"Hypothesis: {' '.join(wer.hypothesis)}")

# Get overall statistics
overall_wer, cluster_wers = calculator.overall()
print(f"Overall WER: {overall_wer}")
for cluster, wer in cluster_wers.items():
    print(f"{cluster} WER: {wer}")

CLI Options

Option Description
--char, -c Use character-level WER instead of word-level WER
--sort, -s Sort the hypotheses by utterance-id or WER in ASC
--case-sensitive, -cs Use case-sensitive matching
--remove-tag, -rt Remove tags from the reference and hypothesis
--ignore-punctuation, -ip Ignore punctuation (except single quotes)
--ignore-file, -ig Path to the ignore file
--max-wer, -mw Filter hypotheses with WER <= this value
--verbose, -v Print verbose output

Output Format

The output includes detailed alignment information:

utt: utt1
WER: 50.00 % N=4 Cor=2 Sub=0 Del=2 Ins=0
ref: 你 好 世 界
hyp: 你 好

===========================================================================
Overall -> 50.00 % N=4 Cor=2 Sub=0 Del=2 Ins=0
Chinese -> 50.00 % N=4 Cor=2 Sub=0 Del=2 Ins=0
SER -> 100.00 % N=1 Cor=0 Err=1 ML=1 MH=0
===========================================================================

Where:

  • N: Total number of reference words/characters
  • Cor: Correct matches
  • Sub: Substitutions
  • Del: Deletions
  • Ins: Insertions
  • SER: Sentence Error Rate
  • ML: Missing Labels (Extra Hypotheses)
  • MH: Missing Hypotheses (Extra Labels)

License

MIT License

Metadata

Release files for compute-wer 0.2.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for compute-wer 0.2.5
File Size Uploaded
compute_wer-0.2.5.tar.gz 14.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for compute-wer 0.2.5
File Interpreter ABI Platform
compute_wer-0.2.5-py3-none-any.whl Python 3 none any Details

Total release size: 27.1 kB

Release files / compute_wer-0.2.5.tar.gz

Download URL compute_wer-0.2.5.tar.gz
Size 14.0 kB
Tags Source
SHA-256 checksum
How to use checksums
fa25ccf18bc6af5cc4a55a35c4b3bbf0475eadef41831d90968bf2ce94a49767
BLAKE2b-256 checksum
How to use checksums
e8638936e81b6413d7ed34f0ce1dc87d994042f64b8248a5e4fc37655e88a9b6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.3

Release files / compute_wer-0.2.5-py3-none-any.whl

Download URL compute_wer-0.2.5-py3-none-any.whl
Size 13.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
87991734c09e226c117dc4c758a6a726065bf77b52ab843bf76715fca18fc117
BLAKE2b-256 checksum
How to use checksums
815e12fc9ab9d50dec4cd764ce9d305076a5fbd0c928e593c9521f3bd7d60ecb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.3

Release history Release notifications | RSS feed

This release

0.2.5 This release

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.9

2 release files

0.1.8

2 release files

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

1 release file

0.1.1

1 release file

0.1.0

1 release file

0.0.9

1 release file

0.0.8

1 release file

0.0.7

1 release file

0.0.6

1 release file

0.0.5

1 release file

0.0.4

1 release file

0.0.3

1 release file

0.0.2

1 release file

0.0.1

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page