Agent-oriented word error rate, named entity F1 score, hallucination, and multi-speaker cpWER/tcpWER error metrics package.
Project description
AgWER: Agent-oriented Word Error Rate
agwer is a simple and efficient Python package to evaluate speech recognition and voice agents. A 40 KB wheel that imports in under 20 ms, fast on CPU and native on Apple Silicon M-series.
- note: this PyPI package was developed organically by a human programmer since 2025. We work together with the great Fable since 2026. 🙂
- we welcome contributions from the community.
agwer supports the classic ASR similarity measures and the agentic ones:
- word error rate (WER) and sentence error rate (SER)
- character error rate (CER)
- recoverable information ratio (RIR)
- harmful edit rate (HER)
- named entity F1 score (NF1)
- word hallucination rate (WHR)
- concatenated minimum-permutation word error rate (cpWER) and its time-constrained variant (tcpWER) for multi-speaker ASR
The agentic measures (3 and 4) evaluate systems that read n-best hypotheses and decide when to edit and when to abstain, such as LLM error correctors and dictation agents. The metric guide explains what question each measure answers, and the example session walks through a complete evaluation from install to CLI.
Installation
With uv:
uv add agwer # as a project dependency
uv pip install agwer # into the active environment
Or with pip (Python >= 3.9):
pip install agwer
The test suite ships inside the package, so any install can verify itself:
pip install "agwer[test]"
python -m pytest --pyargs agwer # 121 tests, a few seconds
Usage
The simplest use case is computing the word error rate of a dictated utterance:
import agwer
ref = "please schedule the quarterly budget review for tuesday march twenty first at nine thirty and invite the design team"
asr_decoded = "please schedule the quarterly budget review for tuesday march twenty first at nine thirty and invite the desire team"
agwer.wer(ref, asr_decoded) # 0.0526, one broken word in nineteen
All measures accept a single string or a list of strings. Lists are pooled
corpus-level (total errors over total reference words), and mer, wil,
wip, cer, ser work the same way:
import agwer
refs = ["send the revised contract to the legal team before the board meeting on friday afternoon",
"remind me to pick up the prescription from the pharmacy after the dentist appointment"]
hyps = ["send the revised contract to the legal team before the bored meeting on friday afternoon",
"remind me to pick up the prescription from the pharmacy after the dentist appointment"]
agwer.wer(refs, hyps) # 0.0345, corpus WER
agwer.cer(refs, hyps) # 0.0116, corpus CER
Evaluating a corrector or voice agent
With n-best input, one call computes everything. Real data end to end: three real Whisper 5-best decodes from the MIT-licensed HyPoradise benchmark (WSJ), each with the verbatim output of a real LLM corrector. The three utterances show the three things a corrector can do: restore a word from another hypothesis, fix a reading that every hypothesis got wrong, and break a name that was already right.
refs = [
"quote obviously we were disappointed we did not get a larger award",
"he retired as a partner in nineteen eighty three and as counsel in nineteen eighty six",
"a senior painewebber official said the firm hopes the job can be cut through attrition since the turnover in such positions tends to be high",
]
year = ("he retired as a partner in one thousand, nine hundred and eighty-three "
"and as counsel in one thousand, nine hundred and eighty-six")
firm = ("official said the firm hopes the job can be cut through attrition "
"since the turnover in such positions tends to be high")
nbest = [ # real Whisper 5-best lists; nbest[i][0] is the 1-best
["obviously we were disappointed if we did not get a larger award",
"quote obviously we were disappointed we did not get a larger award",
"obviously we were disappointed if we did not get a larger award",
"quote obviously we were disappointed we did not get a larger award",
"quote obviously we were disappointed we did not get a larger award"],
[year] * 5, # all five hypotheses agree on the wrong year reading
[f"a senior {x} {firm}" for x in
["painewebber", "payne weber", "payne weber", "paine webber", "pain weber"]],
]
lm_corrected = [ # the LLM corrector's verbatim outputs
"quote obviously we were disappointed we did not get a larger award",
"he retired as a partner in nineteen eighty three and as counsel in nineteen eighty six",
"a senior paine webber official said the firm hopes the job can be cut through attrition since the turnover in such positions tends to be high",
]
out = agwer.evaluate(refs, lm_corrected, nbest=nbest)
out.wer_1best # 0.2264 the raw ASR
out.wer_oracle # 0.1887 o_nb: the best any reranker could reach
out.wer_compositional # 0.0377 o_cp: recombining n-best tokens cannot write "nineteen"
out.wer_corrected # 0.0377
out.rir # 5.0 five times the reranking headroom: generative correction
out.her # 0.3333 one of the three consequential edits broke a correct name
Real decodes also show why reranking alone cannot save a voice agent. In another HyPoradise utterance, the command "leaving after noon" comes back as "leaving afternoon" in all five hypotheses. The query's meaning flips, every reranker is helpless, and only a corrector (and agwer's oracles) can see it.
When the corrector beats every hypothesis (ρ > 1)
Look at the second utterance above. All five hypotheses read the year the same wrong way, so the reranking oracle equals the 1-best and offers zero headroom. Recombining n-best tokens cannot help either, because the word "nineteen" appears nowhere in the list. The corrector still writes "nineteen eighty three", the dictation convention it knows from context. Supplying truth that exists in no hypothesis is generative correction. It is why ρ can exceed 1, and it is exactly what plain WER, reranking metrics, and oracle bounds cannot see. RIR was built to measure it.
Dictating to a coding agent has the same structure: package names and code
terms (uv, pytest, agwer) are exactly the tokens ASR breaks and
exactly the ones the agent can restore from context.
On the full 30-utterance WSJ set these corrections come from, the corrector scores ρ = 2.0 with HER = 0.29: it recovers twice the n-best headroom, and roughly two of every seven consequential edits still break something. That pair of numbers, not WER alone, is the act/abstain profile agwer exists to report.
Entity F1 and Word Hallucination Rate
Two failure modes that plain WER cannot name. First, entity errors hide
inside acceptable WER: one substituted amount in a 21-word transfer request
is a comfortable 4.8% WER and a broken transaction. entity_f1 scores an
explicit token subset with the same alignment engine as WER:
ref = "okay so please transfer five hundred dollars to the savings account before friday and send me a confirmation text right away"
hyp = "okay so please transfer nine hundred dollars to the savings account before friday and send me a confirmation text right away"
agwer.wer(ref, hyp) # 0.0476, looks fine
m = agwer.entity_f1(ref, hyp, entities={"five", "hundred"})
m["recall"], m["entity_wer"] # 0.5, 0.5: it is not fine
agwer.entity_f1(ref, hyp, predicate=agwer.numeric_tokens) # ready-made subset: digits and spelled numbers
Second, LLM correctors make text up, and the two mechanisms differ. A
corrector can invent a word the ASR never produced, or it can loop
autoregressively and repeat words it really heard (one system in
arXiv:2408.16180 repeats a real segment
eleven times, so a vocabulary check sees nothing wrong).
word_hallucination_rate uses occurrence-bounded attribution to catch both,
works with a single 1-best or a full n-best list, and never penalizes
copying an ASR error or a correct generative recovery:
m = agwer.word_hallucination_rate(
["cancel my nine a m meeting"], # references
["cancel my nine a m meeting meeting meeting"], # corrector outputs
["cancel my nine a m meeting"]) # hypotheses (1-best or n-best)
m["whr"] # 0.25: two of eight output tokens hallucinated
m["repetition_hallucinations"] # 2, an autoregressive loop over a heard word
m["novel_hallucinations"] # 0, nothing invented outright
The decomposition also reports passed_through_errors (copied ASR errors,
not hallucination) and generative_tokens (correct words no hypothesis
contained, the ρ > 1 mechanism).
HER comes in two granularities. her_granularity="utterance" (the default)
is the accounting behind the paper's reported values, and "token" follows
the formal per-edit definition. A sentence where the corrector fixes one
token and breaks another is neutral at utterance granularity but
helpful=1, harmful=1 at token granularity, so report which one you used.
Multi-speaker ASR: cpWER
Meeting and conversation transcripts come speaker-attributed, and the
speaker labels a system assigns are arbitrary. cpwer concatenates each
speaker's utterances and finds the label permutation with the fewest word
errors (the MeetEval definition, validated against it):
ref = {"alice": "let us start with the quarterly numbers",
"bob": "sounds good i will share my screen"}
hyp = {"spk0": "sounds good i will share my screen",
"spk1": "let us start with the quality numbers"}
agwer.cpwer(ref, hyp) # 0.0714, labels permuted, one word wrong
agwer.cp_statistics(ref, hyp)["assignment"] # [('alice', 'spk1'), ('bob', 'spk0')]
Values may be single strings or lists of utterances (concatenated in the
given order). A missed speaker counts all its words as deletions, an extra
hypothesis speaker all its words as insertions; cp_statistics reports the
missed and false-alarm speaker counts.
With timestamped segments, tcpwer adds the time constraint of the
CHiME-7/8 official metric: words may only match when their time intervals
overlap (reference words get character-based intervals within their
segment; hypothesis words get their interval center expanded by a collar,
5 seconds by default). Semantics match MeetEval's tcp_word_error_rate,
pinned by a golden fixture, and the pure-Python time-windowed
implementation measures 2.6 to 7.5 times faster on meeting-sized inputs:
ref = [{"speaker": "alice", "words": "let us start with the quarterly numbers",
"start_time": 0.0, "end_time": 3.5},
{"speaker": "bob", "words": "sounds good i will share my screen",
"start_time": 3.5, "end_time": 6.5}]
hyp = [{"speaker": "spk0", "words": "sounds good i will share my screen",
"start_time": 3.6, "end_time": 6.6},
{"speaker": "spk1", "words": "let us start with the quality numbers",
"start_time": 0.1, "end_time": 3.6}]
agwer.tcpwer(ref, hyp, collar=5) # 0.0714, same value with correct timing
agwer.tcp_statistics(ref, hyp) # adds assignment and speaker counts
Normalization
Normalization is the main reason WER numbers are incomparable across papers.
Every entry point takes normalize= (any Callable[[str], str]), and agwer
ships the standards.
A dictation agent that says amounts out loud looks 87% wrong against the written form, until you normalize:
ref = "the invoice total came to $1,250.75 after the 15% discount was applied on march 3rd"
hyp = ("the invoice total came to one thousand two hundred fifty dollars "
"and seventy five cents after the fifteen percent discount was applied on march third")
agwer.wer(ref, hyp) # 0.867 (!)
agwer.wer(ref, hyp, normalize=agwer.EnglishTextNormalizer()) # 0.0
| normalizer | what it does |
|---|---|
None (the general measures' default) |
scores strings exactly as given |
agwer.default_normalize (the agentic default) |
conservative: lowercase, keep apostrophes, strip other punctuation |
BasicTextNormalizer() |
language-agnostic: symbols, brackets, optional diacritic folding |
EnglishTextNormalizer() |
the Whisper English normalizer: spelled numbers to digits, currency, contractions, British to American spelling; cached=True adds an LRU for agent loops |
The Whisper normalizers have their behavior pinned byte-identical to the original by golden tests. Report which normalizer you used; it is part of the metric.
CLI
agwer results.jsonl # {"reference","corrected","nbest"} per line
agwer results.jsonl --json --her-granularity token
Performance
The hot path is batched RapidFuzz, and every aggregate agwer computes is
count-additive, so large corpora parallelize exactly. Results are
identical for any worker count: evaluate(..., workers=8).
Typical performance on Apple Silicon (M-series, 14 cores) for the full agentic evaluation (three WERs, both oracles, RIR, and HER) of 5-best corpora:
| scenario | time |
|---|---|
| 10k utterances (single-threaded) | 0.15 s |
| 100k utterances (single-threaded) | 1.79 s |
| 100k utterances (8 workers) | 0.38 s (4.7×) |
| 1M utterances (single-threaded) | 18.8 s |
| 1M utterances (8 workers) | 4.0 s (4.7×) |
Workers pay process startup, so they win from roughly 100k utterances up. On
macOS and Windows, call from a if __name__ == "__main__" guard, as with any
multiprocessing. Reproduce on your machine:
python -m agwer.bench --workers 8
Apple Silicon
agwer is native on Apple Silicon, with no separate
install: pip and uv select the arm64 wheel automatically, and RapidFuzz
ships compiled arm64-darwin extensions, so the C++ edit-distance core runs
natively on M-series machines. workers= then scales the whole pipeline
across performance cores (see the table above).
- Planned next is an optional
agwer[mlx]extra for embedding-based semantic metrics on the Apple GPU via MLX: semantic inference is the one place extra hardware genuinely helps; edit distance does not need it.
Compatibility & reproducibility
Measure semantics match jiwer, validated bit-identical on a 600-corpus golden set pinned in the package test suite.
agwer gratefully builds on and learns from these projects:
- RapidFuzz provides the C++ bit-parallel edit-distance engine, agwer's only dependency; FuzzyMatch shows the same design done right in Swift.
- jiwer defined the measure semantics and the easy, friendly API style this package follows.
- OpenAI Whisper created the English text normalizer that became the community standard for ASR evaluation. We vendor it faithfully under its MIT license, with attribution.
- NVIDIA NeMo text processing sets the bar for full text normalization across languages. Its semiotic class taxonomy guides our normalization roadmap.
- MeetEval defined the cpWER and tcpWER reference implementations for multi-speaker ASR. agwer matches it on golden sets of 87 (cpWER) and 53 (tcpWER) pinned cases; for ORC/MIMO variants and DER, use MeetEval itself.
License & contact
MIT. If you use or redistribute substantial portions of agwer, please
acknowledge the project rather than republishing it under a different
name. Vendored Whisper normalizers: MIT (see
src/agwer/normalizers/LICENSE_WHISPER).
Contact: Huck Yang huckiyang@ieee.org
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file agwer-0.4.10.tar.gz.
File metadata
- Download URL: agwer-0.4.10.tar.gz
- Upload date:
- Size: 93.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: uv/0.11.30 {"installer":{"name":"uv","version":"0.11.30","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8e376b88c899af3820502e38ae55754ce0bbd78b5d2810912a8045a1ab5d64ce
|
|
| MD5 |
aa97a7a1ede36c5d9e8b9a1c91650c99
|
|
| BLAKE2b-256 |
d53da5884ba49043f506afb9fb0e68cfd697dc4621c7310d8b5fb978853ed15c
|
File details
Details for the file agwer-0.4.10-py3-none-any.whl.
File metadata
- Download URL: agwer-0.4.10-py3-none-any.whl
- Upload date:
- Size: 104.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: uv/0.11.30 {"installer":{"name":"uv","version":"0.11.30","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
326e3a51271b3f5c2557a9fcef64db0e2d5860ed308a2c0269dce39aab3c40f7
|
|
| MD5 |
b939b2c49fcac8a2aa6a6fd9345f0e27
|
|
| BLAKE2b-256 |
1cb189ed6e688b09ff65bf0b5be4d38e3eb19962638f65e4bfcb1af598bb4b0c
|