Key takeaways
- Aggregate word error rate can hide serious performance differences across speaker groups and call conditions.
- Test the full task: names, numbers, intent, entities, correction, tool use, handoff, and outcome—not transcription alone.
- Build the evaluation population from the real service context with consent and careful governance.
- A model update can improve the average while worsening a smaller group; keep subgroup and intersectional regression tests.
01
Start with who the system needs to understand
List expected languages, regional accents, code-switching, age ranges, speech differences, phone networks, background environments, and common domain terms. Use service data and community input rather than choosing a handful of accents the team can imitate.
Collect or commission evaluation material with informed consent, a defined purpose, retention limits, and compensation where appropriate. Do not turn production calls into an unlabeled research dataset because the terms were vague enough to permit it.
02
Word error rate is a diagnostic, not the customer outcome
ASR research repeatedly finds uneven performance across speaker groups, and recent work shows that overall improvements do not guarantee smaller gaps. A transcript can also contain several errors while preserving intent, or one small error can change a medication, address, phone number, budget, or appointment date.
Swipe to compare every column
| Layer | Measure | Example harm |
|---|---|---|
| Recognition | Word and entity error by group and condition | A surname or postcode is repeatedly corrupted |
| Understanding | Intent and slot accuracy | Cancellation is interpreted as confirmation |
| Dialogue | Corrections, interruptions and recovery turns | The system keeps repeating the same mistake |
| Outcome | Completion, handoff, abandonment and complaint | A caller gives up before reaching a person |
03
Test intersections and difficult audio
Compare performance by relevant groups, then examine intersections where sample size permits responsible interpretation. Add mobile and landline codecs, speakerphone, noise, quiet speech, overlapping voices, and unstable connections. Report uncertainty and sample size rather than ranking groups from a tiny test set.
Manual inspection remains valuable. Research on regional dialect adaptation found that established aggregate measures can miss application-specific errors. A reviewer who understands the task can distinguish an inconsequential article from a wrong account number.
04
Give the caller a dignified recovery path
Let people repeat, spell, use the keypad, switch language, receive a confirmation, or reach a person without being blamed for recognition failure. Avoid cheerful loops that say “I didn’t get that” indefinitely. After a bounded number of failures, change the channel or hand off with context.
Run this suite before every meaningful ASR, model, prompt, telephony, or noise-suppression change. Fairness is not a certification the vendor can transfer to your use case; it is a performance question that must remain visible in operation.
Primary sources and further reading
Use the source material to validate details against your own context and current platform configuration.
- Towards Fair Speech Recognition
- Responsible Benchmarking of Fairness for Automatic Speech Recognition
- Adapting Whisper for Regional Dialects
- Sonos Voice Control Bias Assessment Dataset
This field note follows the XenGrowth editorial policy: primary sources where available, visible limitations, material review dates, and no invented first-hand experience.
Stay with the problem



