# mu-bench tests speech transcription in five languages

> Researchers introduced a benchmark based on calls to an AI banking agent. It measures whether transcripts keep callers’ meaning, not just whether their words match exactly.

Oossa · 2026-09-29 · https://oossa.com/en/mu-bench-tests-speech-transcription-in-five-languages

A new benchmark called mu-bench evaluates how well speech-recognition systems transcribe callers giving details such as names, email addresses and confirmation codes. Its 4,270 utterances come from 250 calls to an AI banking agent in English, Spanish, Turkish, Vietnamese and Mandarin.

The researchers also propose Utterance Error Rate, which uses an AI judge to assess whether a transcript preserves meaning. On transcripts rated by people, the measure agreed more closely with human judgments than exact-match word error rate did. In tests of six commercial providers, the best scored 11.9% on Utterance Error Rate; Mandarin was the hardest language for all six.

## The facts

- The dataset contains 4,270 utterances from 250 calls, across five languages.
- Utterance Error Rate agreed with human raters at κ = 0.78, compared with 0.53 for exact-match word error rate on normalized text.

## Why it matters

Testing whether a transcript preserves the caller’s meaning may better show whether a voice agent can correctly handle important details than counting word mismatches.

## Sources & references

1. [mu-bench: A Multilingual Utterance Transcription Benchmark](https://arxiv.org/abs/2609.32082) – arXiv, 2026-09-29

Last updated: 2026-09-29
