Voxvencer

Voice AI

Word Error Rate Explained: Measuring Speech-to-Text Accuracy

Word error rate (WER) is the standard way to measure speech-to-text accuracy. How it's calculated, a worked example, and why a low WER can still mislead you.

Studio microphone on a stand in front of recording equipment
Photo: Maciej Korsan, StockSnap (CC0)

The short version

  • WER = (substitutions + deletions + insertions) / words actually spoken.
  • Two systems with the same WER can perform very differently on the words that matter, like names and numbers.
  • Test on your own call audio. Vendor WER figures are usually measured on cleaner recordings than a phone line delivers.

When a speech-to-text vendor says their model is "95% accurate," they almost always mean the word error rate is around 5%. Word error rate, or WER, is the standard metric for transcription accuracy. It's simple to calculate, useful for comparisons, and easy to misread.

Here's how it works and how to use it to judge whether a model is good enough for phone calls.

The formula

WER counts the edits needed to turn the machine's transcript into the correct one, then divides by the number of words in the correct transcript.

WER = (S + D + I) / N

  • S (substitutions): words transcribed as a different word
  • D (deletions): words that were spoken but are missing from the transcript
  • I (insertions): words in the transcript that were never spoken
  • N: the number of words in the reference (the correct transcript)

The reference transcript is written by a person who listened carefully. Before comparing, both transcripts are usually normalized: lowercase, punctuation removed, numbers written consistently. Without that step, "2:15" versus "two fifteen" counts as an error even though nothing was misheard.

A worked example

Say the caller said:

I need to move my appointment to Thursday at two fifteen

That's 11 words, so N = 11.

The model produced:

I need to move my appointment Thursday at to fifteen please

Lining them up:

  • "to" before Thursday is missing: 1 deletion
  • "two" became "to": 1 substitution
  • "please" was added: 1 insertion

WER = (1 + 1 + 1) / 11 = 0.27, or 27%.

Notice that WER can go above 100% if a model inserts lots of words that were never said. It isn't a percentage of correct words, even though it's often described that way.

Laptop screen showing lines of code in a dark room
Photo: Markus Spiske, rawpixel (CC0)

Why WER is useful

It gives you one number to compare models, track changes over time and catch regressions. If you run the same test set through two speech engines and one scores 8% and the other 14%, the first is very likely better on that kind of audio. That's genuinely helpful when choosing a vendor or deciding whether a model update helped.

Why WER can mislead you

Not all words matter equally

In the example above, the model got "two fifteen" wrong. On a scheduling call, that's the only part that really mattered. A transcript can have a low WER and still miss the appointment time, the account number or the caller's name. Meanwhile a transcript that drops "um" and "like" gets penalized for words nobody cares about (unless the reference transcript removed them too).

For phone agents, it's worth measuring a second thing alongside WER: entity accuracy. Of all the names, numbers, dates, addresses and email addresses spoken, how many came through exactly right? This is usually the number that predicts whether calls succeed.

Test sets aren't your calls

Published WER figures are measured on specific datasets. Many are read speech or high-quality recordings. Phone audio is a different world: it's band-limited, compressed, full of background noise and often recorded on speakerphone. We explain why in why phone audio is hard for speech recognition. A model with a great benchmark score can lose a lot of accuracy on real call recordings.

Normalization choices change the number

Whether you strip filler words, how you treat numbers, and whether "gonna" counts as "going to" can shift WER by several points. When comparing vendors, use the same normalization for everyone, ideally your own.

Averages hide the bad calls

A 7% average WER might be 3% on most calls and 40% on the calls from cars or noisy job sites. Those bad calls are exactly the ones that end in frustrated customers. Look at the distribution, not only the average.

How to run your own test

You don't need a research team to get a useful answer. A practical approach:

  1. Collect 30 to 50 real call recordings. Include a mix: clear calls, speakerphone, accents your customers actually have, background noise. Get consent and handle recordings according to your privacy policy.
  2. Write reference transcripts. Have a careful person transcribe 1 to 2 minutes of each. Mark names, numbers and dates as entities.
  3. Normalize. Lowercase everything, strip punctuation, write numbers as words (or as digits, consistently).
  4. Run each recording through the models you're comparing. Normalize their output the same way.
  5. Calculate WER per call and overall. Many free tools and libraries compute WER. A spreadsheet works for small sets.
  6. Count entity errors separately. For each call, how many of the marked entities were wrong?

The results will tell you more than any spec sheet. You'll also learn where errors cluster, which often points to fixes like adding custom vocabulary for product names.

What's a "good" WER for phone calls?

There's no universal threshold, and anyone who quotes one without knowing your audio is guessing. A rough way to think about it:

  • If entities like names and numbers are coming through reliably, the agent can usually recover from small errors elsewhere by confirming key details.
  • If entity accuracy is poor, no amount of clever prompting downstream will save the call. The agent will book the wrong time or look up the wrong account.

This is why well-designed voice agents read back critical details ("that's Thursday at 2:15, correct?"). Confirmation turns a recognition error into a short correction instead of a failed call.

Try it on your own audio

You can upload your own recordings to our free speech-to-text tool and compare the transcripts with what was actually said. Ten files is enough to see how the model handles your callers.

Reading a transcript like an evaluator

Numbers are useful, but some of the most valuable information comes from reading the errors themselves. When you review transcripts against reference text, sort the mistakes into a few buckets:

Names and proper nouns. Customer surnames, street names, product names, staff names. These are often rare words the model hasn't seen much.

Numbers. Phone numbers, account numbers, dollar amounts, dates and times. Watch for transposed digits and confusions like fifteen and fifty.

Short function words. Dropped "to," "a" or "the." These inflate WER but rarely change meaning.

Homophones. "Their" and "there," "two" and "to." Usually harmless in a transcript, sometimes meaningful in a time or quantity.

Crosstalk and overlap. Stretches where both parties spoke at once, often garbled or missing.

Hesitations and restarts. "I want, um, I need to change the, the delivery address." Different systems handle these differently.

Once you've bucketed errors from a few dozen calls, the pattern usually points to a fix. Lots of name errors suggest custom vocabulary. Lots of number errors suggest the agent should confirm numbers back. Lots of crosstalk errors might mean the agent is talking over callers and needs better turn-taking.

A worked comparison

Here's a simple illustration of why WER and entity accuracy need to be read together. Imagine two speech models transcribing the same short call, where the caller said:

Hi, this is Maria Gonzalez, my account number is 7 4 2 9 1, and I'd like to move my delivery to Friday.

Model A produced:

Hi this is Maria Gonzales my account number is 7 4 2 9 1 and I'd like to move my delivery to Friday

Model B produced:

Hi it's Maria Gonzalez my account number is 7 4 2 1 9 and I like to move delivery Friday

Model A has one substitution (the spelling of the surname). Model B has several errors, including swapped digits in the account number. Both WER and entity accuracy favor Model A here, but you can easily construct cases where they disagree: a model that drops lots of small words but gets every digit right can have a worse WER and better real-world performance. That's why we recommend reporting both.

Improving accuracy after you've measured it

Measurement is only useful if it leads somewhere. The levers you typically have:

  • Choose a model trained on telephone audio if your calls are phone calls. This is usually the biggest single factor.
  • Supply custom vocabulary for names and terms specific to your business.
  • Improve the audio path: better codecs, fewer transcoding steps, less packet loss.
  • Design the conversation to confirm critical details, so recognition errors get corrected before they cause damage.
  • Ask for spelling with examples when names matter.
  • Re-test after every change, with the same reference set, so you know whether it actually helped.

Keeping your test set honest

A test set goes stale. Your callers change, your products change, and if you tune heavily against the same recordings you'll end up optimizing for them rather than for real traffic. Refresh the set every few months with new recordings, keep a portion of it hidden from whoever is tuning, and always include some deliberately difficult audio. The point of the test is to predict how real calls will go, not to produce a flattering number.

Frequently asked questions

Is 95% accuracy the same as 5% WER?

That's usually what vendors mean, but "accuracy" isn't strictly defined, and WER can exceed 100%, so 1 minus WER isn't a true accuracy figure. Ask how the number was measured and on what audio.

What's the difference between WER and CER?

Character error rate (CER) applies the same formula to characters instead of words. It's common for languages without clear word boundaries and is sometimes used for alphanumeric strings like order numbers.

Does punctuation count in WER?

Normally no. Punctuation and capitalization are removed before scoring, which is another reason a low WER doesn't guarantee a readable transcript.

Written by the Voxvencer editorial team. We build and run AI voice agents for call centers and small businesses, and we write about what we see on real phone lines. Questions or corrections: info@voxvencer.com.

Keep reading