NewGovernment-grade service desks, live in days. Hosted in the UAE.See how
VoiFlow
All resources

Technology 6 min read

Speech recognition and accents: how to test before you go live

Word error rate explained simply, a test set built from real calls with names, numbers and noise, and a check that the meaning survived the call.

Engineering
VoiFlow

A microphone with a pop filter in a small treated room, with a person in soft silhouette speaking into it

A voice agent can understand a caller from one city and mishear a caller from another. Speech recognition accents are one of the most common reasons an agent fails on live calls, and the failures are easy to miss until a caller hangs up. The fix starts before launch, with a test built from the people who will actually phone you.

To test speech recognition with accents, record real calls from the people who will phone you, with their accents, places and names. Measure word error rate, the share of words the system gets wrong, by group and call type. Then check whether the meaning survived, since one wrong word can change a booking, a name or a number.

What is word error rate, in plain words?

Word error rate, usually shortened to WER, is the share of words a speech recognition system gets wrong. It counts three kinds of mistake against a correct transcript:

  • a substitution, where one word is replaced by another
  • a deletion, where a spoken word is missed
  • an insertion, where a word appears that was never said

The formula is WER = (substitutions + deletions + insertions) / words in the correct transcript.

For example, a caller says ten words and the system gets one of them wrong. The WER is one divided by ten, which is 10 per cent. That figure is illustrative, not a benchmark. A low WER is useful, but it does not prove the agent understood the caller. The sections below explain why.

Do accents and dialects raise error rates?

The evidence is public. In a 2020 study of five commercial speech recognition systems, the average word error rate was 0.35 for Black speakers and 0.19 for white speakers (Koenecke and colleagues, 2020). The study grouped speakers by race, not by accent, and the authors traced most of the gap to the acoustic models, the parts of a system that map sound to the sounds of speech. They linked it to insufficient audio from Black speakers in training.

That study tested systems available in 2020, so its figures are not a guide to today's products. Its lesson still holds. A system that works well for one group of speakers can perform much worse for another, so measure each group you serve rather than trusting one overall score.

How do you build a test set from real calls?

A good test set reflects your callers, not a studio. Follow these steps.

  1. Pull around 200 recent calls across your main call types. Use only calls you are allowed to review under your own privacy rules.
  2. Ask colleagues which accents, dialects and mixed-language patterns your callers use, and include them. For Gulf dialects, see Arabic voice AI that understands the Gulf. For mixed Malay and English speech, see Manglish, Bahasa and the case for agents that sound local.
  3. Keep the phone audio as it arrived. Clean studio recordings hide the problems a phone line creates.
  4. Have a person transcribe each call word for word. This is the reference the machine is judged against.
  5. Tag each call by language, region or dialect where you already hold that information, and by call type. Do not guess a caller's background from their voice or their name.

Why do names, numbers and addresses need their own test?

Most words in a call matter less than a few specific items: names, phone numbers, dates, addresses and reference codes. One wrong digit in a booking, or a wrong letter in a surname, sends the wrong record to the wrong person. A low overall WER can hide these errors, so they need a separate check.

ItemWhy it mattersHow to test it
NamesA wrong name can point to the wrong customerInclude common names from your market and check each one is captured
Numbers and codesOne wrong digit breaks a booking or a referenceRead a set of numbers back on test calls and score them digit by digit
Dates and timesSimilar-sounding numbers blur on a poor lineTest pairs that sound alike, such as the thirteenth and the thirtieth
Addresses and placesLocal names may be missing from standard dictionariesUse your own street, building and area names, including the way people say them

Why measure meaning, not just words?

A low error rate can still change what the caller wanted. Suppose a caller says, "No, I don't want the morning slot." The system transcribes, "No, I want the morning slot." That is one missing word out of seven, roughly 14 per cent, and the booking now goes the wrong way. The words changed by one. The meaning reversed.

So score each call on the outcome as well as the words. The how AI voice agents work guide shows where recognition sits in each call, and the test should check the points where a wrong transcript turns into a wrong action.

  • Did the agent act on what the caller meant?
  • Did it read back names and numbers correctly?
  • Did it ask again when it was unsure, rather than guessing?

What about noise and phone audio?

Phone audio is harder than studio audio. The line compresses the voice, which blurs the sounds that tell similar words apart. Cafés, cars, open-plan offices and speakerphones add noise and echo on top of that.

Include these conditions in the test set, in roughly the same proportions you see on your own line. A test built from clean audio will score well and tell you very little about the calls you actually take.

How do you re-test after changes?

Accuracy moves whenever something changes. Build the re-test into your routine.

  1. Run the same test set after any change to the speech model, the vocabulary, the voice or the knowledge the agent uses.
  2. Compare the results group by group, not only the overall figure. A gain for one group can hide a loss for another.
  3. Add failures from live calls to the test set, so it grows with your callers.
  4. Keep a dated record of scores, so a drop after a change is easy to see.
  5. Run a fresh sample each month, because callers, products and place names change.

How VoiFlow handles this

VoiFlow supports 30+ languages and regional dialects, including Gulf Arabic, Bahasa Malaysia and Manglish. Each deployment should still be tested on its own callers before launch. Every call is recorded, transcribed, scored and replayable, so recent calls can be reviewed and added to a test set as your callers change. The conversation platform page sets out how the agent handles a conversation from start to finish.

Frequently asked questions

What is a good word error rate?

There is no single threshold that suits every call. Compare the rate across accent and dialect groups, and set a target for each group on the calls that matter most, such as bookings and payments.

How many calls do I need in a test set?

Start with around 200 real calls across your main call types, and add failures from live calls over time. The aim is coverage of the accents and call types you serve, not a single large number.

Can I test accents without recording real callers?

Only in part. Colleagues can read scripts in their own accents, which helps early on, but real callers on real lines show problems that staff recordings miss. Use real calls wherever you have permission to review them.

Is changing the speech system the only fix?

No. Adding names and local terms, and reading back key details, can fix some problems without a change of system. Test first, then decide. You can try a live call in the accent your customers use to hear the result for yourself.

Call it now · UAE line+971 6 558 2137

Hear it run one of your calls.

Call the agent yourself, start a free trial, or bring us your hardest workflow. Most teams are live in days.