NewGovernment-grade service desks, live in days. Hosted in the UAE.See how
VoiFlow
All resources

Technology 5 min read

Latency in voice AI: why half a second changes the call

The pause between a caller finishing and the agent replying: where the delay comes from, how callers feel it and how to measure it on real calls.

Engineering
VoiFlow

Long-exposure photograph of a highway interchange at night, with red and white light trails sweeping through the frame

A caller asks whether they can move their appointment to Thursday, then stops speaking. The silence that follows is where voice AI latency lives. Latency is the delay before the agent starts its reply, and it shapes how the whole call feels, even when the words themselves are right.

Voice AI latency is the delay between the moment a caller stops speaking and the moment the agent starts its reply. People answer each other quickly, so callers expect the same. Every stage of an agent adds to the delay, from detecting the end of speech to producing the first audible word. The caller only hears the total.

Why does a short pause feel long?

Human conversation is tightly timed. In a study of ten languages, speakers avoided overlapping talk and kept the silence between turns short. The mean gap between one person's turn and the next was 208 milliseconds across the full dataset (Stivers and colleagues, 2009).

Half a second is more than twice that gap. On a phone line, a pause of that length is easy to notice. Callers often fill it by speaking again, repeating the question or asking whether the line is still open. Each of those moments costs the call a little trust.

Where does the delay come from?

A reply passes through several stages, described in how AI voice agents work. Each one takes time, and some can overlap if the system is built to stream its work.

  • End-of-speech detection. The system waits to be sure the caller has finished. Waiting longer avoids cutting people off, but it adds delay.
  • Speech recognition. The audio becomes text. Long or noisy audio takes longer to convert.
  • Understanding and deciding. The language model reads the conversation and chooses a reply. Long instructions and long call histories slow it down.
  • Tool calls. A lookup in the calendar, order system or customer record adds a wait before the reply can be completed.
  • Speech output. The reply is turned into audio. The first audible words matter most, not the last.
  • Network. Audio travels from the caller's phone network to the agent and back. Distance and routing add time at both ends.

Perceived latency and measured latency

Measured latency is the clock time between two events. Perceived latency is how long the silence feels to the caller. The two are related but not the same.

A short acknowledgement such as "let me check that" can make a real lookup feel shorter. It does not make the lookup faster. A reply that starts with filler and then pauses again can feel worse than a plain answer. Use filler only when a real lookup is running, and track the time to the first real word.

How do you cut latency without breaking the call?

Most gains come from a few changes. Test each one on real calls before keeping it.

  • Stream everything. Start recognition and speech output before the full sentence is complete, so the reply can begin as soon as the first words are known.
  • Tune end-of-speech detection carefully. A setting that is fastest on paper may cut callers off. Repairing those calls costs more time than the saving.
  • Keep the prompt lean. The prompt is the set of instructions the model reads before each reply. Send the facts needed for this turn, not the whole history of the account.
  • Fetch likely data early. If a caller is booking, load the free slots before they ask for them.
  • Host close to the caller. Shorter network paths help. VoiFlow's approach to hosting is described further down.

The trade-off between a fast cut-off and a patient one is covered in more detail in barge-in and turn-taking.

A worked example with round numbers

The figures below are illustrative, not measured. Suppose an agent takes 150 milliseconds to decide the caller has finished, 100 milliseconds to transcribe the last words, 400 milliseconds for the language model to produce its first words, 200 milliseconds to produce the first audio and 100 milliseconds of network time. The caller waits 950 milliseconds.

Now suppose streaming and a tuned end-of-speech setting cut the decision wait to 100 milliseconds and the model time to 250 milliseconds. The total falls to 750 milliseconds. That is a real gain, but it is still more than three times the 208-millisecond gap that people use in conversation.

How do you measure latency on real calls?

Measure on live phone calls, not only in a test environment. A test harness cannot show how the phone network behaves.

  1. Log four timestamps for every turn: when the caller stops speaking, when the first words are recognised, when the reply's first words are generated, and when its first audio leaves the system.
  2. Track the gap from the caller stopping to the first audio leaving the system. That is the figure most callers feel.
  3. Report the median and the slowest tenth of replies. The slow replies are the ones callers remember.
  4. Review the figures weekly by call type, because a single average hides the slow cases. The wider set of measures is in voice agent KPIs.

Do not measure the model on its own. A fast model still produces a slow call if the audio path or a tool call is slow, so time the whole turn from the caller's side.

StageWhat to logWarning sign
End-of-speech detectionTime from the caller stopping to the system deciding they have finishedCallers talk over the agent or get cut off
Recognition and reasoningTime from end of speech to the first reply wordsSilence grows on longer or multi-step questions
Speech outputTime from the first reply words to the first audioReplies start late even when the words are ready

How VoiFlow handles this

VoiFlow is hosted in-region in the markets it serves, which keeps the audio path close to the caller. Every call is recorded, transcribed, scored and replayable, so the timing of each turn can be checked against the real call rather than a demo. The conversation platform page describes how the agent handles a live conversation. If you compare vendors, ask for latency measured on live phone calls in your market, with the method explained.

Frequently asked questions

What is a good latency for a voice AI agent?

There is no single figure that suits every call type. Set a target from real calls, and watch how often callers talk over the reply or ask whether the line is still open.

Does faster always mean better?

No. A faster reply that cuts callers off, or answers before the agent has understood the question, makes the call worse. Balance speed against accurate turn-taking and answer quality.

Can filler words such as "let me check" hide latency?

They can make a real lookup feel shorter, but they do not reduce the time it takes. Use them only when a tool call is genuinely running, and keep the time to the first real word on your dashboard.

How do I test latency before I buy?

Ask for a live call in your market and time the pause yourself with a stopwatch. Then ask the vendor for its own measurements on real phone lines. You can try a live call now, or talk to us about your phone setup.

Call it now · UAE line+971 6 558 2137

Hear it run one of your calls.

Call the agent yourself, start a free trial, or bring us your hardest workflow. Most teams are live in days.