NewLondon office open. British English agents now live on UK numbers.See our offices
VoiFlow
All resources

Technology 6 min read

How AI voice agents work: from the first ring to the follow-up

The phone line, listening, deciding, acting in your systems, speaking and the follow-up, explained in plain words with the two main designs compared.

Engineering
VoiFlow

Close-up of a telephone handset's microphone grille and coiled cable, lit with soft blue light on a dark background

When a customer rings your business at nine at night, something has to answer, understand the request and keep the conversation moving. So how do AI voice agents work when a real person is waiting on the line? This article follows one call from the first ring to the written follow-up, explaining each stage in plain words so you can judge a vendor's answers.

An AI voice agent handles a call in a loop. It turns the caller's speech into text, works out what they want, checks the rules, acts in your systems if needed and speaks a reply, then listens again. Some designs do the speech steps inside one model, while others run them in sequence. Each choice trades speed against control.

What happens in the first second

The call reaches the agent in one of two ways. It may come to a local phone number that the agent answers, or it may arrive over a SIP connection (the standard way phone systems set up calls over the internet) from your existing phone system, such as a PBX (the private branch exchange that routes calls inside a business). Either way, the agent receives the caller's voice as a live stream of audio and sends its own voice back the same way.

Before anything else, the system must know when the caller is speaking. A component called voice activity detection, or VAD, separates speech from silence and background noise. It runs all the time, so the agent knows when to start listening and when a pause might mean the caller has finished.

Listening: turning speech into text

Speech recognition, also called automatic speech recognition or ASR, converts the audio into words. It works while the caller is still talking, so the text is nearly ready when they stop. This is the step most affected by accents, background noise, weak phone lines and names. A wrong word here travels through every later step.

For that reason, test this stage on the way your own callers speak. We explain how in speech recognition and accents.

How does the agent know what to do next?

The language model is the part most people picture when they think of AI. It is a model trained on large amounts of text to produce language. In the agent, it reads the conversation so far, works out what the caller wants and chooses the next move: answer a question, ask for a missing detail, check a record or pass the call to a person.

Good designs limit that choice. The model gets a short list of approved actions and the approved knowledge it may use. It should not make up a policy on the spot. If a caller asks about a refund, the agent checks the written refund policy and whether the amount sits within its limit.

What happens when the agent acts?

Understanding a request is only useful if the agent can do something with it. Actions happen through tools, which are small, defined connections to your calendar, customer records, order system or task list. A booking on a typical call runs like this:

  1. The caller asks for an appointment on Thursday.
  2. The agent checks the calendar for free slots.
  3. The caller picks a time.
  4. The agent confirms the details it needs and books the slot.
  5. The agent reads the time back and sends a written confirmation.

Before it reveals or changes anything, the agent should confirm who it is speaking to. Our guide to caller verification covers how light and strong checks differ.

How does the reply reach the caller?

The reply is written as text, then a text-to-speech (TTS) component turns it into speech. This is where the voice, pace and pronunciation of names are set. Good systems start speaking before the whole sentence is finished, which is one of the main ways to keep the wait short. The effect of delay on the call is covered in voice AI latency.

Turn-taking: knowing when to speak

Turn-taking is the timing layer. The agent has to decide when the caller has finished, when it may speak, and what to do when the caller cuts in. Get this wrong and the agent talks over people or waits too long. Barge-in and turn-taking explains the details.

Two designs: speech-to-speech and pipelines

There are two common ways to arrange these steps. A speech-to-speech model takes audio in and produces audio out inside one model. A pipeline runs separate components in order: speech recognition, then a language model, then speech synthesis.

Speech-to-speechPipeline
How a turn worksOne model takes audio in and produces audio outSeparate components run in sequence, each passing text to the next
StrengthsCan carry tone and timing into the reply, with fewer hand-offsYou can read the text at each stage, swap one part and check rules on the text
Trade-offsHarder to inspect what it heard, to insert a business rule mid-turn or to trace an errorEach hand-off adds time unless the stages overlap, and there are more parts to run
Better fit whenThe conversation is open-ended and tone matters mostActions, records and audit trails must be clear at every step

Neither design is automatically better. The right one depends on how much of the call must be checked line by line.

Where do the guardrails sit?

Guardrails are the rules that decide what the agent may do, whatever the model wants to say. They should sit outside the model, in code the business controls. That includes permissions, limits such as the largest refund the agent can approve, consent checks, redaction (removing sensitive numbers from stored records) and an audit trail, which is a log of every action. The model proposes. The rules decide. The platform's control layer is where those rules are set.

What happens after the call?

The call does not end when the caller hangs up. The agent writes a summary, saves the transcript and records the outcome: a booking, a changed address, or a promise to call back. Where a person must act, it creates a task. A confirmation can follow by WhatsApp, SMS or email, so the caller has a written record of what was agreed. Next time they call or message, the agent picks up the matter instead of starting again.

What is kept, and for how long, should be a rule you set, not a default you accept.

What should you ask a vendor?

  • Which design does the product use, speech-to-speech or a pipeline, and why?
  • Can I see a full call transcript that shows what was heard, what was decided and which action ran?
  • Where do the facts in a reply come from, and what stops the agent from guessing?
  • Which rules are enforced outside the model, and who can change them?
  • What happens when a tool fails in the middle of a call?

How VoiFlow handles this

VoiFlow's AI voice agents answer and make business calls end to end. They understand the caller, act in your systems and follow up on WhatsApp, SMS and email in the same conversation. Permissions, limits, consent, redaction and audit trail are enforced outside the model. Every call is recorded, transcribed, scored and replayable, and VoiFlow works with existing phone systems through SIP or PBX, or with new local numbers.

Frequently asked questions

Do AI voice agents work on an ordinary phone line?

Yes. They receive calls from the phone network or from a SIP connection to your phone system. Line quality still affects recognition, so test on the phone lines your customers actually use.

Is speech-to-speech better than a pipeline?

Neither is better in every case. Speech-to-speech can sound more natural, while a pipeline is easier to read, control and audit. Choose based on how much of each call you need to check.

Does the agent have to follow a word-for-word script?

Usually not. Good designs give the agent approved facts, rules and goals, and let it phrase its replies. Fixed wording suits only short steps, such as a required notice.

Can I see what an agent said and did on a call?

You should be able to. Look for a system that records the call, transcribes it and logs each action. You can hear a live agent for yourself before you decide.

Call it now · United States line+1 (831) 282-9226

Hear it run one of your calls.

Call the agent yourself, start a free trial, or bring us your hardest workflow. Most teams are live in days.