NewAgents that speak Manglish and Bahasa, now live on Malaysian numbers.Call it now
VoiFlow
All resources

Playbooks 5 min read

How to measure an AI voice agent: the metrics that matter

Track a short list of measures that show whether the agent works: calls answered, resolved, handed over and followed up, plus cost per resolved call.

Deployment team
VoiFlow

An operations manager in a glass-walled office studies printed charts pinned to a wall, arms folded in afternoon light

Switch on an AI voice agent and a month later someone asks the obvious question: is it working? Without agreed AI voice agent KPIs, the answer depends on who you ask. The sales lead counts leads, the operations manager counts queue length and finance looks at the invoice. You need one short list that everyone reads the same way.

Measure what happened on each call and what it cost: answer rate, resolution rate, handoff rate and quality, promises kept, customer satisfaction, response time and cost per resolved call. Calculate each one the same way every week from the call records, and set targets from your own baseline rather than from someone else's benchmark.

Start with a baseline, not a target

Before you can say whether the agent improved anything, you need to know how things worked before it. A baseline is one or two weeks of current performance, measured the same way you will measure the agent later. If you do not have call records, start a simple log now.

A baseline shows, for example, how many calls were missed in a busy hour and how many were about bookings. Those figures belong to your business. Do not copy them from anyone else.

Which AI voice agent KPIs matter most?

The table lists the contact centre AI metrics most teams need, with a simple way to calculate each one.

MeasureWhat it tells youHow to measure it
Answer rateWhether calls get picked up at allCalls answered divided by calls received, from the phone system
Resolution rate (first call resolution)Whether the caller's request was completedCalls with a recorded outcome divided by all calls, checked on a sample
Handoff rateHow often a person was needed, and whyHandoffs divided by calls, split by reason
Quality scoreHow well each call was handledAutomatic score against your rubric, checked by people on a sample
Promises keptWhether callbacks and follow-ups happened on timeTasks done before their due date divided by tasks created
Customer satisfactionHow callers felt about the outcomeSurvey responses divided by surveys sent, tracked as a trend
Response timeHow long before a caller is greeted and helpedTime from connection to greeting, and to the first useful answer
Cost per resolved callWhat each completed request costsRunning cost for the period divided by resolved calls

For handoffs, the warm handoff approach explains what the summary should contain. A clear summary makes the reason codes easier to set, and a reason code is what turns a handoff count into something you can act on.

Measure the time to the first useful answer, not only the greeting. A fast hello followed by a long wait for an answer feels slower than a slightly longer greeting that gets straight to the point.

For the cost measure, the ROI of AI call answering article sets out a simple model you can fill in with your own figures.

Why containment is a trap

Containment rate is the share of calls the agent finishes without a person. It is popular because it is easy to count and it climbs quickly. It is also easy to inflate without meaning to. An agent that never hands over will show high containment, while callers give up, call back or write a complaint.

Pair containment with resolution rate, repeat calls within a few days, and complaints. If containment rises while repeat calls and complaints rise too, the agent is keeping callers away from a person rather than helping them.

The vanity numbers to ignore

Some numbers look impressive in a demo and tell you very little about results:

  • Total minutes handled, which shows how long calls ran, not what they achieved.
  • The number of topics in the knowledge base, which shows effort, not results.
  • One reviewer's tone rating after a demo call.
  • Calls answered with no recorded outcome.
  • Sentiment scores with no reason behind them.

None of these is useless. They cannot tell you whether a caller got what they needed, so they should not sit at the top of your weekly review.

A simple weekly dashboard

Keep one page with seven numbers, and review it every week with the people who own the calls:

  1. Calls answered, and the answer rate.
  2. Resolution rate, with the sample check.
  3. Handoffs, and the top three reasons.
  4. Promises due, and promises kept.
  5. Complaints, and repeat calls within seven days.
  6. The satisfaction trend.
  7. Cost per resolved call.

Each week, read the calls behind any number that moved, not only the number itself. A handoff rate that rose on one day may be a single new rule that needs checking, not a trend. The QA for AI calls article shows how to choose which calls to read.

How do you set targets for the first 90 days?

Do not set targets in the first month. Use weeks one to four to measure the agent and to check that the measures are recorded correctly. In weeks five to eight, agree a target for each measure, moving it against the baseline by an amount the team thinks is realistic. In weeks nine to twelve, decide whether to widen the scope to another call type.

For example, a clinic taking 300 calls a day might find from its baseline that a large share of those calls are booking requests. The first target could be to resolve most booking calls without a transfer, while checking that the no-show rate does not rise. Set the target in the same units as the baseline, so the comparison is fair.

The 90-day review is the point to decide whether to keep, change or stop. The voice AI pilot in 30 days article sets out how to run that decision with clear success criteria.

How VoiFlow handles this

VoiFlow records, transcribes, scores and replays every call, and gives supervisors a live operations view. Promises made on a call become tracked tasks until they are done, so promises kept can be counted rather than estimated. Those records are the raw material for every measure above. The intelligence page describes the reporting in more detail.

Frequently asked questions

What is a good answer rate for an AI voice agent?

There is no single figure that applies to every business. Compare your own numbers: how many calls you missed before going live, and how many the agent answers now, measured the same way.

Is containment rate a useful KPI?

Only alongside other measures. A high containment rate can mean callers are being kept away from a person. Read it together with resolution rate, repeat calls and complaints.

How often should I review the numbers?

Look at the weekly dashboard every week, and do a deeper review each month, reading the calls behind any number that changed.

What should I measure in the first month?

Set up the records first: the answer count, the outcome of each call, the handoff reasons and the promises. Use the first month to build a baseline and check that the data is right, not to judge the agent.

To see these measures on a real call before you commit, try a call and ask for a person to see a handoff, or contact us to talk through your baseline.

Call it now · Malaysia line+60 3 2774 4929

Hear it run one of your calls.

Call the agent yourself, start a free trial, or bring us your hardest workflow. Most teams are live in days.