NewGovernment-grade service desks, live in days. Hosted in the UAE.See how
VoiFlow
All resources

Playbooks 6 min read

QA for AI calls: review every conversation without listening to all of them

Score every AI call automatically, sample a few for human review, replay the ones that went wrong and fix the cause. A practical QA process for voice agents.

Deployment team
VoiFlow

A quality analyst in headphones with eyes closed, listening in a quiet office corner, a notebook open on the desk

An AI agent that takes a few hundred calls a day cannot be checked by listening to every call. A quality analyst who reviews ten calls a week will miss the pattern that matters: the same wrong price quoted on forty calls, or the same question handled badly every Tuesday. Call quality assurance AI means covering every call with a score, then spending human judgement where it counts.

Score every call automatically against a short rubric, send a sample to a person each week, and replay the calls that scored low or were handed over. Turn each repeated fault into a fix in the knowledge, the instructions or the rules, then check the next calls. Keep the scores and reviews as the record of how calls were checked.

Why call quality assurance AI needs to score every call

Many teams start with sampling: pick a handful of calls at random and listen. That is a fair start, but a random sample misses rare failures and is slow to show a trend. A fault that happens on one call in fifty may never appear in a sample of five, and it will still cost you customers.

Contact centre QA automation starts with scoring every call automatically, which closes that gap. The software reads each transcript against the same checks, so you can see which faults are common, which are new and which are getting worse. People then read the calls that matter most.

How do you score an AI call?

A rubric is a short list of areas. Each area has a plain description of a good call and a common fault. Keep it to five or six areas. Score each one as pass, partial or fail, and write one line of evidence, such as the moment in the call where it happened.

AreaA good callA common fault
Identity and permissionsChecks identity in proportion to the request, before sharing detailsShares an account detail before the check is complete
AccuracyAnswers from approved information and says when it does not knowGives a price or policy that is not in the knowledge base
Task completionCompletes the request, or books the follow-up and records itEnds the call with no outcome recorded
Handling and toneShort, clear and polite, with names and numbers repeated backLong replies to simple questions, or a mispronounced name
HandoffHands over when the rules say so, with a summaryHands over with no summary, so the caller repeats everything
Notes and recordThe summary and outcome are written to the recordThe outcome is missing or wrong

The table is a starting point. Change the rows to match your call types, and keep the wording plain enough that two reviewers would score the same call in the same way. Agree on what a good call looks like before the first review, not after the first disagreement.

What software can score automatically

AI call scoring is good at counting and matching. It can flag the following, and a person then judges what the flag means:

  • Whether the call reached a recorded outcome.
  • Whether a handoff happened, and what reason was given.
  • Whether a required notice was given, using the wording your rules specify.
  • Whether a prohibited phrase, topic or promise appeared in the agent's words.
  • Long silences, overlapping speech and repeated questions.
  • Calls that ran much longer than a typical call of the same type.
  • A caller who asked for a person and did not get one.

Treat these as signals, not verdicts. A long call can be a good call with a worried caller.

What should people still review?

Some judgements cannot be automated, and they are the ones customers feel most. Set aside time each week for these:

  • Every complaint, and every call where the caller asked for a person.
  • Every handoff, to check whether it should have happened sooner or later. The warm handoff approach explains what the summary should hold, which makes the review faster.
  • Calls with a low score, read in full rather than judged by the flag alone.
  • A random sample each week, so you can check that the automatic scores agree with people.
  • The first calls of any new call type, during its first two weeks.
  • Calls where the caller may be vulnerable, or upset in a way that affects what they can understand or agree to.

Reviewers judge fairness, tone and whether the agent should have transferred. A score cannot tell you that a caller was upset in a way that mattered to them. Post-call survey answers can add a second view of the same call.

Replay, coach and fix the cause

A low score is a symptom. The cause is usually one of a few things: missing knowledge, a weak instruction, an unclear rule, a speech recognition error or a delay in the call. Replay the call with its transcript and find out which one it is.

  1. Group the low scores by reason, not by call. Ten calls failing for the same reason is one problem.
  2. Replay the three most common cases in full, listening to the moments the transcript flags.
  3. Name the cause in one line, such as "the agent does not have the return window for gift cards".
  4. Fix it in the right place: the knowledge, the instructions, the rules or the speech settings. Fix the cause, not the single call.
  5. Check the next fifty calls of that type, and note whether the fault has come back.

If the fault is in what the agent heard, not what it said, the cause is in speech recognition. Speech recognition and accents explains how to build a test set from real calls, so the fix can be checked before it goes live.

Keep the evidence

If your business needs to show how calls are checked, the evidence is what you keep from your call review process. Keep the rubric and its version number, the scores, the human reviews with their dates, and the list of fixes with the date each went live. A record that shows the checking happened is more useful than a report that claims it did. Pair it with the measures in how to measure an AI voice agent, so the review shows results as well as activity.

Regulated sectors have their own expectations. For UK financial services, the FCA Consumer Duty and AI phone agents article sets out what firms need to be able to show, and it is worth reading before you design the review.

How VoiFlow handles this

VoiFlow records, transcribes and scores every call, makes each one replayable, and gives supervisors a live operations view. Reviewers work from the same recording and transcript, so when a fix goes in, the next calls of that type can be checked against the same rubric. The intelligence page describes the scoring and reporting in more detail.

Frequently asked questions

How many calls should a person review each week?

There is no fixed number. Review every complaint, every handoff and every low-scoring call, plus a random sample large enough to show whether the automatic scores agree with people. Adjust the sample size after the first month.

Can automatic scoring replace human review?

No. Automatic scoring finds the calls worth reading and counts patterns. People judge fairness, tone and whether a handoff should have happened. Use both.

How long should a QA rubric be?

Five or six areas, each with a plain description of a good call. Keep it short enough that reviewers use it on every call, not only on the calls they already suspect.

How do I know a fix worked?

Track the same score on the next calls of the same type. If the fault does not come back in a fair sample, the fix worked. If it returns, go back to the cause, not the call.

To see scoring and replay on a live call, try a call and ask for a person to see a handoff, or contact us about setting up review for your team.

Call it now · UAE line+971 6 558 2137

Hear it run one of your calls.

Call the agent yourself, start a free trial, or bring us your hardest workflow. Most teams are live in days.