
The short version
- Containment alone is misleading. Pair it with resolution, repeat calls and caller feedback.
- Track why calls transfer, how accurately key details are captured and how long callers wait for replies.
- Review transcripts every week, and change one thing at a time so you can see what helped.
Launching an AI voice agent is the start of the work, not the end. The first weeks tell you what callers really ask, where the agent struggles and what's worth fixing first. To learn that, you need the right measures and a routine for reviewing them.
Here's a practical scorecard for AI phone agents, and how to use it.
Containment, and why it's not enough
Containment rate is the share of calls the agent handles from start to finish without transferring to a person. It's the number most vendors lead with, and it matters: it's where cost savings come from.
But containment can look great while customers suffer. A caller who gives up and hangs up counts as "contained." So does a caller who gets a wrong answer and calls back tomorrow. Never report containment on its own. Pair it with the measures below.
Resolution
Did the caller get what they needed?
- Task completion: for structured tasks, did the booking, lookup or update succeed? This is usually measurable from system data.
- Repeat contacts: did the same caller contact you again about the same thing within a few days? Track this exactly as you would for people. See first call resolution.
- Abandonment: how many callers hung up mid-conversation, and at what point? Clusters of hang-ups at the same step point to a problem.
Transfers
Track transfer rate, and more importantly, transfer reasons:
- Caller asked for a person
- Out of scope by design
- Agent couldn't understand or complete the task
- Policy or verification failure
The first two are expected. The third is where improvement lives. Each week, look at the top reasons in that bucket. Often a few missing facts or permissions account for most of them. Our article on designing the handoff covers what a good transfer looks like.
Accuracy
- Information accuracy: were the answers correct? Sample transcripts and check the agent's statements against your policies and knowledge base.
- Capture accuracy: were names, numbers, dates and addresses recorded correctly? Compare a sample of captured details against the recordings. This depends heavily on speech recognition quality, discussed in word error rate explained.
- Action accuracy: were bookings made for the right time, the right service, the right customer?
A small number of serious errors (wrong appointment time, wrong account) matters more than a lot of minor wording issues.
Caller experience
- Response latency: the gap between the caller finishing and the agent replying, measured on real calls. Look at the slowest turns, not just the average. See voice agent latency.
- Turns per task: how many back-and-forths it takes to complete common tasks. Rising numbers suggest confusion.
- Repeats: how often the agent asks the caller to repeat themselves.
- Feedback: a short post-call survey ("Did we solve what you called about?"), or sentiment detected from transcripts.
Compliance
For every call, check that required items happened:
- Identification and AI disclosure at the start
- Recording notice, where applicable
- Verification before account details
- Opt-out requests recognized, confirmed and logged (critical for outbound)
- No prohibited statements
These can be checked automatically from transcripts on every call, as described in call center QA with transcripts. Any miss here deserves immediate attention.
Cost
- Cost per handled call for the agent
- Cost per resolved call, which accounts for repeats and transfers
- Human minutes saved on the call types the agent handles
- Review and maintenance time your team spends on the agent
See cost per call for building these into a simple model.
A weekly review routine
A routine that works for most teams, taking about an hour a week:
- Look at the dashboard: volume, containment, resolution, transfer reasons, latency, compliance misses.
- Read 20 to 30 transcripts: a random sample plus all compliance misses, all abandoned calls, and calls that transferred because the agent failed.
- List the problems: missing facts, unclear instructions, integration errors, recognition issues.
- Pick one or two fixes. Update the knowledge base or instructions.
- Test the fixes against saved conversations before publishing.
- Note what you changed, so next week you can see whether it helped.
Changing one thing at a time is slower, but it's the only way to know what worked.
What good looks like over time
In the first weeks, expect a long list of small fixes: questions you didn't anticipate, phrasings the agent misunderstands, edge cases in your rules. That's normal. Over the following months, the list should shrink, containment and resolution should rise together, and the transfer reasons should shift toward "caller asked" and "out of scope by design."
If containment rises but repeat calls or complaints rise too, the agent is closing calls it shouldn't. Tighten the rules on when to hand off.
Expanding the agent
Once the agent handles its first call types well, transcripts show you what to add next. Look for frequent out-of-scope reasons that are structured and low-risk. Add one new capability at a time, with the same measurement routine.
Example dashboard
A one-page weekly dashboard for an AI agent might include:
| Metric | This week | Last week | Target |
|---|---|---|---|
| Calls handled | 2,140 | 2,010 | |
| Containment | 64% | 61% | 65% |
| Task completion (bookings, lookups) | 91% | 90% | 90% |
| Repeat contact within 7 days | 9% | 11% | under 10% |
| Transfers: caller asked | 14% | 15% | |
| Transfers: agent failure | 6% | 8% | under 5% |
| Median reply latency | 0.9 s | 0.9 s | under 1 s |
| Slowest 10% of replies | 1.8 s | 2.2 s | under 2 s |
| Compliance misses | 2 | 5 | 0 |
The figures here are illustrative. What matters is the shape: a handful of metrics, trends against last week, and clear targets. Anything off target gets a transcript review that week.

Segment everything
Averages hide problems. Break your metrics down by:
- Call reason: an agent can be excellent at order status and poor at returns
- Time of day: late-night calls may come from different callers with different needs
- Channel or number: calls from a marketing campaign may behave differently from existing customers
- New vs returning callers
- Language, if the agent handles more than one
When a metric moves, the segment breakdown usually tells you why within minutes.
Running controlled changes
When you want to know whether a change actually helped, compare like with like. Options include:
- A/B testing: send a share of calls to the new version and the rest to the current one, then compare outcomes for the same call types.
- Before and after with care: compare the same weekdays and hours, and watch for outside changes like a marketing campaign or a holiday.
- Replay testing: run a fixed set of saved conversations through both versions before launch.
Small, measured changes add up. Large untested changes are how agents get worse without anyone noticing.
Reporting to leadership
Leaders usually want to know three things: is it working, is it safe, and what's next. A short monthly report can cover all three:
- Impact: calls handled, human time freed, after-hours calls captured, any revenue effects
- Quality and risk: resolution, repeat contacts, compliance results, notable incidents and what was done
- Plan: the top issues being fixed and the next capability to add
Include two or three short transcript excerpts, one good and one that shows a problem being fixed. Real examples build more confidence than charts alone.
Frequently asked questions
What's a good containment rate?
It depends entirely on the call types and the agent's scope. A narrowly scoped agent for a simple task should contain most calls. A general front-door agent will transfer more. Track your own trend alongside resolution.
How long until an agent is "done"?
Never quite. Your business changes, and the agent needs to keep up. The review effort drops a lot after the first couple of months.
Who should own the agent internally?
Someone in operations or customer service who understands the calls, with access to change the knowledge base and instructions. Without an owner, agents slowly go out of date.


