An AI voice or chat agent goes live and the reports start arriving. Calls attempted, calls answered, conversations handled, a satisfaction score. Everything looks busy. Three weeks later the operations head asks a simple question, which is whether it is actually working, and nobody can answer it from the dashboard in front of them.
The problem is not a shortage of numbers. It is that measuring AI agent performance needs a small set of metrics that connect to a business outcome, read together rather than separately, and checked against real transcripts every week.
This post sets out those metrics, how to sample and score conversations, and the one-page report a manager should get every Monday.
Start with the funnel, not the dashboard
Every voice or chat agent runs the same funnel, and each stage has a different failure mode. Write yours down before you choose metrics.
For voice: attempted, connected, conversation held, outcome reached, hand-over. For chat: message received, conversation started, outcome reached, hand-over, abandoned. A number only means something once you know which stage it belongs to. A poor result at “connected” is a data or timing problem. A poor result at “outcome” is a script or model problem. Treating both as “the AI is not working” wastes a month.
Measuring AI agent performance: the metrics that matter
Contact rate
Of the numbers you attempted, how many produced a real conversation with the right person. This is mostly a data quality and timing metric, not an AI metric. Wrong numbers, switched-off phones and calls placed at the wrong hour all land here.
Timing is partly set for you. Under TRAI’s commercial communication framework, four time bands are default off for all customers, covering 00:00 to 10:00 and 21:00 to 24:00, which in practice leaves the middle of the day open unless a customer has set something different. For lenders the Reserve Bank’s 2022 circular is tighter still: recovery agents must not call a borrower before 8:00 a.m. or after 7:00 p.m.
Containment
Of the conversations that started, how many the agent completed without a person. Read alone, containment rewards an agent that refuses to transfer, so always pair it with repeat contact within 72 hours and with complaints. A contained conversation the customer has to repeat tomorrow is a failure recorded as a success.
Hand-over quality
This is the metric most teams skip and the one customers feel most. Three parts: how quickly the transfer happens after the customer asks or the rule triggers, whether the human receives the full context, and whether the customer had to repeat themselves. Measure all three, because a fast transfer with no context is not a good transfer.
Outcome rate
The only metric the business actually asked for. It differs by process: promises to pay captured, appointments confirmed, documents collected, queries resolved. Every other number on this page exists to explain movement in this one.
| Metric | What it counts | What a bad number usually means |
|---|---|---|
| Contact rate | Real conversations per attempt | Stale numbers, wrong calling window, poor list hygiene |
| Containment | Conversations closed without a person | Knowledge gaps, or an agent refusing to transfer |
| Repeat contact in 72 hours | Same customer coming back | Answers that sound complete but are not |
| Time to human | Seconds from the ask to the transfer | Loops in the flow, or no clear exit route |
| Context carried over | Transfers arriving with transcript and record | Broken integration between agent and helpdesk |
| Outcome rate | Promises, bookings, resolutions | Script problem, or the wrong process for automation |
| Complaints and opt-outs | Customers objecting or withdrawing | Wrong list, wrong tone, or wrong hour |
How to sample and review calls
Dashboards tell you where to look. Only a transcript tells you why. Build sampling into the week rather than doing it when something goes wrong.
- Fix the sample size and keep it fixed. Twenty conversations per language per week is enough to spot patterns and small enough that a supervisor will actually do it.
- Stratify the sample. Take some at random, some from hand-overs, some from the shortest conversations and some from the longest. Random-only sampling hides both failure modes.
- Always include every complaint and every opt-out. These are not a sample, they are the whole population of things that went wrong.
- Review end to end. Listen or read to the last turn. Most failures happen after the point where a reviewer would have stopped.
- Score against a written sheet, not an impression. The sheet is below.
- Turn every failure into a line in the backlog with the exact wording that should have been used. Vague feedback such as “sounded robotic” changes nothing.
A quality score a supervisor can apply
Keep it to a page with yes or no answers, so two reviewers reach the same score.
- Did the agent identify itself and the organisation at the start, and say the conversation is automated and recorded?
- Was the information given correct against the source system?
- Did the agent avoid stating anything it could not verify?
- Did it follow the hand-over rule where one applied?
- Did it reach a person on the first request, with no loop?
- Was the language and register right for the customer, not a translated circular?
- Was the outcome recorded correctly in the system afterwards?
- Would you be comfortable if this recording were played back to the customer?
The last question catches more than the other seven together.
The weekly report a manager should get
One page, same shape every week, so trends are visible without analysis.
- Volume by stage of the funnel, this week against last week and against the pre-launch baseline.
- Contact rate, containment, repeat contact and outcome rate, each as a trend line rather than a single figure.
- Hand-over reasons, ranked. This is the build backlog.
- Quality score from the sample, with the number of conversations reviewed stated.
- Complaints and opt-outs, listed individually, not summarised.
- Three transcripts: the best, the worst and one typical.
- What changed in the configuration this week, and by whom.
That last line matters more than it looks. Without a change log you cannot tell an improvement from a coincidence. It is the same discipline the NIST AI Risk Management Framework describes under its Measure function: measure, document, and keep measuring after deployment.
Numbers that mislead
Average handling time on its own rewards rushing. Read it beside outcome rate or not at all.
Satisfaction scores collected only from completed conversations quietly exclude everyone who gave up, which is the group you most need to hear from.
Accuracy quoted by a vendor from their own test set tells you nothing about your documents, your accents or your customers. Ask for accuracy measured on your data, during your pilot.
And a containment rate that rises while complaints also rise is not an improvement. It is an agent that has learnt to hold on to people.
Consent and records
Recordings and transcripts are personal data. Under the Digital Personal Data Protection Act, 2023 the notice must state what is collected and why, consent must be specific and unambiguous, and withdrawal must be as easy as giving consent. Say at the start of the call that it is automated and recorded, keep the retention period written down, and make sure an opt-out actually stops future calls rather than just closing one ticket.
Frequently asked questions
What is a good containment rate?
There is no universal figure, and anyone quoting one without knowing your question mix is guessing. Compare against your own baseline from before launch, by question type. That comparison is the only one that means anything.
How soon should we expect numbers to settle?
Give it a few weeks of steady running. Early weeks contain configuration changes, so read trends rather than daily figures, and avoid changing the script and the list in the same week.
Who should own these metrics?
The manager who owned the process before it was automated. If measurement sits with IT, the report measures software. If it sits with operations, it measures the business.
Do we need a separate tool for this?
Usually not at the start. What you need is every outcome logged, transcripts retrievable, and a change log. A spreadsheet built from those three will outperform a dashboard built on none of them.
Where to start
Pick four metrics, one sample size and one weekly slot, and hold them steady for a month. The discipline matters more than the tooling, and it applies to any agent you deploy, whether it answers calls or works through a process end to end, as our note on agentic AI beyond the chatbot describes.
The AI voice calling in AI Solutions by AIMatric logs every outcome and supports human takeover, which is the minimum a measurement routine like this needs.
Sources
- Telecom Commercial Communications Customer Preference Regulations, 2018, TRAI (Schedule II time bands)
- Outsourcing of Financial Services: Responsibilities of regulated entities employing Recovery Agents, RBI, 12 August 2022
- The Digital Personal Data Protection Act, 2023, Ministry of Electronics and Information Technology
- AI Risk Management Framework, National Institute of Standards and Technology
