Home / Guides / Call quality assurance
Method · 8 min read
Call quality assurance when you can only listen to 2% of calls
Almost every call quality assurance programme in the country works the same way: a team leader listens to a few calls per person per month and scores them against a form. It is better than nothing. It is also, statistically, close to nothing.
Start with the arithmetic
Take a modest operation: fifteen people on the phones, twenty-five calls each per day, twenty working days a month. That is 7,500 recorded calls a month.
A diligent QA programme reviews perhaps five calls per person per month. That is seventy-five calls. One per cent.
Now ask the question that matters: if two of those 7,500 calls contained a customer quietly deciding to leave, what is the chance your QA programme found either of them?
The number is bad, but the shape of the problem is worse than the number. Sampling assumes the thing you are looking for is evenly distributed. Service failures are not. They cluster around a broken process, a particular supplier, a system that fails on Tuesdays. A random one per cent sample is exactly the wrong instrument for finding something that clusters.
What the sample is actually good at
This is not an argument that QA sampling is worthless. It is good at things full-coverage review is bad at:
- Coaching individuals. A team leader listening to a real call with a real person is how skill actually transfers.
- Calibration. Getting managers to agree on what "good" sounds like requires them to listen to the same call and argue about it.
- Tone and manner. Whether someone sounded warm is a human judgement.
What it is bad at is finding the calls that matter in the first place. Those are two different jobs, and most QA programmes only have a tool for the second one.
The failure QA scoring is structurally blind to
Score a call on a standard QA form - greeting, verification, empathy, ownership, close - and a call can pass every line and still be a disaster.
The agent greeted the customer properly. Verified them. Sounded genuinely warm. Took ownership. Closed politely. And the thing the customer needed did not happen, because a system failed after the call ended, or an approval never came through, or the promised callback went to nobody.
That call scores well. It should score well - the agent did their job. The business failed. QA forms measure agents, so they cannot see it.
This is the same blindness that makes sentiment scoring miss the worst calls. In a measured 30-day review at an Australian corporate travel agency, the calls that were eventually raised as genuine service failures had a median sentiment score of 82 out of 100. Only three of forty-five scored below forty. They sounded fine, because they were fine, right up until the part where nothing happened.
Full coverage changes what you can ask
When every call is read rather than one in a hundred, the useful questions change. You stop asking "how did Sarah do on these five calls" and start asking:
- How many customers called back about something we had already been told about?
- Which supplier or system generates the most repeat contact?
- Is this one bad week or a pattern that has been running for four months?
- Where did a customer say something that should have escalated and did not?
These are operational questions, not people questions, and they are the ones that actually change outcomes. Coaching one agent fixes one agent. Finding the workflow that fails silently fixes everybody.
What full coverage found in practice
In a later month at that same agency, more than 5,600 calls were reviewed automatically and 114 were raised for a manager - about two per cent. The breakdown of what those 114 were actually about is the part worth sitting with:
| What the alert was about | Share |
|---|---|
| Technology or system failure affecting the customer | 67% |
| Customer complaint raised on the call | 13% |
| How the call itself was handled | 12% |
| Booking identified as at risk | 7% |
| Escalation requested | 1% |
Roughly two thirds pointed at systems and processes rather than at people. A QA programme built entirely around scoring agents would have found almost none of it - not because the reviewers were not good at their job, but because they were looking at the wrong unit of analysis.
Designing the two-tier programme
The version that works keeps both instruments and gives each the job it is good at:
- Automated review over everything, surfacing the small number of calls where the customer was left worse off. This is detection.
- Human review on what surfaces, plus a continuing random sample for coaching and calibration. This is judgement.
- A route for the systemic findings that is not the coaching route. If two thirds of what you find is process and technology, sending it all to a team leader as coaching material guarantees it never gets fixed.
That last point is where most programmes fall over. The alert arrives, the manager reads it, and the only lever they have is a conversation with the agent - who did nothing wrong. Give the operational findings somewhere else to go.
What to measure
Coverage-based review makes a handful of metrics available that sampling cannot produce honestly:
- Repeat contact rate on the same issue - the single most useful number most businesses do not have.
- Time to resolution measured from first contact rather than from ticket creation.
- Share of raised issues actually closed out, which measures the programme rather than the team.
- Cause mix - people versus process versus technology - which tells you where to spend.
Resist the temptation to turn any of these into a per-agent league table. The moment a coverage-based system is used for ranking, the team learns what makes it trigger, and you lose the visibility that made it worth having.
Common questions
What percentage of calls should we review?
For coaching, a small sample per person per month is enough and always has been. For detecting service failures, any sample is the wrong tool, because failures cluster rather than distributing evenly. The practical answer is to review everything automatically for detection and keep a human sample for coaching.
Does automated QA replace team leaders?
No. It replaces the search, not the judgement. Deciding whether an issue is genuine, whether the customer needs contact and whether anything should change is human work. Finding the twenty calls worth looking at out of seven thousand is not.
Will staff see this as surveillance?
That depends entirely on how it is introduced and what it is used for. A system that surfaces broken processes and is explicitly not used for individual scoring reads very differently from one that produces a ranking. In the measured example above, 67% of what was found was technology failure, not agent behaviour - which is usually the most persuasive thing you can tell a team.
See what a full review of your calls turns up
WiseSentry reads every recorded call and raises only the ones where your business let the customer down. Talk to us about a review of your own call history.