Case study · AI engineering

They stopped choosing which calls to listen to.
Life Insurance Direct has compared life cover for Australians since 2006, across nine insurers and six product types, and almost all of it happens on the phone. We built the system that reads every one of those calls, scores both sides, and hands the team a review queue in risk order instead of a sample.
What it changes: They used to choose which calls to hear, then listen. Now every call is read and scored first, and only then does a person decide what to open. The call that most needed a listen was almost never the one that looked worth picking.
- Client
- Life Insurance Direct
- Sector
- Life insurance distribution, Sydney
- Status
- Live in production
- What we built
- Metricsense, call quality and compliance QA
- Coverage
- Every call, scored on the customer side and the adviser side
- Deployment
- Their own cloud account, Australian region, personal information stripped before any transcript reaches a model
- Reads every call
- Both sides scored
- Risk ordered queue
- Plain English checks
- Pinned to the transcript
- Human in the loop
- In their own cloud
9,218
Calls read, both sides, 16 Apr to 10 Aug 2026
2 to 5%
What hand picking reached before
26
Checks on every call
0
Findings reported on its own
Why not just listen to more calls
The cheapest option was the one they were already running, and running properly. So the honest question was never whether hand picking works. It was what happens to the calls it never reaches.
Do the arithmetic
Around five hundred calls a week, so roughly two thousand a month. A careful manual process reaches two to five of every hundred. Reading the rest by hand is not one more reviewer. It is a department.
The lag is the other half
Feedback reached an adviser three to four weeks after the call it was about. By then they had taken hundreds more, the same way, with the same habit.
A sample only ranks what you suspect
Risk based selection is still selection. It sorts the calls a person thought to look at, so the one nobody thought to look at stays exactly where it was.
What the first round changed
The brief moved off whether an adviser might cross the compliance line, and onto whether they stop dead at it, say what they cannot do, and lose a customer who only wanted to understand what they were buying. That is the difference between compliance policing and compliant selling, and the client thought of it, not us. Our job was to notice it changed the product and rebuild the checks around it.
The hard part: the obvious build was the wrong one
Why that fails
The hardest check was a comparison against the insurer’s application. Hand a model a whole application form and attention lands on the front and the end. The middle gets skimmed, and on an insurance application the middle is where the medical questions sit, exactly the part that has to be checked word for word.
What we did instead
The check never sees the document. It sees one marked line range, with personal information already removed, and a single purpose built check runs against that section and nothing else. As the client’s QA lead put it: the less we give it, the less it makes up.
What it does now
Risk order, not date order
The queue is ranked. The highest risk calls get a person first, then the middle band, then the ones simply worth a look.
A new check is one sentence
Checks are written in plain English. No tagging exercise, no retraining. Deploy one and it runs across every call from then on.
Every score links to the moment
Open a score and it lands on the exact point in the transcript that produced it, with the audio playing alongside.
A quality index per adviser
A composite out of a hundred, across accuracy, satisfaction, resolution, compliance and communication. The spread inside one person is what tells you which axis to coach.
Where each person is on their journey
Not just how someone performs, but which areas they are weaker in than the floor. Onboarding and refreshers get aimed at the specific thing, not the whole team.
In their cloud, in their region
The whole thing runs inside their own account, in Australia, with personal information stripped before any transcript reaches a model.
The line that matters most
A flag is a candidate issue with evidence attached, not a finding. A person confirms it, and the team kept their own scorecards and their own judgement. Metricsense reports nothing to a regulator, and nothing to anyone else, on its own.
How we worked
- Concepts, then scripts
- Universal checks first, because they run on every call from day one. Call type scripts layer on top of something already working.
- Their QA became the checks
- Their own scripts and standard procedures came across marked three ways: said word for word, covered in your own words, or left to the adviser. Each item became one plain English instruction.
- They marked it
- Their QA lead reviewed calls beside the system, weekly, and told us where it was right and where it was not, with the reason. We tuned against real disagreements, not imagined ones.
- The library grows
- A review turns up something worth watching. It becomes a new check, written as a sentence rather than a retrain, and from then on it runs on every call. The loop closes in days.
What we left out on purpose
The version of the application check that runs automatically on every application call is not built. It stayed a pilot on two marked up forms. Everyone assumed there was a clean structured record to check the call against, and there was not: their customer record holds around ten fields where the application itself collects a hundred.
A half connected version would have looked like a feature and behaved like a guess, on the one check where a guess does the most damage. It stays a pilot until there is a way to do it we would defend in front of their regulator, not just in front of them.
“Before Metricsense, checking calls was a manual process: we selected files for review. Now the platform analyses every call, and we review by risk. It has identified customer experiences we would not have picked up before, and it lets us scale our quality assurance program without increasing headcount, with our team still the human in the loop. It helps us protect our customers, our business and our agents.”
The job transfers
Anyone who sells on the phone
Advice licensees
Mortgage brokers
Intermediaries
Life insurance distributors
Anyone audited on the call
Any regulated contact centre
Any team reviewing a sample by hand
More ai engineering work
OnlyFacts.ioAI engineering · Climate data · Australia
The climate agent that will not answer from memory
OnlyFacts publishes the Australian climate data newsrooms quote by name. We built the agent that answers readers in plain English, grounded in a live query and cited every time.
- 190 data tools, two agents
- What we built
- Live in production
- Status
Next
Tell us which part of this looks like your problem.
We will walk you through how it was built, who built it, and what it took to keep it running afterwards.