What we build

Six services, each one live for a named client.

What each one is, what it takes, and what you keep. Every service has its own page with the method, the deliverables and the questions buyers ask.

  1. Our people are already using ChatGPT. Nobody has said what they may put in it.

    AI literacy and safe adoption

    Your team is already using AI. This decides what they may put in it.

    Most people in your business have pasted company information into a chat tool they signed up for themselves. Two days of prompt training does not fix that. It makes it faster. We start by finding what is already in use, and draw the line on what may go where. Then two days teaching your team to work inside it on their own operations. Then we stay while they use it on real work. You leave with a policy, a shared set of skills that people actually run, and a written call on what to build properly.

    Length
    Six to eight weeks
    Format
    Two days in a room, the rest on real work
    Group size
    Twelve to thirteen a batch, hands on
    We step back
    When the skills are used without us
  2. We have a dozen AI ideas, and the budget to build one.

    AI discovery and value mapping

    Find the use case worth building, before anyone writes code.

    Most AI budgets are spent on the first idea someone had in a meeting. We map the jobs your business needs done, weigh each against what it would take and what it would pay, and leave you with a ranked list and one pilot you can defend to your board.

    Length
    Two weeks to start
    Your time
    About six hours, from two or three people
    You leave with
    A ranked list and a priced pilot
    Afterwards
    Build it with us or with anyone else
  3. Our chatbot answers questions. We need something that does the work.

    AI agent development

    Agents that take real actions in your systems, not only chat.

    An agent here opens a ticket, checks a record, routes to the right team, or chases a deadline. It plans, calls your tools, recovers when a step fails, and hands over to a person where a wrong action would cost something. Every action is logged.

    First agent
    Under a month, for a narrow pilot
    Model
    Chosen on your evaluation set
    Runs in
    Your cloud account
    You own
    Code, prompts, evaluations, documentation
  4. It sounds sure of itself, and nobody can tell where the answer came from.

    RAG and knowledge systems

    Answers from your own documents and data, with the source shown every time.

    A retrieval system that pulls the figure from the document that actually contains it, cites it, and says what it does not hold rather than guessing. Built for your material, not pointed at a folder.

    Starts with
    A corpus audit
    Answers
    Cited, or declined
    Runs in
    Your account and region
    Your data
    Never trains a shared model
  5. Our callers switch from Hindi to English in the middle of a sentence.

    Voice and conversational AI

    Agents that take the call, and systems that read every call your people take.

    Two things, on one discipline. A voice agent that answers the line at any hour, follows a caller who switches from Hindi to English mid sentence, and has the ticket written before they hang up. And a call analytics system that reads every call your team takes, scores both sides, and hands them a review queue in risk order instead of a sample.

    Languages
    Hindi, English and Hinglish in production
    Channels
    Phone, WhatsApp and chat
    Recordings
    In your cloud and region
    Fallback
    A person, whenever the caller asks
  6. Every time the model changes, we hold our breath.

    Evaluation, observability and guardrails

    Look at your data first. Then measure what you found.

    Most AI evaluation starts with a dashboard of generic scores and ends with nobody trusting them. We start by reading your real traces with someone from your side who knows the domain, name the ways the system actually fails, and only then build the checks. The result is an evaluation suite you own, that runs before every change and grows from production.

    Starts with
    About a hundred traces, read
    Your time
    One domain expert, a few hours a week in month one
    Verdicts
    Pass or fail, with a written reason
    Existing systems
    First findings usually within a week
  7. We check two calls in a hundred and hope the rest are fine.

    Data analytics, structured and unstructured

    Your dashboard tells you what happened. Your customers and your documents already said why.

    Most of what a business knows is in the calls, chats, tickets, reviews, survey answers and scanned documents nobody has time to read. We build systems that read all of it, score it against rules your own people write in plain English, and pin every finding to the exact moment or line that produced it.

    Coverage
    Every conversation, not a sample
    Checks
    Anything you can describe in a sentence
    Every flag
    Confirmed by a person
    Your data
    Your cloud, personal information stripped first

How an engagement starts

Small, on your own material, judged by your team.

Every AI engagement starts the same way, whichever service it is.

  1. 01

    A narrow pilot, under a month

    One question your data should already answer, on your own material. Built to be judged before it is trusted.

  2. 02

    You mark it

    Your team scores the answers right or wrong, with the reason. That marked set becomes the evaluation suite.

  3. 03

    Then it ships

    Nothing goes live without passing the suite. It runs in your cloud and your region, and you own the code, the evaluations and the documentation.

The evaluation suite is the deliverable that outlives the model. It is why our clients can move to a newer, cheaper model the week it lands and show the answers did not change.

The stack, and why we know it

Model agnostic by design.

Our Labs team built its own agent platform, AgentOS, from the ground up, so we know every layer of this stack and what breaks at each one. The model is a setting, not a rebuild.

  • Models

    • Claude (Anthropic)
    • GPT (OpenAI)
    • Gemini (Google)
    • Open source, self hosted
  • Agents and retrieval

    • Claude Agent SDK
    • LangGraph
    • MCP servers and tools
    • pgvector, Pinecone, Qdrant
  • Voice and channels

    • LiveKit
    • Plivo
    • Sarvam
    • WhatsApp Cloud API
  • Evaluation and observability

    • Client marked evaluation suites
    • LangSmith, Langfuse
    • Cost and latency tracing
    • Your cloud, your region

What clients say

In their words.

01 / 04

Life Insurance Direct
Before Metricsense, checking calls was a manual process: we selected files for review. Now the platform analyses every call, and we review by risk. It has identified customer experiences we would not have picked up before, and it lets us scale our quality assurance program without increasing headcount, with our team still the human in the loop. It helps us protect our customers, our business and our agents.

Russell Cain

Founder and CEO, Life Insurance Direct

Questions

What buyers ask us first.

AI literacy and safe adoption

Why not just run the two days?
Because a room full of people who enjoyed a workshop is not adoption. Everyone leaves keen. Within a fortnight most are back to the old way, because the first time the AI got something wrong there was nobody to ask. The weeks either side are what turn it into something your team still uses in month three. If you only want the two days, say so on the call. It is not what we recommend, and it is not what this is.More on AI literacy and safe adoption

AI discovery and value mapping

What if you find nothing worth building?
Then we say so. It happens, and it is a cheaper way to learn it than a failed pilot.More on AI discovery and value mapping

AI agent development

Can the agent do damage?
Only what its policy allows. Anything that costs money or touches a customer goes through a person until the evaluation record says it can be trusted, and every action is logged either way.More on AI agent development

RAG and knowledge systems

Will it hallucinate?
It answers only from the retrieved material and cites it. When the material is not there, it says so. The evaluation set is how we prove that, on your questions.More on RAG and knowledge systems

Voice and conversational AI

Which languages?
Hindi, English and Hinglish in production today, with Sarvam for speech. Other languages depend on the speech models available for them, and we test before we promise.More on voice and conversational AI

Evaluation, observability and guardrails

Can we not just use an evaluation dashboard?
Generic scores like helpfulness or coherence measure nothing about the ways your system fails. They look reassuring and they move for reasons nobody can explain. The checks that matter are the ones that came out of reading your own traces, which is why we start there.More on evaluation, observability and guardrails

Data analytics, structured and unstructured

Does this mean a machine grades our people?
No. It reads everything, attaches the evidence, and hands your team a shortlist. A person confirms every flag, and your team keeps its own scorecards and its own judgement. The right measure is not how many flags it raises but how many your team agrees with once they open the evidence.More on data analytics, structured and unstructured

Start with one question

Tell us the question your data should already be able to answer.

Thirty minutes with one of our architects. We will tell you whether it is a pilot, and what it would take.

See product development