AI engineering

What is AI evaluation? A guide for teams putting AI into production

By Chirag••20 min read
A deep teal cover. On the left, the title "What is AI evaluation?" above the line "finding out whether it is right, on your own data, before a customer does". On the right, a run of five named checks with pass and fail verdicts, one of them failed, under the note "release blocked on one check"

AI evaluation is the practice of measuring whether an AI system produces the right output on your own data. It covers accuracy, grounding, safety, cost and what happens at the edges.

It applies to language models, retrieval systems and agents. Anywhere the same input can produce a different output twice.

Traditional software testing checks whether code does what it was written to do. The answer is the same every time you run it. AI evaluation checks whether a system that can answer differently on Tuesday is still answering well enough to ship.

Key takeaways

  • Evaluation exists to settle decisions. Do we ship this. Can we take the new model. Is this feature safe to widen. Do we need a guardrail here. A suite that answers none of those is a dashboard
  • An eval is an assertion you write after looking at how your system behaved on your own data. Not a benchmark, and not a single score
  • Benchmarks decide which model to start from. Evals decide whether your feature works. The first is a procurement question, the second is a release question
  • The measure that settles the decision is usually one you had to name yourself, not one that shipped with a framework. Nobody can sell you "did it route to the right team"
  • The bottleneck is judgement, not tooling. Reading traces and writing the rubric is the work, and the people who should do it are the ones who know what a right answer looks like
  • An eval suite outlives the model it was written for. That is what makes it the deliverable rather than a by-product

In this guide

  1. What is AI evaluation, and why does it matter now
  2. How is an eval different from a benchmark
  3. How is evaluating an agent different
  4. What should an eval actually measure
  5. How do you build your first eval set
  6. How are evals actually run
  7. Where does evaluation sit in the pipeline
  8. What does this look like in production
  9. What tools do teams use, and where each one fits
  10. What do most teams get wrong
  11. How often should evals run

1. What is AI evaluation, and why does it matter now

AI evaluation is how a team answers one question: is this thing right, and how would we know if it stopped being right.

The question is new because the failure is new. Software that breaks usually breaks loudly. It throws, it times out, it returns a 500. A language model that is wrong returns a fluent, well formatted, confident answer in the same shape as a correct one.

Nothing in your monitoring will flag it. The user may not flag it either.

This matters more now because of who is shipping. Most teams putting AI into production this year are not research teams. They are product teams who added a feature. The gap between a working demo and a system anyone will sign off is almost always evidence rather than engineering.

That gap is what evaluation closes.

The decisions it has to settle

An eval suite is not a quality ritual. It exists to let somebody say yes or no to a specific question, with a reason they can defend afterwards.

Four of those questions come up in every engagement:

  • Do we ship this? Is it right often enough, on the cases that matter, to put in front of a customer
  • Can we take the new model? A provider deprecates yours, or releases something cheaper. Can you move without finding out the hard way
  • Is it safe to widen? It works for one use case. Does it hold when you point it at the next one
  • Do we need a guardrail here? Which failures are rare enough to live with, and which need a check in front of them

If your evaluation cannot answer those, it is producing numbers rather than decisions. That distinction decides everything about how the suite gets built, which is the rest of this guide.

2. How is an eval different from a benchmark

A benchmark is generic and external. MMLU, HumanEval, SWE-bench, GAIA. Someone else chose the problems and your users are not in them.

An eval is specific to your product, your prompts, your documents and your edge cases.

The practical difference is who can write it. Anyone can run a benchmark. Nobody outside your team can write your evals, because nobody else has your traces.

A benchmark tells you which model is strong in general. An eval tells you whether your feature works for the people using it.

Both are useful. Only one of them is yours.

A worked example. A climate data agent we built answers public questions from a live database and cites a source every time. No benchmark on earth measures "did it answer from the database rather than from memory", because that failure mode only exists in that product.

That assertion had to be written by the people who knew the data.

3. How is evaluating an agent different

Most writing about evaluation assumes one input and one output. An agent is neither.

A model answers. An agent decides: which tool to call, with what arguments, in what order, when to stop, and what to do when a step fails. By the time it produces an answer it has made a dozen choices you did not see.

So the unit of evaluation changes. For a model you judge the output. For an agent you judge the path.

That matters because an agent can arrive at a correct final answer through a sequence no one would sanction. It called an expensive tool nine times. It looked up a record it had no business reading. It got there by luck after two failures it did not report. An eval that only reads the last message passes all three.

Four things worth measuring on an agent that have no equivalent in model evaluation:

  • Tool choice. Was the right tool called, with the right arguments
  • Recovery. When a step failed, did it notice and adapt, or carry on regardless
  • Termination. Did it stop when the job was done, rather than looping or quitting early
  • Cost of the path. The same answer can cost four calls or forty

There is a practical consequence for how you store traces. If you only log the final response, you cannot build any of this later. Capture the whole trajectory from the beginning, including the calls that failed.

4. What should an eval actually measure

There is no universal list. There is a list for your system, and you find it by looking at what your system got wrong.

That said, five dimensions recur in almost every production system we have built, and a sixth decides whether the system is any good.

4.1 Groundedness

Did the answer come from the source, or did the model recall it?

This is the dimension that matters most in regulated work. An investor assistant that produces a plausible return figure from memory rather than from the factsheet is not slightly wrong, it is a licence problem.

4.2 Accuracy against a known answer

For questions with one right answer, does the system give it?

This is the easiest to measure and the easiest to over-index on. Many of the questions worth asking do not have one right answer.

4.3 Task completion

For an agent, did the job actually get done?

A support agent that holds a polite conversation and logs the wrong ticket type has failed, and every text-level metric will say it did well.

4.4 Safety and refusal

Does the system decline what it should decline?

For an assistant under a financial services licence, that means no advice and no forecasts, however the question is phrased.

4.5 Cost and latency

What does a correct answer cost, and how long does the user wait?

A system that is right and takes nine seconds has failed a phone caller. These belong in the eval set, not in a separate performance ticket.

4.6 Whatever is specific to your task

The five above recur. The one that decides whether your system is worth having is usually not on anybody's list, including ours.

Real examples, each from a system we built:

  • Did the agent route to the right specialist team on the first pass, out of eight
  • Did the answer come from an approved document, rather than from any document
  • Can a reader work out which individual contributed a given piece of advice
  • Did the summary need correcting before the person acting on it could use it

None of those appear in an evaluation framework's defaults. Every one of them was the thing the client actually cared about.

This is also the dimension that is hardest to buy. A vendor can sell you faithfulness scoring. Nobody can sell you "did it route to the right team", because that metric did not exist until your system did.

If your eval set contains only the five dimensions above, you have bought someone else's checklist. The measure that matters most is usually the one you had to name yourself.

5. How do you build your first eval set

The order matters, and most teams do it backwards by starting with a tool.

5.1 Collect real traces

Not questions you imagined. Questions people actually asked, with what your system actually answered.

If you are pre-launch, run an internal pilot and use those. Synthetic questions produce synthetic confidence.

5.2 Read them and name the failures

Sit with the traces and sort them into groups. This answer was wrong because the retrieval missed. This one because the model guessed. This one because the question was ambiguous and we picked the wrong reading.

You are building a named list of the ways your system fails. This step is manual and it is the one that cannot be skipped.

5.3 Turn each named failure into an assertion

Now the tooling earns its place. Each failure mode becomes something that produces a number on every run: a code assertion, a model-as-judge prompt, or a human review rubric.

5.4 Have the right people mark it

This is the step teams get wrong most often, and it is the cheapest to get right.

The person who knows whether an answer is correct is usually not the engineer. For a fund assistant it is someone who knows the fund. For a support agent it is the specialist who picks up the ticket.

6. How are evals actually run

Four methods, and a production system usually needs all four. They differ in what they can judge and what they cost.

6.1 Code assertions

A deterministic check written in code. Is the citation present. Does the figure match the database. Is the JSON valid. Did the agent call the tool with the right arguments.

Fast, free, and completely reliable. If a check can be code, it should be code.

The limit is obvious: most interesting questions about a language model's output are not mechanically checkable.

6.2 A model judging against a rubric

A second model scores the output against a written standard. Commonly called LLM as judge.

This is how teams evaluate the things code cannot reach: was the answer faithful to the source, was the tone right, did it refuse when it should have.

A judge needs its own evaluation before you trust it. Score a set of outputs with the judge, have a human score the same set, and measure how often they agree. If the agreement is poor the judge is not measuring what you think, and tightening the rubric is the fix. Teams routinely skip this and then trust a number that was never checked.

6.3 Human review

A person marks the output, usually against the same rubric the judge uses.

Expensive and irreplaceable. Human marks are the ground truth that validates everything else, and they are where new failure modes get discovered. The volume needed is lower than teams expect, because the job is to calibrate rather than to cover.

6.4 Online evaluation on live traffic

Everything above runs offline, against a fixed set. Online evaluation measures the system as real users hit it.

That means sampling production traces for review, tracking implicit signals like whether the user rephrased or escalated, and A/B testing a change against the current behaviour.

Offline tells you whether you broke something you already knew about. Online tells you what you did not know to check.

Which to use where

MethodJudgesCostRuns
Code assertionsAnything mechanicalNothingEvery change
Model as judgeQuality, faithfulness, toneLowEvery change
Human reviewEverything, and the truthHighWeekly sample
OnlineWhat you did not anticipateMediumContinuous

A reasonable starting shape: code assertions for everything that can be one, a validated judge for the rest, a human reading a sample each week, and online measurement once there is enough traffic to read.

7. Where does evaluation sit in the pipeline

In the same place tests do, and for the same reason.

In application code the test is the specification. You write it before the code and it runs on every change. In an AI system the eval suite is the specification. You write it from real behaviour and it runs on every prompt, model or retrieval change.

That parity is the whole argument. A team that already practises test driven development and continuous delivery does not need a new discipline for AI. It needs the same discipline pointed at a different artefact.

Three practical consequences.

Evals run in CI, not in a notebook. If the suite only runs when someone remembers, it is documentation.

A failing eval blocks the release, exactly as a failing test does. Otherwise it is a dashboard.

The suite is version controlled with the prompts. A prompt change and the eval that covers it belong in the same commit.

There is a fourth consequence that takes longer to feel. An eval suite is the only artefact in an AI system that does not go stale when the model changes.

Prompts get rewritten. Models get deprecated. Retrieval gets rebuilt. The set of questions your users actually ask, and the judgement about which answers were right, survives all of it.

That is why we treat the suite as the deliverable rather than a by-product. The model is rented. The evaluation is owned.

8. What does this look like in production

Three systems we built, and what evaluation meant in each.

An investor assistant under a financial services licence

India Avenue's assistant answers questions on three funds from roughly 200 documents.

Accuracy was poor at first and we spent longer on the prompt than we should have. The prompt was not the problem. Much of the data sat inside images of charts and tables. Retrieval could not read it, so the agent answered from whatever words happened to sit nearby.

No tool surfaced that. We found it reading traces one at a time.

Their own team then sent fifty real investor questions and marked every answer right or wrong, with the reason. Those fifty are now the suite. Every change to the agent's instructions runs against them first. That is what lets the firm adopt a newer model and still show the answers did not change.

A voice agent on a support line

Talent Carriage's agent answers the phone and writes the ticket while the caller is still talking.

The measure here is not how many calls it handles. It is whether the specialist who picks up the ticket has to fix it. 70 of the first 71 needed no correction.

The measurement mechanism is the interesting part. The specialist flags a bad summary in one field, and that field is the accuracy number. The people doing the work are the eval.

A call quality system in life insurance

A system reading every sales call, both sides, against 26 checks.

The number worth noting is zero. It reports nothing on its own. Everything it finds is ranked and queued, and a person decides what to open.

That is an evaluation decision as much as a product one. A system that acts on its own findings needs to be right. A system that ranks findings for a human needs to be useful, which is a lower and more honest bar.

9. What tools do teams use, and where each one fits

Most people have heard of one or two of these and cannot place them. That is not ignorance, it is the market's fault: almost every tool does several jobs at once, and nobody sells "the second of four things you need".

There are four jobs. Work out which one you are buying for.

9.1 Capture, so there is something to look at

Recording what went in, what came back, what the retrieval returned, what the model cost and how long it took.

Without this you have nothing to read, and the reading is the work. If you adopt one thing before you have any evals, adopt this one.

Names you will meet: LangSmith, Langfuse, Arize Phoenix, Helicone, Weights and Biases Weave, Braintrust. OpenTelemetry conventions for AI now exist, so some teams send traces to the observability stack they already run.

9.2 Datasets and annotation, so judgement gets recorded

Somewhere to keep the examples you are evaluating against, and an interface for a human to mark them with a reason.

This is the job teams most often do in a spreadsheet, and the spreadsheet mostly works until two people disagree and there is no record of why.

Names: LangSmith has annotation queues, Argilla and Label Studio are dedicated labelling tools, Braintrust keeps datasets alongside results.

9.3 Running the assertions

The part that takes your named failure modes, runs them on every change, and produces numbers you can compare across versions.

Names: promptfoo is configuration driven and sits comfortably in CI, DeepEval is written like pytest, Ragas is specific to retrieval systems, Inspect comes from the UK AI Safety Institute, OpenAI Evals is a registry and framework, and EleutherAI's LM Evaluation Harness is the standard for benchmarking base models rather than products. LangSmith and Braintrust both include runners.

9.4 Judging

Grading an output against a rubric when no code can.

Almost nobody buys this separately. It is a feature of the platforms above, and the thing that matters is not which one you use but whether you validated the judge against human marks.

9.5 The platform you are already on

If your system runs on a hyperscaler, evaluation is probably already sitting in the console you have access to.

Amazon Bedrock has model and RAG evaluation built in. Databricks ships Mosaic AI Agent Evaluation. Google Vertex AI and Azure AI Foundry both carry their own evaluation services.

These are rarely the best of the four jobs individually. They are frequently the fastest route to a first number, and they involve no new vendor, no new contract and no data leaving an account you already approved. For a regulated client that last point decides it more often than feature comparisons do.

So where does LangSmith fit

This is the question we get most, so it is worth answering directly.

LangSmith is not one of the four. It is three of them in a bundle: capture, datasets and annotation, and a runner, with prompt versioning alongside. It came out of LangChain and works best if you already use LangChain or LangGraph, though the SDK traces plain code too.

That is why it is hard to place. You are trying to fit a suite into a category, and it is not a category.

The same is true of Braintrust and largely of Langfuse. promptfoo, Ragas and Argilla are the opposite: each does one job well and expects you to bring the rest.

Neither shape is better. Buying a suite means fewer decisions and more lock in. Buying parts means the reverse.

One warning worth more than any tool recommendation

The model is a setting, not an architecture. On the voice agent described earlier we tested cheaper and faster models and rejected them. They could not reliably work out the right next question.

Because model choice was a configuration value rather than something welded into the design, that decision can be revisited the week a better option lands. Build it so swapping the model is a setting, and your eval suite is what tells you whether the swap was safe.

A caveat on this whole section: it dates fast. The four jobs will not change. The names will.

10. What do most teams get wrong

Buying the tool first. The tool runs assertions. It cannot tell you what to assert, and that is the work.

Writing the evals themselves. Engineers write evals that test what the system does rather than what it should do. Those pass.

Measuring once. An eval run at launch is a launch report. An eval run on every change is a gate.

Treating a score as the output. A single number hides which failure mode got worse. Keep them named and keep them separate.

Skipping it because the demo worked. The demo is the most favourable sample that will ever be drawn.

11. How often should evals run

On every change that could affect behaviour. In practice that means four triggers.

Every prompt change, including the ones that look cosmetic. Reordering instructions changes outputs.

Every model change, including minor version bumps from the provider. This is the trigger teams forget, because it happens without a commit on their side.

Every retrieval change. New documents, a different chunking strategy, a reranker swap.

On a schedule, against live traffic. The offline suite tests what you knew about when you wrote it. A weekly sample of real production traces is how you find the failure modes that arrived since.

The first three are cheap and automatable. The fourth needs a person and it is the one that keeps the suite from going stale.

Where to start

If you have an AI system in production and no way to tell whether it is still right, start with the traces. Read a hundred of them with someone who knows the domain, name what went wrong in each, and count. The suite comes out of that, and nothing you buy will do it for you.

See how we build evaluation into delivery.

Frequently Asked Questions

What is AI evaluation?

AI evaluation is the practice of measuring whether an AI system produces the right output on your own data, before your users find out that it does not. It is made of assertions you write after looking at how your system actually behaved, run on every prompt, model or retrieval change.

Collapse

What is the difference between an eval and a benchmark?

Expand

How many examples do you need in an eval set?

Expand

Who should write the evals?

Expand

Can AI evaluation be automated?

Expand

How is evaluating an agent different from evaluating a model?

Expand

When should you start evaluating?

Expand

Share this article

Help others discover this content

TwitterLinkedIn
Categories:AI engineering