AI engineering

23 articles in AI engineering

A factsheet page with a bar chart drawn as an image, a magnifying glass over it finding no text, and beside it three separate tool boxes labelled by document type feeding one cited answer.

Before you touch the prompt, check what reached the model

When an AI assistant answers wrongly, the first instinct is to rewrite the instructions. Half the time the instructions are fine and the model never saw the figure at all. Here is how retrieval works, the failure that looks like a model failure, how to tell them apart, and the funds firm where the returns turned out to be pictures.

By Chirag••6 min read
A question on the left, a wall of 190 small tool tiles in the middle, and a single cited answer on the right, with the open web crossed out below.

Give the model buttons, not a database

Ask an AI model for a figure it does not have and it will give you one anyway, fluently, and wrong. If your name goes on the numbers, that is not a risk you can prompt away. Here is the pattern we use so an assistant cannot invent a figure, explained from the start, and the climate data publisher it runs for today.

By Chirag••6 min read
A pyramid of three evaluation levels — unit tests at the base, human and model eval in the middle, A/B testing at the top — beside three cross-cutting axes: offline to online, reference-based to reference-free, end-to-end to per-component

You Don't Have Evals. You Have One Kind of Eval.

Most teams build one corner of the eval map — offline, end-to-end, on the inputs they imagined — and call it "evals," the way you'd say "we have tests." Eval is plural. Here's the map: Hamel Husain's three levels, and the three questions that decide which kind of eval you're actually writing.

By Harsh Parmar••11 min read