Articles.
Notes from the work: product thinking, engineering quality, and what AI tools do well and what they leave to you.
- articles
- 59
- series
- 3
- topics
- 5
Series
Read a topic end to end.
Multi-part deep dives, written in order. Start at part one.
All articles
Page 1 of 7
AI engineering
Replace old tool results with a placeholder, and the context stops rotting
Our data agent got worse the longer a conversation ran. No prompt change fixed it. When we profiled the context window, stale tool output was eating most of it. We replaced each old tool result with a one line placeholder and cut input tokens by 38% on the first turn and 37.6% across a 25 step run.
Chirag6 min read
AI engineering
Give the model buttons, not a database
Ask an AI model for a figure it does not have and it will give you one anyway, fluently, and wrong. If your name goes on the numbers, that is not a risk you can prompt away. Here is the pattern we use so an assistant cannot invent a figure, explained from the start, and the climate data publisher it runs for today.
Chirag6 min read
AI engineering
Three Questions I Ask Before Trusting AI Output
AI generates answers in seconds. Confidence is not accuracy. Three questions that separate effective AI users from ones who amplify mistakes.
Vishvjitsinh Vanar7 min read
AI engineering
Error Analysis Is the Eval Work. Here's How to Actually Do It.
Everyone agrees you should "look at your data." Then they open a hundred traces, scroll for ten minutes, feel vaguely worried, and reach for a tool. The looking has a method — open coding, then axial coding into a ranked taxonomy — and the taxonomy, not your assumptions, is what decides which evals to write.
Harsh Parmar12 min read
Engineering practices
The Test That's Hardest to Write Is Telling You Something About Your Design
When a test requires 40 lines of setup and three mocks for four lines of assertion, engineers blame the test. The test isn't the problem. It's the first thing honest enough to say the design is wrong.
Shivani Sutreja9 min read
AI engineering
You Don't Have Evals. You Have One Kind of Eval.
Most teams build one corner of the eval map — offline, end-to-end, on the inputs they imagined — and call it "evals," the way you'd say "we have tests." Eval is plural. Here's the map: Hamel Husain's three levels, and the three questions that decide which kind of eval you're actually writing.
Harsh Parmar11 min read
Engineering practices
"Done" Doesn't Mean It Works. It Means Someone Else Can Change It.
The feature shipped. Tests passed. Ticket closed. Three months later someone needs to change it — and the original author is the only person who understands it. That is not done. That is a liability with a time delay.
Shivani Sutreja10 min read
AI engineering
You Can't Look at Data You Don't Have: Building Your First Eval Set
The first eval post said pull a hundred production traces and read them. But what if you haven't launched? Here's how to build your first eval set from dimensions, scenarios, and synthetic data — and why that set is scaffolding, not the building.
Harsh Parmar11 min read
Engineering practices
The Flaky Test Is the Most Expensive Test You Have
Your CI went red. You clicked Retry. It went green. You merged. This happens dozens of times a month on most teams. Nobody is counting. The cost is not the retries — it is what the retries teach your team about what red means.
Shivani Sutreja8 min read
Working through one of these problems?