Articles.
Notes from the work: product thinking, engineering quality, and what AI tools do well and what they leave to you.
- articles
- 59
- series
- 3
- topics
- 5
All articles
Page 2 of 7
AI engineering
Evals Aren't a Benchmark Suite. They're a Habit of Looking at Your Data.
Most teams buy an eval tool, run it once, and call it done. They confuse benchmarks with evals — and ship AI that confidently produces wrong outputs no one will catch. The work the tool can't do for you is the work that matters: looking at your traces, naming the failure modes, and writing assertions that fire when they recur.
Harsh Parmar12 min read
AI engineering
AI Is Making Your Junior Engineers Worse At Their Jobs
Your juniors are shipping more code and learning less from it. The agent answers the question before they form the question. The debugging muscle never gets built. The cost lands eighteen months from now when those juniors are the mid-levels and nobody can reason from first principles.
Vishvjitsinh Vanar11 min read
Engineering practices
The Refactoring That Never Made It Into the Sprint
Someone wrote "we should refactor this" in a PR comment nine months ago. Three more features have been built on top of it since. The comment is still there. The code is still there. This is not a discipline problem. It is a system that produces the wrong outcome reliably.
Shivani Sutreja5 min read
Metrics and code health
God Classes, Circular Dependencies, and the Architecture That Kills Your Lead Time
Your lead time is not slow because of code review or approvals or sprint ceremonies. It is slow because three files do everything, four modules depend on each other in a circle, and every meaningful change has to pass through the same engineer who understands the wiring. Process is the symptom. The architecture is the disease.
Vishvjitsinh Vanar13 min read
How we build
CODEOWNERS Was Built for Humans You Could Trust. Your Committers Are Now Agents.
CODEOWNERS, blame, and branch protection were built around named humans with reputation. When the committer becomes an agent, the trust scaffolding stays standing on a foundation that quietly disappeared — and the load migrates to a single reviewer's signature.
Harsh Parmar12 min read
How we build
Evals for Engineering Agents: How We Test the AI That Tests Your Code
Your engineering agent reviews PRs, writes tests, and ships code. Nobody can tell you whether it is better than rubber-stamping. Without evals on the agent itself, you are flying blind on the tool that is making 40% of your engineering decisions.
Vishvjitsinh Vanar9 min read
AI engineering
The Half-Life of AI-Generated Code
AI code passes review on day one and ages worst by month six — not because it's wrong, but because the design intent that makes it refactorable was never part of the diff. The fix is encoding intent in artifacts the next iteration must read.
Harsh Parmar15 min read
Metrics and code health
The Onboarding Clock: What Your New Engineer's First Week Reveals About Your Codebase
The number of days between a new engineer's start date and their first feature in production is the most honest readout of your codebase's health. Every friction point they hit — environment chaos, unclear boundaries, noisy tests, PR limbo — is a structural problem your team learned to route around. The onboarding clock is not running on your new engineer. It's running on your codebase.
Shivani Sutreja4 min read
Engineering practices
Your Services Talk to Each Other. Nobody Wrote Down What They Agreed To.
When two services integrate, an implicit agreement is formed — a field name, a response shape, an error code. That agreement was never written as a test. Both services keep shipping. The agreement does not update. Contract testing makes the assumption explicit, verifiable, and enforced by CI before a breaking change reaches production.
Shivani Sutreja7 min read
Working through one of these problems?