Harsh Parmar

Insights on product thinking and engineering excellence • 17 articles published

A pyramid of three evaluation levels — unit tests at the base, human and model eval in the middle, A/B testing at the top — beside three cross-cutting axes: offline to online, reference-based to reference-free, end-to-end to per-component

You Don't Have Evals. You Have One Kind of Eval.

Most teams build one corner of the eval map — offline, end-to-end, on the inputs they imagined — and call it "evals," the way you'd say "we have tests." Eval is plural. Here's the map: Hamel Husain's three levels, and the three questions that decide which kind of eval you're actually writing.

By Harsh Parmar••11 min read
Cover for "The Half-Life of AI-Generated Code" — a Mineral Green canvas with a decay curve showing two trajectories from the moment of merge: an AI-written module with no encoded intent decays steeply, while an AI-written module with tests-as-spec, decision notes, and fitness functions stays flat

The Half-Life of AI-Generated Code

AI code passes review on day one and ages worst by month six — not because it's wrong, but because the design intent that makes it refactorable was never part of the diff. The fix is encoding intent in artifacts the next iteration must read.

By Harsh Parmar••15 min read
Diagram showing what an AI agent actually reads from a repository — code and tests — versus the artifacts of intent that never reach it

Your Tests Are the Only Spec AI Reads

When an AI agent works on your codebase, your test suite is the only artifact of intent it consistently sees. Tickets, Notion pages, and Slack threads do not enter the loop. If your tests describe what the system should do, the agent is constrained. If they only describe what the code already does, the agent ships whatever it generated.

By Harsh Parmar••11 min read
Two columns. Left column labelled "the thinking — keep in-house" lists specification, architecture, test design, critical review. Right column labelled "the typing — delegable" lists scaffolding, boilerplate, mechanical refactors, glue code & migrations. A thin coral line labelled "the harness" runs between them.

Don't Outsource the Thinking

AI agents and offshore teams break in exactly the same way. They are execution accelerators, not thinking substitutes. The work that doesn't compress — specs, architecture, test design, critique — is the work you have to keep in-house. Outsource the typing. Keep the thinking.

By Harsh Parmar••11 min read