Engineering practices

What is AI-assisted engineering? A guide to shipping faster without losing the quality

By Chirag••14 min read
A deep teal cover. On the left, the title "What is AI-assisted engineering?" above the line "shipping faster is the easy half". On the right, four delivery metrics with throughput rising and change failure rate rising with it

AI-assisted engineering is the practice of having models write production code inside controls that decide whether it is safe to ship.

The tools are the easy half. The practice is everything that has to absorb a higher volume of change: testing, review, deployment, and the architecture underneath.

It helps to be precise about what is new. Software teams have always had more ideas than capacity, so the constraint was how fast code could be written. That constraint is largely gone. What replaced it is the constraint that was always behind it: how fast code can be verified.

Key takeaways

  • This exists to settle real decisions. Do we roll assistants out, to which teams, behind what controls, and how will we know if it went wrong
  • The research is not ambiguous. AI adoption now correlates with higher throughput and, separately, with higher instability. Both at once
  • AI-written code fails differently. It is fluent, so it passes the kind of review that catches code that looks wrong
  • Measure all four delivery metrics or none. Deployment frequency alone will tell you the rollout succeeded while production gets worse
  • The gates come before the tools. A team that cannot safely absorb more change does not get safer by producing more of it

In this guide

  1. What is AI-assisted engineering, and what actually changed
  2. What does the research say
  3. Why does AI-written code fail differently
  4. What should you measure
  5. What has to exist before the tools go out
  6. How do you roll them out
  7. What does this look like in production
  8. What tools do teams use
  9. What do most teams get wrong
  10. How long does this take

1. What is AI-assisted engineering, and what actually changed

Three things sit under the term, and they are different jobs.

Completion. The assistant suggests the next few lines while you type. Low risk, and the smallest share of the benefit.

Generation. You describe a change and the model writes it. The output is a diff a person reviews.

Agentic work. The model plans, edits several files, runs the tests, reads the failures and tries again. Nobody watched the middle of it.

The risk rises steeply across those three, and so does the value. Most teams are somewhere between the second and the third, and most of their process was designed for a world with only the first.

What changed is not the writing. It is the ratio. One engineer can now produce several engineers' worth of diff. Everything downstream of writing was sized for the old ratio: how many pull requests a reviewer reads in a day, how long the suite takes, how often you deploy, how much a release contains.

Raise the input and the bottleneck moves to whichever of those is weakest. That is the whole problem, and it is not a problem about models.

2. What does the research say

Google's DORA programme has surveyed software delivery for over a decade. Its 2025 report focuses on AI-assisted development.

The headline numbers:

  • 90% of respondents use AI at work
  • More than 80% say it has increased their productivity
  • 30% report little or no trust in the code it generates

That last one deserves a moment. Nearly a third of people using these tools daily do not trust what comes out, and are shipping it anyway. That is not a tooling problem, it is a controls problem.

On delivery, the 2025 report found two things at once.

Throughput improved. AI adoption now correlates positively with software delivery throughput. This is a change from the previous year, when the relationship was not positive. Teams are genuinely shipping more.

Stability did not. AI adoption continues to correlate with instability: more change failures, more rework, and longer to resolve issues.

DORA's own framing of the cause is worth quoting rather than paraphrasing: without robust control systems, such as strong automated testing, mature version control practices and fast feedback loops, an increase in change volume leads to instability.

So the finding is not "AI makes things worse". It is that AI raises change volume, and change volume finds whatever is weakest in your delivery system. If you had a slow flaky test suite before, you now have one failing more often.

3. Why does AI-written code fail differently

Three ways, and the first is the one that catches teams out.

3.1 It is fluent

Code a person wrote badly usually looks like it. Odd naming, inconsistent structure, a comment that trails off.

Code a model wrote badly reads like a senior engineer wrote it. Consistent naming, tidy structure, sensible comments, plausible abstractions.

Review is pattern matching against what wrong code looks like. This does not look wrong.

3.2 The intent is not in the diff

When an engineer writes a non-obvious thing, the reasoning exists somewhere: in their head, in the ticket, in a comment, in the conversation at review.

When a model writes it, there is often no reasoning anywhere. The abstraction is there because that shape appears frequently in training data, not because this problem needed it.

Six months later, nobody can tell you why it is like that. The code has no author to ask.

3.3 It amplifies what is already in the codebase

A coding agent writes more of whatever is already there. Give it a codebase with clear boundaries and good tests and you get more of that. Give it one with tangled dependencies and decorative tests and you get more of that, faster.

The codebase is now an input to the tooling, not just an output of the team. That is new, and it is why two teams buy the same assistant and get opposite results.

3.4 It changes who learns

This one is about the team rather than the code, and it shows up last.

An engineer becomes senior by making decisions and living with them. Choosing the wrong abstraction, feeling the cost six months later, and choosing differently next time. That loop is how judgement forms.

When the model proposes the abstraction and the tests pass, the loop does not run. The code ships, the engineer moves on, and the thing that would have taught them never happens.

The effect is small on anyone who already has ten years of those decisions behind them. It is not small on someone in their second year, and it compounds quietly, because nothing about their output looks wrong.

Recall the DORA number: 30% report little or no trust in the code the tools generate. Working all day with something you do not trust, and shipping it anyway, is corrosive in a way no delivery metric captures.

Two things help, and neither is a tool. Review that asks why rather than whether, so the reasoning gets said out loud by somebody. And deliberately leaving some work unassisted, chosen because it teaches something rather than because it is hard.

4. What should you measure

Four metrics, and the trap is measuring half of them.

MetricTells you
Deployment frequencyHow often you ship
Lead time for changeHow long from commit to production
Change failure rateWhat share of releases cause a problem
Time to restoreHow long to recover when one does

The first two move first and move fast when assistants land. They are also the two that get reported upward, because they look like progress.

The second two are what the throughput cost. A rollout that doubles deployment frequency and doubles change failure rate has not improved anything, and a dashboard showing only the first half will say it has.

Take the baseline before the tools go out. After the rollout there is no way to construct it, and no way to answer the only question leadership will ask, which is whether it worked.

Two more worth having alongside:

Code health, scored continuously rather than audited occasionally. If agents amplify what is there, you need to know what is there.

Review latency. When generation outpaces review, the queue is the first thing to break, and it breaks quietly. Time from pull request opened to first substantive review is the number to watch.

5. What has to exist before the tools go out

Four things. A team missing any of them does not get safer by producing more code.

A test suite that fails for the right reason. Not coverage. Coverage says lines ran. The question is whether a test fails when the behaviour it claims to protect is broken. Mutation testing answers that directly, and most suites score worse than their owners expect.

A pipeline that can reject work without a human. If the only thing standing between a bad change and production is somebody noticing, that person is now reviewing several times the volume.

Small batches. The blast radius of a mistake is the size of the change it arrived in. Higher volume with unchanged batch size means bigger, more frequent, harder to diagnose failures.

Architecture with findable boundaries. An agent works well in a codebase where the right place to make a change is discoverable, and badly in one where it is not. This is the slowest to fix and the most decisive.

The order matters and teams get it backwards. The tools are procurement and take a week. The gates are engineering and take a quarter. Doing the quick one first is how throughput and instability rise together.

6. How do you roll them out

6.1 Read the repository first

Before any rollout, find out what the agents will be amplifying. Test quality, dependency structure, where change actually concentrates. This tells you which parts of the codebase are safe to point tools at and which are not yet.

6.2 Take the baseline

All four delivery metrics, plus review latency. A fortnight of data before anything changes.

6.3 Start where failure is cheap

Internal tools, test code, migrations, documentation. Not the payments path. Teams reach for the highest value area first and it is the worst place to learn.

6.4 Pair, do not train

A workshop produces people who have seen it done. Pairing on real tickets produces people who have done it. The difference shows up about three weeks later, when the workshop cohort has quietly reverted.

6.5 Write down what the tools may and may not do

Which repositories, which data, what has to be reviewed by a person regardless, what gets logged. Teams adopt these tools whether or not there is a policy. The only choice is whether the policy arrives before or after the incident.

6.6 Re-measure, and be willing to read it honestly

Same four metrics, same method. If change failure rate rose, that is the finding. The response is to fix the gate it exposed, not to argue with the number.

7. What does this look like in production

A property platform rebuilt while serving millions. We have engineered with Australia's number three property portal since 2019. In 2023 the whole platform was rebuilt on a new design in under four months with the old one still serving users throughout.

The thing that makes that possible is unglamorous and it predates AI. Visual regression and test automation in the pipeline, open source application monitoring rather than a paid black box, alarms and health checks around the clock, and over six hundred production deploys since January 2024.

When assistants arrived, none of that had to change. That is the point. A team with the controls already in place absorbs the extra volume. A team without them discovers what it was missing.

A support agent rebuilt after the first attempt failed. We built a voice agent for an HR firm, then threw it away and rebuilt it. The first version categorised the call and then asked that category's questions. Tidy, and wrong: real callers give details before they reach the subject and change direction halfway.

That was found because the output was measured by the people who used it, not by us. Their specialists flag a bad summary in one field, and that field is the accuracy number. Seventy of the first seventy-one tickets needed no correction.

The general point transfers. Whatever writes the code, somebody downstream has to be able to tell you whether it was right, in a way that produces a number.

8. What tools do teams use

Four categories, and most products sit in more than one.

In-editor assistants. GitHub Copilot, Cursor, Windsurf, JetBrains AI. Completion and generation inside the editor.

Agentic tools. Claude Code, Codex, Devin, Aider. These plan and edit across files and run commands. The risk and the value both sit here.

Review and quality. CodeRabbit, Greptile, Graphite, Qodo for review. SonarQube, CodeScene and mutation testing frameworks such as Stryker and PIT for the health of what lands.

Delivery measurement. DORA metrics tooling, whether that is LinearB, Swarmia, DX, or something built on your own pipeline data.

The order to buy in is the reverse of the order teams buy in. Measurement first, because without it no later decision can be evaluated. Then review and quality, because that is the capacity that gets overwhelmed. Then the assistants, which are the part everyone wants to start with.

This section dates fast. The four categories will not.

9. What do most teams get wrong

Rolling out before measuring. Then the honest answer to "did it work" is that nobody knows.

Counting lines or acceptance rate. Neither says whether anything got better. Accepted code that gets rewritten next sprint counted as a success.

Treating review as the gate. Review is a human reading at human speed, and the volume just multiplied. It is a filter, not a gate. A gate is something that fails without anyone being available.

Pointing agents at the worst code first. The legacy module is where the value seems highest and where the agents are least able to help, because it has the fewest boundaries and the weakest tests.

Banning the tools. The DORA numbers put adoption at 90%. A ban produces unmanaged use, no logging, and no policy.

10. How long does this take

Six to twelve weeks for a team that already has reasonable practices. Longer where the test suite or the architecture has to be worked on first, and that is usually what decides it.

A workable shape:

WeeksWhat
1 to 2Read the repository, take the baseline, agree the policy
3 to 6Gates first. Tests that fail correctly, a pipeline that can reject, smaller batches
5 to 10Roll the tools out by pairing, starting where failure is cheap
9 to 12Re-measure, widen what worked, write the governance down

The gates overlap the rollout deliberately. Waiting for perfect controls is its own failure mode, and the team is already using the tools regardless.

Where to start

Take the baseline this week. Four metrics, a fortnight of data, before anything changes. It costs almost nothing and it is the only thing you cannot reconstruct later.

Then look at the test suite honestly: not coverage, but whether a test fails when the behaviour breaks. That answer decides whether the rest is a rollout or a rescue.

See how we build AI-assisted delivery into a team.

Frequently Asked Questions

What is AI-assisted engineering?

AI-assisted engineering is the practice of having models write production code inside controls that decide whether it is safe to ship. The tools are the easy part. The practice is the testing, review and deployment discipline that has to absorb a higher volume of change without the failure rate rising with it.

Collapse

Does AI actually make teams ship faster?

Expand

Why does AI-written code fail differently?

Expand

What should we measure?

Expand

What has to be in place before rolling the tools out?

Expand

Share this article

Help others discover this content

TwitterLinkedIn
Categories:Engineering practices