AI-assisted delivery
AI is already writing your code. This is the gate it writes inside.
Google’s DORA study of AI-assisted teams found the pattern: throughput up, stability down. Code arrives faster than review, testing and deployment can absorb it. The roadmap moves and the release calendar slips in the same quarter. We put the gates in first, then roll the tools out behind them, so your team keeps the speed and you get the evidence that the code still holds.
At a glance
- Length
- Six to twelve weeks
- Format
- Pairing on real tickets, not training
- Baseline
- Your four DORA numbers, before and after
- We step back
- When the numbers hold without us
What it covers
Gates first, then tools.
The gates, before the tools
Tests, coverage thresholds, mutation testing on changed code and a code health score that runs on every change. AI can draft whatever it likes. What survives is decided by the pipeline, not by a tired reviewer.
The tools, rolled out properly
Claude Code, Cursor or Copilot set up on your repositories with your context, your conventions and your guardrails committed alongside the code. Configured once for the team rather than twenty five times in private.
Evaluations for what you prompt
The prompts, skills and agents your team writes are code too. They get a test suite, so a change that improves one case and quietly breaks three others is caught by the suite rather than by a customer.
A governance story you can hand over
What the tools may touch, where the code goes, what is logged and who approved it. Written for the audit and the procurement questionnaire, not for a slide.
How we do it
From the baseline to the same numbers again.
01
Read the repository first
A diagnosis of where delivery actually stalls, and a baseline of the four DORA numbers. Without that baseline there is nothing to compare against when the tools go in.
02
Put the gates in
The pipeline gets the tests, the thresholds and the health scoring before a single licence is handed out. Speed without a gate is the thing that broke the stability figure in the first place.
03
Roll out, paired not lectured
Our engineers sit with yours on the tickets already in the sprint. The conventions that work get committed to the repository so the whole team inherits them.
04
Re-read the numbers, then leave
The same four DORA figures against the baseline, plus the health trend. We step back when they hold without us.
What you get
Running on your pipeline, owned by you.
- A baseline of your four DORA numbers, and the same numbers after
- Quality gates running on every change, on your own pipeline
- AI coding tools configured on your repositories, with conventions committed
- An evaluation suite for the prompts, skills and agents your team wrote
- A written policy on what the tools may touch and what is logged
Questions
Questions we get asked.
- Our developers already use Copilot. What does this add?
- Almost certainly the gates, and the evidence. Most teams adopt the tool and keep the review process they had before. That is why throughput rises and stability falls. This puts the mutation gate, the coverage threshold and the health score in the pipeline. Whether the generated code is good enough stops being a judgement call at merge time.
- Is this the same as Practices coaching?
- No, though they end up in the same place. Practices coaching is for a team without test driven development and continuous delivery yet. This is for a team already shipping fast with AI, watching the change failure rate go the wrong way. If you are not sure which you are, the repository diagnosis at the start of either will tell you.
- Which tools do you roll out?
- Whichever you have bought, or whichever the diagnosis suggests. We use Claude Code, Cursor and Copilot ourselves and have no commercial interest in any of them. The gates matter far more than the choice, and they are the same gates whichever tool your team types into.
- Do you need access to our codebase?
- For the diagnosis, read access to the repository and its history. For the rollout, the same access any engineer joining your team would get. Everything runs on your own pipeline and your own accounts, and nothing is copied out.
What clients say
In their words.
01 / 04
Before Metricsense, checking calls was a manual process: we selected files for review. Now the platform analyses every call, and we review by risk. It has identified customer experiences we would not have picked up before, and it lets us scale our quality assurance program without increasing headcount, with our team still the human in the loop. It helps us protect our customers, our business and our agents.
Russell Cain
Founder and CEO, Life Insurance Direct
Also under product development
Where it leads next.
Practices coaching
Coaching your engineers in TDD and continuous delivery until they run it.
Platform and DevOps
CI/CD, observability and infrastructure that is ready for the day the traffic arrives.
Legacy modernisation
Rebuild without stopping the business.
Not sure this is the right place to start?




