Measuring AI ROI takes more than one number
Chances are your AI tool spend has gone up this year. Chances are also that someone has asked you what you’re getting for all those tokens.
In principle, this is easy:
In practice, there are two problems.
Problem 1: Gains in what?
Lines of code? People figured out in the 80’s that this isn’t a great metric to optimize alone. Reward lines of code, and you’ll start solving simple problems in complicated ways. Plus, the whole point is to solve a problem with as little code as possible. Next.
Pull requests? Better, but not all pull requests are equal. Some are one-line fixes, some are week-long refactorings. And even a ‘big’ pull request isn’t worth much if it doesn’t solve a user problem or create more revenue. A one-line fix could save a key customer account, which is far more valuable than a ‘bigger’ pull request.
Stories and epics? These are closest to user problems and business value. But they’re not a standard unit, and some work never makes it to the issue tracker. Hotfixes and small improvements often ship as PRs with no linked issue.
Flow metrics? PR cycle time, review time, deployment frequency. These are useful leading indicators. If PRs are moving faster, something is working. But they measure how smoothly work flows, not how much of it gets done or whether it mattered.
If you’ve been around this space, you know where this is going: there’s no single productivity metric in software development. And even if you found one, Goodhart’s law would ruin it for you. When a measure becomes a target, it stops being a good measure.
Developer experience? You can survey your developers, ask how they feel about the AI tools, and even turn that into a composite score. I’m convinced developer experience is worth investing in. But I can’t really tell you what a 1.4-point bump is in dollars.
Time saved? This one avoids most of the pitfalls above. If AI saves 20% of your engineers’ time, you can assume they use that time for something useful. And time is easy to measure. You have a clock.
The catch, however: to know how much time you saved, you need to know how long the task would have taken without AI. You don’t have that number, which brings us to the second problem.
Problem 2: Compared to what?
Self-assessment? This is the most common approach. After a task, developers estimate how much time AI saved them. Easy done.
The downside is that people are rather bad at this. In METR’s randomized controlled trial last year, experienced developers estimated AI made them about 20% faster. In reality, it made them about 20% slower.
Segmenting people? Compare heavy AI users, light users, and non-users. Say your heavy users ship 40% more PRs. Great, but does spending more tokens make you a more effective developer? Or do more effective developers spend more tokens? Or do more effective developers actually spend less tokens because they know which models to use? You can’t tell which way the causality runs.
Segmenting work? Compare PRs built with heavy AI use to PRs built without. Same issue. Did spending tokens on a PR make it bigger, or did you spend more because it was a harder task? You end up with an inconclusive result.
The honest answer is that there’s seldom a randomized control group, unless you set one up yourself. That’s a big project, and few organizations are willing to do it.
It’s also getting harder to find “no AI” anywhere. METR ran a follow-up study and struggled to find volunteers who would work without AI, even when they got to pick their own tasks, and even when they got paid $50 an hour.
The past? You used AI less before and more now, so you compare the two periods.
This is the most practical option, but you need to be careful with alternative explanations. Between the two periods, you probably also reorganized a team or two, hired some people, changed your development process, and had a summer holiday season. Any of those could explain a change in throughput.
Then again, AI coding tools are probably the biggest change in your development process lately. So they surely explain at least some of what you see.
One way to do it anyway
I don’t want to only complain, so here’s my best attempt at an ROI formula you can actually use:
For throughput, I wouldn’t pick just one metric. Look at lines of code, pull requests, and issues side by side, normalized per full-time engineer. This way, there are three variables instead of one.
For the baseline, I’d compare to the past. It’s not perfect science, but it’s a real number, and at least you’re aware of the flaws and caveats. You know what changed and when, so you can at least argue about what else changed.
Let’s run the numbers for an imaginary company. It has 20 engineers who cost $3M a year in total. It spends $500 per engineer per month on AI coding tools, so $120k a year. Compared to a year ago, lines of code per engineer are up 40%, PRs are up 10%, and completed stories are down 15%.
Depending on which lens you pick, AI is either the best investment the company ever made or a money pit. Which number is right? All of them, sort of. The spread itself tells you how AI is changing the shape of work: code per PR went up, so pull requests are getting bigger, and each one still needs a human to review it. PRs per story went up too, which means the constraint has moved somewhere outside of coding, like scoping or waiting on a release. So you should probably talk about it.
Turning the formula around
The before-and-after comparison has a clear flaw. Your “no AI” period might be two years ago by now. It’s hard to pick a fair time frame, and it’s hard to spot small changes in a continuous process.
So I prefer to take the same inputs and organize them differently. Instead of ROI, look at the total cost of throughput: developer cost plus AI spend, divided by throughput. Then follow the trend.
Back to our company. Last year, 4,000 merged PRs cost $3M, so $750 per PR. This year, 4,400 PRs cost $3.12M, so $709 per PR. AI spend went up, but PR throughput went up more. Cost per PR went down 5%.
Stories tell a different story. Last year, our company had 1,000 completed stories that cost $3,000 each. This year, 850 stories cost $3,671 each, an increase of 22%.
Three trend lines like this are much easier to follow than a one-off ROI percentage. And they raise better questions. Why are we merging more PRs but finishing fewer stories? Is the work getting split differently? Is more of it going untracked? Or are we just busier without getting more done?

Two caveats
Control for quality. Going faster while producing worse code is not what I’d call productivity. Keep an eye on change failure rate, bug backlog, and the share of keeping-the-lights-on (KTLO) work in your investment balance. If cost per PR drops 5% but the KTLO share climbs from 10% to 25%, the savings are borrowed, not earned.
Expect a dip first. AI tools aren’t a light switch. Many organizations see productivity go down before it goes up. People need to learn the tools, and the team needs to change its process and practices to get the gains out.

That’s why I’d hold off on the ROI calculation until you’ve figured out whether you’re set up for success at all. We just published the Build framework to help with exactly this. It has three areas:
- AI leverage. Are you creating the conditions for AI tools to work? Measured by adoption, readiness, autonomy, and culture.
- Effectiveness. Are you delivering value faster, and is the pace sustainable? Measured by productivity, quality, and developer experience.
- Return on engineering. What business value does your engineering investment produce? Measured by business outcomes and engineering cost.
Return on engineering is where the numbers we’ve just worked out together live. However, alongside cost per unit of work, we also recommend looking at business outcomes: how much you invest in new features versus maintenance, how much work links to organizational initiatives, and whether you actually finish those initiatives.
The number nobody asked for
Ok, back to the person who wants to know what you’re getting for all those tokens.
You could give them one number. 9x ROI, say. It sounds pretty impressive, and it’s easy to whack on a slide deck.
But what happens when a couple of months later that same person notices the backlog has doubled, and now your one number is wrong? And oh no, you only had one number!
This is exactly why a few numbers, tracked over time, ideally that are in tension with each other, are a much better bet.
Sure, you might have to translate engineering language into business language (like saying “we’re releasing features to customers at a lower cost than previously” instead of “cost per story is down”) but it’s better for everyone to see the tradeoffs, not just the headline.
So instead of 9x, you’d go back to that person with something like: cost per PR is down 5%, cost per story is up 22%, and here’s what we think is happening between the two. It makes for a less exciting slide, admittedly, but it’s one you won’t have to take back a couple of months later, and it gets you into a conversation about what’s actually going on with the work.
Subscribe to our newsletter
Get the latest product updates and #goodreads delivered to your inbox once a month.




