Part 1
AI leverage
Are we creating conditions for success for AI tools?
Leading software organizations focus on adoption depth, system readiness, and whether organizational processes adapt. That means getting deliberate about delegation: raising autonomy step by step, handing more complex tasks to agents as the quality of your specs, context and infrastructure allows.
It also means changing what AI is used for, from following specs to solving open-ended problems and iterating on product ideas before committing to delivery. That’s a culture change as much as a tooling one, and you can tell it’s happening when engineers get pulled into product problems earlier and non-developers start testing ideas themselves.
1.1 AI adoption
As the share of AI-assisted PRs grows — 82% of merged PRs in the top quartile of organizations (p75), rising to 90% in the top decile (p90) — Q2 2026, n=352 — depth of adoption and usage patterns are the number one thing to pay attention to. Each new tool and model has its own adoption cycle, so breaking the numbers down by tool is important for seeing enablement opportunities. This section helps you understand how licenses are used, who’s active in each tool, and how AI tools are used in completed work.
| Metric | What it tells you |
|---|---|
| License adoption | Share of engineers with an assigned AI seat, overall and per tool. |
| Active usage (DAU/WAU) | Share of engineers using AI daily and weekly. |
| Share of AI-assisted pull requests | Merged PRs where AI contributed at any stage. Work where AI isn’t used can point at a readiness or autonomy gap, or at a deliberate choice. Breaking the share down by tool shows which tools are earning their seats and where enablement is worth the effort. |
1.2 AI readiness
Your agents are only as effective as the system they work in. These metrics show whether that system gives agents what they need to do the work, and whether it can take the output once they’re done.
| Metric | What it tells you |
|---|---|
| Agent configuration maturity | Share of repos scored low / medium / high on the instructions agents need: AGENTS.md, CLAUDE.md, in-repo docs and skills that tell agents how the codebase works and what “done” means. |
| CI/CD pipeline health (build time, failure rate, test flakiness) | Fast, trustworthy pipelines are the safety net that lets you trust work you didn’t watch happen. Agents also run them far more often, so slow CI now taxes every agent iteration, not just every engineer. |
| Custom agent environments | Whether purpose-built harnesses exist (e.g. Ramp’s Inspect, Stripe’s Minions) and how much they carry, measured as the share of changes created through them. |
1.3 AI autonomy
As trust in agentic output grows, more work moves to agents running at higher levels of autonomy. Start by understanding the share of completed work authored by fully autonomous agents, then add task complexity, which tells you whether you’re moving past the easy wins into harder product and engineering problems.
| Metric | What it tells you |
|---|---|
| Share of completed work delegated to autonomous agents | The signal for whether engineers trust autonomous agents enough to delegate — and whether the codebase makes that possible: 3.9% of merged PRs in the top quartile of organizations (p75), rising to 12% in the top decile (p90) — Q2 2026, n=486. Break it down further by level of autonomy, focusing on task agents, and autonomous agent loops running without additional human input. More autonomy is not always a good thing, and should be task-appropriate. |
| Complexity of delegated tasks (modeled from change size, plus survey) | Whether you’re still taking the easy wins or the work is genuinely getting harder. Track it alongside delegation failures — delegated work abandoned or rewritten by a human — which is what tells you where AI breaks down and why: readiness, culture, or the current limit of the tooling. |
1.4 AI development culture
The value of agents is limited if the process around them doesn’t change. The organization has to change too, and that’s where most of the difference between 10% and 10x comes from.
The savings usually aren’t in the step you sped up. A bug that now takes an engineer thirty minutes to fix using an agent still takes the organization ten hours: to reproduce it, loop in the CSM, discuss it in Slack, triage it, fix it, review it, merge it, and communicate with the customer.
Three recommendations for building the culture around AI development:
Prototype before you prioritize. Spend a few hours prompting an end-to-end version before the planning session, so the meeting decides what to cut, rewrite, de-risk or ship rather than what to build. Use AI in research and discovery, not only in delivery — and make prototyping available to designers and PMs as well as engineers.
Hand engineers problems, not specs. Open-ended product problems are the work now; writing a spec for someone else to implement adds a dependency without adding information. If a problem arrives already specified, someone spent time deciding things that should have been decided against a running prototype.
Build together, not in sequence. Fast prototyping makes small-group sessions far more productive than they used to be. One version of this is a campfire: three to five people from product, design and engineering working together live — with one person driving and sharing their screen, and the rest commenting and helping with review.
Part 2
Effectiveness
Are we delivering value faster, and is the pace sustainable?
It’s tempting to read the speed AI brings as progress. But while AI makes individuals, or even teams, faster at producing code, it doesn’t make an organization more effective at delivering software.
What makes an organization effective largely stays the same: build around outcomes, work fast without interruptions, and keep improving how you work.
This section covers the three areas that determine whether you’re delivering value faster and sustainably: developer productivity, quality and maintenance, and developer experience.
2.1 Developer productivity
These metrics show how quickly work moves through the system, where it slows down, and whether review is keeping up with the extra code.
| Metric | What it tells you |
|---|---|
| Throughput (combined measure of coding and issue throughput) | Signals whether you’re completing more work over time. Raw counts inflate with AI usage, so read coding and issue throughput together. |
| Delivery (WIP, cycle time, batch size) | Where work slows down. WIP often runs higher with agents, since starting work costs almost nothing. Batch size usually trends up as AI generates more code, which is the biggest confound in any before/after comparison: bigger PRs stretch cycle time and review on their own. |
| Code review (time to first review, review load per engineer, AI-assisted review share, AI findings per PR) | Review is where the added volume lands first, and where the bottleneck usually moves. AI-assisted review share is climbing fast — 88% of merged PRs in the top decile of organizations (p90) — Q2 2026, n=719 — but it complements human review rather than replacing it. Review stays a cultural practice: it’s how engineers grow and how codebase knowledge spreads. |
2.2 Quality and maintenance
These metrics show whether the faster pace is sustainable: how much engineering time goes to maintenance and rework, and which way bugs, incidents, and change failure rate are trending.
| Metric | What it tells you |
|---|---|
| Share of new feature development vs. KTLO (maintenance, firefighting and rework) | Where engineering effort actually goes (measured in FTEs). If AI is creating capacity, it should show up as a shift toward new work. A ratio bending back toward KTLO after a heavy shipping period is usually rework surfacing — check incidents and bug inflow before concluding that. |
| Bug backlog, number of incidents and change failure rate | The fast-moving quality signals: problems show up here in weeks, and in the investment balance over quarters. Worth more attention as a growing share of changes originates with agents. |
2.3 Developer experience
These metrics show how the work feels from the inside — whether engineers have what they need, how heavy the cognitive load is, and how much the system itself is getting in the way.
| Metric | What it tells you |
|---|---|
| Developer satisfaction | Survey measure spanning clarity of direction, focus, tooling, work management, and system health. |
| Cognitive load | The mental overhead of the work: context-switching, unfamiliar tools, and more code and actors in the system (agents and non-developers). |
| Organizational & system friction (ease of release, release confidence) | Perceived ease of getting code to production, and how much confidence engineers have when they ship. |
Part 3
Return on engineering
What business value does engineering investment produce?
This section connects engineering effort to business value: whether work is going towards your organization’s priorities, and what it costs to get there now that AI is part of the bill.
3.1 Business outcomes
These metrics show whether engineering effort is going toward the organization’s priorities, and whether you’re delivering on them.
| Metric | What it tells you |
|---|---|
| Investment in new feature development and productivity improvements | How much effort goes to new customer value and to internal productivity work, measured in FTEs. |
| Share of work linked to organizational initiatives | How much of what engineering does connects to an initiative (and, on the highest level, OKRs) rather than unplanned or reactive work. |
| Progress towards organizational initiatives, initiative completion | Whether stated priorities are moving and landing: steady-progress rate (is work advancing continuously, or stalling between bursts), initiative completion, and share of initiatives delivered on time. |
3.2 Engineering cost
These metrics show what engineering costs against what it produces, from the whole org down to a single unit of work.
| Metric | What it tells you |
|---|---|
| Combined engineering cost (headcount + AI) | Fully loaded headcount cost plus AI spend, in one number. |
| AI cost per FTE per month (total and share) | Whether per-engineer AI spend is justified, and where it’s concentrated. Combines license cost and model usage. Use this as input for AI budgeting. |
| Cost per unit of work (combined cost per user storyA user story. You can also look at the same metric per epic, or any other scoped unit of work., per pull request) | Total engineering cost divided by what was shipped. PRs represent coding contributions, and stories are the closest proxy for business value delivered. To identify savings, attributing cost to individual work items helps you focus on outliers: PRs or stories whose cost is out of line with their complexity, as well as engineers with the highest AI spend per unit of completed work. Newer models burn budget faster per task, so it’s worth reviewing more often whether token and model use is task-appropriate. |
If you’ve made it this far with more questions than you started with, that’s a good sign. You’re not meant to have all the answers now — they’ll come in time, and from doing the work.
And the work is simply to start — with figuring out your baseline, choosing one problem your organization is facing, and working on it. Then do it again. That’s the whole idea: build a little more, get a little better, on purpose, over and over.