Why we created the Build framework

Roman Musatkin, Head of Product and Design · Oct 6, 2026

Depending on who you ask, AI has made software organizations anywhere between ten percent and ten times more productive.

Both answers are right. Some companies really are shipping at a pace that would have sounded made up two years ago, and others have rolled out the same tools to the same kind of engineers and are getting something closer to a rounding error. You’ll often find both inside one company, too, where one team has rebuilt how it works around agents and yet the org average barely budges.

What separates those two groups has very little to do with the model they picked, how many licenses they bought, or how good their engineers are. An individual is only ever as productive as the systems, processes, and culture around them allow them to be, and that was true long before anyone had an agent in their terminal. What AI has done is make the gap between a good system and a mediocre one very large, very quickly.

That’s the idea the Build framework is centered around: the organizations getting the 10x are building two things, the product and the system that builds the product — and they design that system on purpose.

A lot of what’s in the framework comes directly from our work with customers going through this transformation, some of whom are getting the largest gains we’ve seen anywhere, and that work is what puts us in a position to lead this change. None of it is specific to their size or their industry, so any engineering organization can learn from what they’re doing and make the same kind of progress, whether it has 50 engineers or 5,000.

The gains come from removing friction, not from the tools

In most large organizations, agents are still used almost only for writing code, and most of the effort goes into getting more people to use them. In our experience, that on its own doesn’t get you very far. Engineers write code faster, but the code then gets stuck somewhere else in the process, so the organization as a whole doesn’t get much faster until it changes how it works around the tools.

Some of what needs to change is about creating the conditions for AI in the first place, like development environments that are quick to set up, CI that’s fast and trustworthy, and codebases with enough context for agents to do useful work. For most organizations, though, the biggest constraint sits outside the AI implementation entirely, in organizational friction that builds up in the same three places: code review, the way decisions get made, and maintenance.

Code review is usually the first one people notice, because more code from more contributors, human and otherwise, ends up in front of the same reviewers. Decision-making doesn’t show up in delivery data in the same way, but it limits delivery just as much, especially when the process for deciding whether something should be built takes longer than building it. Maintenance is the one most organizations are least prepared for, and it’s where we see the biggest gap between how people talk about the problem and what’s happening in reality.

What we included in the framework, and why

Those three places are why the framework covers much more than AI. A framework built only around AI adoption, autonomy, and agent readiness could tell you how good an organization is getting at using AI, but you’d also want to know whether the organization itself has improved, wouldn’t you? This is why the Build framework has three parts, each built around its own question.

It’s also where Build differs from most of the other frameworks that have come out in the past couple of years. Many tend to describe what you can measure today. Engineering organizations, and enterprises in particular, are trying to change how they build software, and Build is about exactly that: how software gets built, how that’s changing, and whether the organization is getting better at it, including parts like product development culture that don’t fit nicely into a dashboard.

AI leverage asks whether you’re creating the conditions for AI to succeed, and it’s where most of the new material is. Adoption is part of it, but deliberately a small part. Across Swarmia customers, AI contributed to 82% of merged pull requests in the top quartile of organizations in Q2 2026, and 90% in the top decile. At Swarmia, adoption is close to 100%, with a growing share of our throughput currently going through custom agents.

At that level, another percentage point of adoption doesn’t tell you much. The more useful questions are whether repositories contain good instructions for agents, whether tests and CI can be trusted, what kinds of work teams can delegate successfully, and whether AI is used while people are still exploring a problem or only once somebody has already decided what the answer should be.

Effectiveness asks whether you’re delivering value faster, and whether that pace is sustainable. It contains the more familiar engineering measures (throughput, delivery, code review, quality and maintenance, developer experience), and it’s the part we’d tell most organizations to focus on first, because that’s where the main bottlenecks still are. Without it, whatever you invest in AI leverage doesn’t turn into gains, and once you work on it, everything else in the framework multiplies.

Return on engineering asks what business value the engineering investment produces. It connects everything else to company priorities, customer outcomes, and the total cost of engineering, because more pull requests, more stories, or cheaper units of work are outputs, and the company still needs to know whether engineering effort is going toward its priorities and whether that work is getting done.

The full Build framework page has the metrics, benchmarks, and questions under each part, and you don’t need all of them to get started: the ones that will be useful for your organization depend on what you’re trying to understand.

Maintenance and the cost of shipping more

Effectiveness is where we’d start, and within it, maintenance is the part that gets underestimated most often. The usual way to talk about AI and maintenance is in terms of rework: does the new code stay in the product, and does it cause more bugs and more incidents? Those are the right things to watch, but they miss something more basic, which is that if the goal is to ship more, you’ll get more of those things anyway, however good the code is.

One of our customers described it like this: if you used to deploy 100 times a week and have one incident, and now you’re ten times faster and deploying 1,000 times a week, you still get the same number of incidents per deployment, which means ten of them. Each one needs the same amount of human attention as before, and it’s the same team handling them. Trading speed for quality doesn’t help here, because speed and quality aren’t a dial you turn one way or the other. You have to build in a way that aims to avoid accumulating maintenance in the first place.

That means baking quality into how every change is made, with automated tests as part of every change and a round of bug fixes right after something is released. It also means knowing how much time goes to maintenance in each part of the product, and whether you’re accumulating it at an unsustainable pace or steadily working it down, which is difficult to see without data and one of the things we’ve put the most work into quantifying.

Incidents and how fast you respond to them are part of that picture, along with how much of the team’s time goes to work that doesn’t add anything new, like firefighting, cleaning up old code, or redoing things that were already done once. Look at it from the other direction and it tells you which parts of your product are unstable, because they’re the ones where the share of new work keeps shrinking.

What’s sustainable looks very different depending on who you are. Large organizations carry a legacy maintenance burden that a startup doesn’t have, and a startup can reasonably take on more maintenance in exchange for more speed, which is a much harder call for a company working to a five-year plan.

How development culture is changing

Maintenance is something you can measure. Development culture, which sits under AI leverage, is harder to pin down, and it’s where the difference between the organizations getting the largest gains and everyone else is easiest to see. AI multiplies whatever culture is already there, the good parts and the bad ones, so the same tools that make a team with quick decisions and close customer contact much faster will also help a team that waits on approvals produce more work that waits on approvals.

The first thing we see changing is the line between product and engineering. In the organizations furthest along, engineers are taking on more of the open-ended product problems alongside architecture and systems, while product people spend less of their time on roadmaps and stakeholder management and more on making decisions quickly, bringing in customer context, and keeping the long-term vision clear. Product managers and designers are increasingly making working changes in the code themselves, too.

The second is the order in which things happen. When building a working version of something takes less time than deciding whether to build it, the decision should be made with the working version in front of you, and more teams are starting to prototype before they prioritize, building an end-to-end version they fully expect to throw away and using it to decide what to take further. We’ve seen the same thing in our own product meetings, where a discussion about whether something was a good idea and how much effort it would take ended with someone building it before the meeting was over.

The third is what gets handed to engineers. A spec written before anyone has tried the solution in a working product often locks in decisions that nobody has had a chance to test. Handing engineers the customer problem, the context, and the constraints leaves room for them and their agents to find a better answer, and sometimes that answer is much bigger than the original request: when one of our engineers was asked to build an SSO integration, they looked at the problem from a different angle and realized they could solve all of our SSO integrations at once.

And more of the work happens together rather than in sequence. Campfires, the format Anthropic uses where a few people from product, design, and engineering build the same thing at the same time, are one version of this. They’ve also turned out to be one of the best ways for a team to learn how to work with agents, because prompting techniques and ways of structuring agent work are much easier to pick up by watching someone than by reading about them.

AI ROI is one part of the picture

The third part, return on engineering, is where AI spend comes in. It’s large enough now that nobody can ignore it, and whether it’s paying off comes up in almost every conversation we have with engineering leaders, but AI ROI is only one part of what return on engineering covers.

The most basic question is whether people are using the licenses you pay for, and if they aren’t, you cut them. Beyond that, a useful approach is to look at cost per unit of work and how it’s changing over time, with headcount cost and AI cost considered together, and then to look at whether AI use is task-appropriate: whether some people or teams are using a disproportionate amount of AI for the kind of work they’re doing. This doesn’t only apply to people, either: repositories without AGENTS.md files, or skill files that are inefficient (with a scope that’s too broad, for example), can drive up AI spend significantly, and fixing them is one of the bigger cost levers you have.

What you want to build is a loop that keeps finding the places where usage isn’t task-appropriate, regardless of the author — like spending $1,000 on tokens for a problem that would have cost less to do manually, could have been handled by a different model for a fraction of the price, or didn’t need to be done at all.

Cost is only part of return on engineering, though. You can ship a lot more with AI, but it doesn’t count for much if it isn’t helping the business grow, so you also need to connect engineering work back to business outcomes: where effort goes, whether it’s going toward the company’s priorities, and what that work costs with headcount and AI together. That’s what Swarmia’s effort model is for. It’s our own model, and a unique part of Swarmia’s data platform: it builds a picture of each developer’s effort from their actual work, with every number traceable back to the work and contributors behind it, and it’s the part Swarmia does better than anyone else in this space.

What progress looks like

Put together, the three parts describe an organization that’s getting better at building software, and there’s no single score for that. Progress looks more like a set of practical characteristics developing together:

  • Teams can turn an assumption into something testable within a day.
  • Work goes through reliable pipelines in small enough batches that one change doesn’t occupy the whole system.
  • Agents have enough context and clear enough boundaries to take on work that somebody needed done, and do it well.
  • Code review continues to protect quality and spread knowledge without becoming an ever-growing queue.
  • Maintenance is visible per product area, and it isn’t accumulating faster than teams can work it down.
  • Leaders can see where engineering time and AI spend go, and teams understand how their work connects to company priorities and customers.

Underneath all of these is one capability that’s harder to put in a list, which is how quickly the organization can adapt to change. Getting the benefit from these tools depends on being able to change the way you build, again and again, without a huge fuss, because almost everything else in the framework reflects the tools and ways of working we have today.

The tools will be different next year, models will improve in some areas and disappoint in others, and new forms of agent work will create problems that don’t have names yet, while some of today’s concerns turn into routine infrastructure questions.

We built the Build framework on the same assumption. It describes what engineering organizations need to focus on now, based on what we’re seeing with the organizations leading this change, and we’ll keep revising it as the tools and the work change.

See the full Build framework
AI leverage, effectiveness, and return on engineering, with the metrics, benchmarks, and questions to understand whether your organization is getting better at building software with AI.
Read the framework
Roman Musatkin
Roman Musatkin heads up product and design at Swarmia. He's known for having the perfect balance of ambitious product vision, wild ideas, and a relentless commitment to thoughtful design.

Subscribe to our newsletter

Get the latest product updates and #goodreads delivered to your inbox once a month.

More content from Swarmia
Otto Hilska · Nov 9, 2021

Measuring software development productivity

The software world gave up too soon on measuring development productivity, deeming it impossible. A few years ago a new wave of research arrived that proved otherwise. After discussing the foundations for this measurement approach, this article will share some practical tips for instrumenting a…
Read more→
Pinja Dodik · Sep 21, 2026

Inside an AI-native engineering organization

Nobody here writes that much code by themselves anymore. That’s Juho Eräste, co-founder and CTO of Taito.ai , and former software engineer at Swarmia, opening a demo he gave our engineering team. He’s about to show us how his company builds software, so he shares his screen. He doesn’t open a code…
Read more→