When agents take over the coding work of building software, the first number most leaders look at is team size. The more useful thing to look at is what each person on the team now spends their day doing, because that is where the real change happened when we rebuilt our own teams.
We did this on our own AI transformation projects, working to real deadlines, and the change we did not anticipate was in the roles.
It helps to separate the two kinds of work a delivery team does.
-
The first is producing the work: writing the code, building the screens, and setting up the tests. This used to be most of the hours spent in a normal week.
-
The second is reviewing and directing it. Is this worth building? What does good look like here? Does what we just built actually meet that? Would a customer accept this trade-off?
In our trial runs, a working prototype that took two to three people over a week was done by one person in one to two days. What agents could not do was direct the work. Those decisions rest on knowing your business, your customers and where your company draws the line on risk.
So the people on the team moved from producing the work to directing it, and fewer of them were needed to carry the same scope.
Our designer used to draw screens one at a time. Now she writes down the rules for what a screen should look like, the spacing, the colours and how a button behaves, and sets up automated checks that catch anything an agent produces that breaks them. Her decisions now shape every screen the team ships, instead of the handful she could have drawn herself.
Our engineer used to write code. Now they decide how the pieces of the system connect, write the instructions the agents work from, and review what comes back before any of it goes near a customer.
Three principles for sizing the team
Size the team by the decisions it owns, then give each decision one owner. Count decisions rather than output: what should exist, what good quality means here, what passes review, and which trade-offs a customer would accept. That load does not shrink as fast as the production work. What changes is the speed at which those decisions must be made. When an agent can return a working build overnight, the people become the slowest part of the system, and every extra person with a say adds to the wait. With more judges, the debate drifts into domain-level detail, the bigger picture gets lost, and the team tends to settle on what everyone can live with instead of the trade-off the project needs. So keep the number of judges as small as the decision load allows. This only works if the people you keep are experienced enough, and trusted enough, to make the call within the project’s scope without escalating it.
Expect the roles to converge. When agents do the producing, specialists can write down what they know so colleagues and agents can use it: front-end design patterns, a design review checklist, the rules for staying on brand. One person’s judgment then reaches work they never touch. Narrow specialist roles collapse into fewer, broader ones. This is the change your people will feel first, and the one you have to manage deliberately.
Reviewing work becomes the most senior responsibility. Encoding a skill shares the know-how; accountability for the result stays with a named person. When producing work is quick and cheap, the hard part is knowing whether the result is right. The mistakes these tools make look plausible, complete and confident. Your most experienced people should spend their time here.
We went in expecting one person could run a whole build. That did not hold. With nobody else in the loop, a wrong call goes unchallenged, and the agents repeat it in everything they build next. What held was a small team with real structure around the agents and a named person accountable for every result. Fewer judges per team, that is, and nothing stops you running more teams.
The five roles that emerged
Each one keeps a question that stays with a person.
The product owner. Is this worth building?
Their job has always been to stop the team from going in the wrong direction. That job gets more challenging, because a team with agents can go in the wrong direction much faster than before. Most of the role now involves getting the right context in front of the team quickly and keeping it current, so that people and agents are working from the same understanding. The Product Owner shows working demonstrations rather than writing documents about them and remains responsible for the product after it ships.
The technical lead. Will this hold up?
Used to spend the day writing code. Now spends it deciding what needs to be done and in what order, preparing the work so an agent can pick it up, reviewing what comes back, and designing the structure on which the whole system rests. How much reviewing depends entirely on how well the work was broken down first. A small, clearly defined step needs almost none. A vague instruction requires thorough review because you cannot predict what the agent will hand back.
The AI Engineer. What do the agents work on tonight, and is their output good enough?
This is where software engineers and data scientists converge. The AI Engineer masters the agentic development process: defines the work for the agents, sets up the harnesses they run in, and brings the technical judgment to know when their output can be trusted.
The Design Steward. Does everything we ship look and behave like us?
Instead of designing screens one at a time, this role sets design standards and builds automated checks to catch any failures. There is a reason it matters more than it used to. AI-generated interfaces tend to look alike because they are built from the same underlying pieces. Without a defined component library of your own, everything your teams ship drifts toward looking like everyone else’s.
Tools and Standards. What guardrails and automations does every team need?
This goes beyond what a typical platform engineering team does. The role embeds best practices and guardrails for every function into the tools the company already uses, so that the checks run whether or not anyone remembers to run them.
Your organisation may not land on exactly these five. What holds is the shift underneath them. Every one of these roles shifted from producing the work to deciding what good work is and confirming that it happened.
What this asks of your team
Every role here asks someone experienced to stop doing the thing they are known for and start doing something they have never been measured on. Your designer stops producing the work they are best at. Your engineer stops writing the code that got them promoted. That is a real cost, and any plan that leaves it out has understated itself.
None of these forces a smaller headcount. The same structure runs the other way: keep all your people and take on several times the work by running more of these small teams in parallel. Add a person only when there is a new set of decisions to own.
The second cost is one we did not see coming, and we are still working through it. The routine work that used to teach junior people their craft is now done by agents. Experience no longer accumulates on its own, so growing your next set of experienced people becomes something you fund deliberately rather than something that happens.
We worked this out on our own projects, and the roles above are what we ended up with. The page sets out the full findings, including the numbers, the rules for deciding what agents can lead, and how to run the same trial in your own organisation.
Read what we learned from rebuilding our own teams. Reorganise your team in the agentic era


