Before you give AI agents any work, ask these 3 questions

Content

Most organisations have already bought the tools, and their teams are actively using them. Individuals appear noticeably more efficient. However, it’s still unclear whether the organisation as a whole is delivering better software sooner.

The main cause is often that no one has determined which tasks should be handled by agents and which by people. When the tools get implemented, everyone uses them on whatever is available at the moment, resulting in different answers depending on the person and the day.

We rebuilt our own delivery teams around AI agents on our AI transformation projects, and that showed us where the approach failed. What we took away was a simple rule for deciding who leads a piece of work, built on three questions. An agent leads only when all three hold.

1. Has anyone decided what the job is?

An agent needs to know exactly where a job starts and stops, and it needs a test it can use to evaluate against its own work. A task with a settled scope, one outcome and no open decisions meets that standard, and an agent can lead it, with a named person still reviewing what comes back.

If nobody has yet decided precisely what the job is, a person will have to lead. This is where costly failures begin, because the gaps are not visible at the time. Give an agent an incomplete plan, and it will not stop to ask. It fills the gaps with something plausible and returns what appears to be a finished result.

2. Are the requirements and systems readable?

An agent can act only on what it can read, and it needs two things to be legible: your request and the system it is working in. The less clear of the two sets the ceiling.

A precisely written requirement for a codebase nobody has documented in four years will struggle. A well-documented system will struggle just as much if the instructions are vague. Leaders tend to focus on the first issue and overlook the second, even though documentation is often the bigger constraint, and the one only you can fix.

3. Are mistakes cheap?

When an error stays bounded within a sandbox or a pull request and cannot reach production on its own, an agent can run for a long time before a person needs to step in. When a mistake could reach a customer, a regulator or the brand before anyone catches it, a person leads, however clear and readable the task is.

The same work can fall on either side depending on what it connects to. A change to a reporting screen and a change to a payment flow may look identical in the backlog and belong on opposite sides of this line.

What the answers tell you about your own organisation

Apply the three questions to the work one of your teams has queued up for the next few months. What comes back is a count of how much of that work an agent cannot yet lead, with a specific reason for each piece.

That count describes your organisation more than the tools you’ve implemented. Two companies with the same tools and the same backlog will get different answers because one has documented its systems, contained its mistakes and trained its people to work this way, and the other has not. Any company can buy the same tools. What decides the outcome is whether your people change how they work.

We rebuilt our own delivery teams to find this out and measured what happened. Read what we learned from rebuilding our own teams, we set out the full rule, where the bottleneck is located, and how to run the same trial in your own organisation. Reorganise Your Team in the Agentic Era

Recommended articles